H Company Holo · Flagship · 2026-08-13

Holo3 122B A10B vision benchmark

How Holo3 122B A10B performed on screenshot-based UI grounding, element-presence, and visible layout-defect tasks relevant to test automation.

Exact API modelholo3-122b-a10b

Requests were sent directly through the H Company Models API with structured output and thinking disabled. Fallbacks were disabled for every retained request.

Grounding hit rate92%
Mean box IoU0.478
Presence accuracy100%
Defect accuracy58%

Overall result

Holo3 122B A10B across all three UI tasks

Holo3 122B A10B hit 22 of 24 requested UI targets. Its balanced accuracy was 100% for element presence and 58% for layout-defect detection.

This dated pilot retained 24 grounding responses and 24 detection responses. No detection requests failed. Results describe this exact API model and test corpus, not general model quality or a consumer chat product.

Check out the complete LLM vision benchmark →

Relative performance

Where Holo3 122B A10B sits among the tested models

Each scale compares the selected model with the strongest and weakest other model. All three metrics come from the same completed 24-case grounding run; stronger is always to the right.

Box IoUHigher is stronger
0.478
Weakest other
Claude Sonnet 50.043
Strongest other
GPT-5.6 Luna0.771
SpeedLower latency is stronger
1.68 s
Weakest other
Gemini 3.1 Pro Preview11.89 s
Strongest other
Grok 4.51.67 s
CostLower total cost is stronger
$0.0126
Weakest other
Claude Opus 5$0.2474
Strongest other
GPT-5.6 Luna$0.0057
Check out the complete LLM vision benchmark →

UI grounding

Can Holo3 122B A10B locate actionable elements?

Twenty-four targets span six deterministic desktop and mobile screens. A schema-valid point inside the requested target counts as a hit; box IoU measures localization precision.

commerce catalog desktop / clean / grounding / add productPassedHit0.7573.56 s$0.00071
commerce catalog desktop / clean / grounding / searchPassedHit0.5721.73 s$0.00070
commerce catalog desktop / clean / grounding / catalog queryPassedHit0.4631.69 s$0.00072
commerce catalog desktop / clean / grounding / stock filterPassedHit0.1251.65 s$0.00071
analytics dashboard desktop / clean / grounding / search fieldPassedHit0.4461.66 s$0.00071
analytics dashboard desktop / clean / grounding / compare toggleInvalid responseMiss0.0001.70 s$0.00072
analytics dashboard desktop / clean / grounding / audience tabPassedHit0.4121.76 s$0.00072
analytics dashboard desktop / clean / grounding / revenue cardPassedHit0.3861.63 s$0.00070
iot control desktop / clean / grounding / auto modePassedHit0.2162.51 s$0.00072
iot control desktop / clean / grounding / devices navPassedHit0.7441.83 s$0.00072
iot control desktop / clean / grounding / pump cardInvalid responseMiss0.0001.65 s$0.00069
iot control desktop / clean / grounding / restart devicePassedHit0.8391.65 s$0.00071
banking ios / clean / grounding / transferPassedHit0.8811.44 s$0.00034
banking ios / clean / grounding / profilePassedHit0.4591.47 s$0.00035
banking ios / clean / grounding / transaction searchPassedHit0.2571.42 s$0.00032
banking ios / clean / grounding / hide balancePassedHit0.1811.38 s$0.00031
flight pass ios / clean / grounding / trips tabPassedHit0.2671.43 s$0.00035
flight pass ios / clean / grounding / boarding passPassedHit0.5771.37 s$0.00031
flight pass ios / clean / grounding / check inPassedHit0.8711.41 s$0.00034
flight pass ios / clean / grounding / profile tabPassedHit0.7371.46 s$0.00037
smart home android / clean / grounding / thermostat cardPassedHit0.6871.47 s$0.00033
smart home android / clean / grounding / add devicePassedHit0.8811.47 s$0.00034
smart home android / clean / grounding / morePassedHit0.4611.46 s$0.00033
smart home android / clean / grounding / room searchPassedHit0.2471.46 s$0.00034

Detection results

Presence and visible layout defects

Separate balanced tasks measure whether the model rejects absent elements and distinguishes clean controls from screenshots containing one seeded layout defect.

Can the model tell when an element is not there?

Six present and six absent target queries per model. This track measures classification only.

Balanced accuracy100%
Present recall100%
Absent specificity100%
Coverage12/12
Expected
commerce catalog desktop / clean / element presence / presentValidCorrectPresentPresent1.41 s$0.00055
commerce catalog desktop / clean / element presence / absentValidCorrectAbsentAbsent1.52 s$0.00055
analytics dashboard desktop / clean / element presence / presentValidCorrectPresentPresent1.43 s$0.00055
analytics dashboard desktop / clean / element presence / absentValidCorrectAbsentAbsent1.30 s$0.00055
iot control desktop / clean / element presence / presentValidCorrectPresentPresent1.38 s$0.00055
iot control desktop / clean / element presence / absentValidCorrectAbsentAbsent1.60 s$0.00055
banking ios / clean / element presence / presentValidCorrectPresentPresent1.32 s$0.00017
banking ios / clean / element presence / absentValidCorrectAbsentAbsent1.10 s$0.00017
flight pass ios / clean / element presence / presentValidCorrectPresentPresent1.07 s$0.00017
flight pass ios / clean / element presence / absentValidCorrectAbsentAbsent1.11 s$0.00017
smart home android / clean / element presence / presentValidCorrectPresentPresent1.11 s$0.00017
smart home android / clean / element presence / absentValidCorrectAbsentAbsent1.11 s$0.00017

Can the model distinguish clean and broken layouts?

Six clean controls and six screenshots with one seeded defect per model. Transport failures remain operational failures, not accuracy answers.

Balanced accuracy58%
Defect recall17%
Clean specificity100%
Coverage12/12
Localization IoU
commerce catalog desktop / clean / layout defectValidCorrectNot applicable1.87 s$0.00062
commerce catalog desktop / defect / layout defectInvalidIncorrectIncorrect0.0004.55 s$0.0018
analytics dashboard desktop / clean / layout defectValidCorrectNot applicable1.42 s$0.00062
analytics dashboard desktop / defect / layout defectInvalidIncorrectIncorrect0.0004.57 s$0.0018
iot control desktop / clean / layout defectValidCorrectNot applicable1.48 s$0.00063
iot control desktop / defect / layout defectInvalidIncorrectIncorrect0.0004.42 s$0.0018
banking ios / clean / layout defectValidCorrectNot applicable1.33 s$0.00025
banking ios / defect / layout defectInvalidIncorrectIncorrect0.0004.32 s$0.0014
flight pass ios / clean / layout defectValidCorrectNot applicable1.22 s$0.00024
flight pass ios / defect / layout defectValidCorrectIncorrect0.0001.69 s$0.00038
smart home android / clean / layout defectValidCorrectNot applicable1.28 s$0.00025
smart home android / defect / layout defectValidIncorrectIncorrect0.0001.17 s$0.00025

Why this research matters

Using the right kind of vision for each test

Top vision LLMs can reason about complex visual meaning, but they are still too slow and expensive for every test step. This research helps us identify where their semantic capabilities justify that trade-off.

01

Fast local evaluation

Repeato’s own local vision algorithms handle frequent interactions, visual matching, and regression checks with fast feedback, predictable behavior, and no per-request model cost.

02

Semantic reasoning when needed

Top vision LLMs are useful for more complicated questions: understanding meaning, checking dynamic content, describing defects, or evaluating concepts that cannot be captured by a simple visual fingerprint.

03

An evidence-based hybrid

Benchmarking accuracy, latency, reliability, and cost helps us decide which checks should remain local and which benefit enough from an LLM to justify a slower, more expensive evaluation.

Our product strategy

Use Repeato’s local computer vision as the fast foundation, then bring in top vision LLMs selectively for complicated or semantic reasoning where they provide meaningful additional value.

Check out the whole vision benchmark →

Test coverage

Six screens, three task types

Every model is evaluated against the same frozen corpus, prompts, response schemas, and normalized 0–1000 geometry.

01

Commerce catalog

Desktop web with a clean control and a paired container clipping defect.

02

Analytics dashboard

Desktop web with a clean control and a paired text overflow defect.

03

IoT control center

Desktop web with a clean control and a paired element overlap defect.

EN