Benchmark

Which model reviews best, and what it costs

Kind of work

1 model on 1 harness against the same 7 pull requests, scored on 83 defects written down before any of them ran · measured since 2026-08-15, each model at the measurement that read best

Which model reviewed best

Initially ordered by blocking recall · 1 of 1 model carries a ranked recall
Value of F1 0 → 1
Every model on record; sortable by every column, initially ordered by blocking recall
glm 5.3 flash 5/7 (35.9% to 91.8%) 52.9% (31% to 73.8%) N/A N/A 100% (20.7%–100%) Unavailable $0.0131 5m 26s OpenCode · Z.AI · high 1 over 1 PR

How each model has moved

F1 by the day it was measured · each model at effort high on the harness and vendor of its row above, whichever rung read best there
Nothing measured since 2026-08-15

This chart starts at 2026-08-15 and fills in as sweeps are recorded. No earlier day is on record.

The same models, one pull request at a time

1 pull request measured · a model does not read the same on every diff
What to read each model on
higher is better
0%50%100%feat(web): shareable findings reports dashboardglm 5.3 flash glm 5.3 flash
  • glm 5.3 flash OpenCode · Z.AI · high
Each model's reading on each pull request
Model feat(web): shareable findings reports Trials on record Completed
glm 5.3 flash OpenCode · Z.AI · high5/7 (35.9% to 91.8%) 1 100%
How these models were measured

Which harness a model ran under is a variable held beside it, not the thing this page compares. These sections read the harness: what it costs to find a blocking defect with one, how each has moved campaign by campaign, and what every configuration of one was priced at.

Cost of a blocking defect, by harness

Cheapest first · 20 of 83 defects are merge-blocking · campaign ci-glm-5.3-flash-high
OpenCode Z.AI · glm 5.3 flash · effort high · 7 runs
$0.005490 $0.01 per must-fix bug found per review
Must-fix bugs found
75% (53.1% to 88.8%)
All bugs found
45.8% (35.5% to 56.4%)
Time per review
2m 58s
Extra findings
13.6% of findings
Completed
100% (64.6%–100%)
Spent
$0.0823

Every configuration we could price

Colour is reviewer · size is effort · filled markers are on the frontier

What a configuration costs against what it catches

What a review is priced in
Which bugs count

One marker per configuration, each priced over every pull request it was run on. Filled markers are on the efficient frontier: nothing in the set is both cheaper and better. 4 of 5 configurations sit on it

64%73%82%91%100%$0.02Z.AI1234Cost per review (USD, log scale)Must-fix bugs found
The efficient frontier for this measure, best first. Highlighted rows are efficient on cost and on time, and are the ones badged in the chart above
#ReviewerModelEffortPer reviewTimeMust-fixMatchedCompleted
1 OpenCodeZ.AI · glm 5.3 flashhigh$0.0048361m 27s67%86%100%
2 OpenCodeZ.AI · glm 5.3 flashhigh$0.01315m 26s71%86%100%
3 OpenCodeZ.AI · glm 5.3 flashhigh$0.01634m 11s75%86%100%
4 OpenCodeZ.AI · glm 5.3 flashhigh$0.02012m 9s100%86%100%

What kind of defect it finds

Severity comes from the answer keys, not the reviewer

By severity

How much each bug matters

0%25%50%75%100%75%MUST-FIX20 bugsmust fix before merge38.8%ISSUE49 bugsshould fix28.6%SUGGESTION14 bugsoptional

By difficulty tier

How hard each one was to see

0%25%50%75%100%TIER 113TIER 232TIER 327TIER 411Tier 1 floor: every bug in the diff

What a review costs you, in money and in time

Both panels start at 52% · bars are 95% intervals over the blocking defects

Cost

Cheaper is left, better is up. Dashed rays are equal cost per must-fix bug, priced at the label

What is written beside each mark
60%70%80%90%$0.01$0.01$0.01$0.01OpenCodeCost of one reviewbars are 95% intervals over the blocking defects, 1 trials per pull requestMust-fix bugs found

Time

Faster is left

60%70%80%90%2m 50s3m 0s3m 10sOpenCodeTime for one reviewbars are 95% intervals over the blocking defects, 1 trials per pull requestMust-fix bugs found

Hollow rings are the 7 pull requests behind each solid marker, each one against its own bug count. 9 readings sit outside these frames and are not drawn

What more effort buys

Same reviewer, same role stack, same pull request · only the effort knob moved
No reviewer ran two effort settings on the same pull request, so there is no move to show
benchee benchee-dashboard-2 built from e9576b6b Static benchmark evidence ·