Benchmark

Which model reviews best, and what it costs

Kind of work

5 models on 1 harness against the same 4 pull requests, scored on 75 defects written down before any of them ran · measured since 2026-08-15, each model at the measurement that read best

Which model reviewed best

Initially ordered by blocking recall · 10 of 10 pairs overlap
Value of F1 0 → 1
Every model on record; sortable by every column, initially ordered by blocking recall
glm 5.3 3/3 (43.9% to 100%) 55.6% (33.7% to 75.4%) N/A N/A 100% (20.7%–100%) Unavailable $0.1197 38s OpenCode · Z.AI · high best of 4 efforts1 over 1 PR
glm 5.3 flash tied not separated 6/7 (48.7% to 97.4%) 58.8% (36% to 78.4%) N/A N/A 100% (20.7%–100%) Unavailable $0.0170 3m 59s OpenCode · Z.AI · high 1 over 1 PR
kimi k3 tied not separated 5/7 (35.9% to 91.8%) 58.8% (36% to 78.4%) N/A N/A 100% (20.7%–100%) Unavailable $0.2204 2m 46s OpenCode · Moonshot AI · high 1 over 1 PR
deepseek v4.1 flash tied not separated 2/3 (20.8% to 93.9%) 35% (18.1% to 56.7%) N/A N/A 100% (20.7%–100%) Unavailable $0.0226 41s OpenCode · DeepSeek · high 1 over 1 PR
deepseek v4 flash tied not separated 0/3 (0% to 56.1%) 22.2% (9% to 45.2%) N/A N/A 100% (20.7%–100%) Unavailable $0.007961 2m 13s OpenCode · DeepSeek · high 1 over 1 PR

4 of 4 neighbouring pairs have overlapping intervals on this column, so the order between them is not a measured difference.

How each model has moved

F1 by the day it was measured · each model at effort high on the harness and vendor of its row above, whichever rung read best there
Nothing measured since 2026-08-15

This chart starts at 2026-08-15 and fills in as sweeps are recorded. No earlier day is on record.

The same models, one pull request at a time

4 pull requests measured · a model does not read the same on every diff
What to read each model on
higher is better
0%50%100%feat(web): shareable findings reports dashboardfeat(registry): saved views with CSV export registryfeat(scanner): scan the transitive graph scannerglm 5.3 glm 5.3glm 5.3 flash glm 5.3 flashkimi k3 kimi k3deepseek v4.1 flash deepseek v4.1 flashdeepseek v4 flash deepseek v4 flash
  • glm 5.3 OpenCode · Z.AI · high
  • glm 5.3 flash OpenCode · Z.AI · high
  • kimi k3 OpenCode · Moonshot AI · high
  • deepseek v4.1 flash OpenCode · DeepSeek · high
  • deepseek v4 flash OpenCode · DeepSeek · high
Each model's reading on each pull request
Model feat(web): shareable findings reportsfeat(registry): saved views with CSV exportfeat(scanner): scan the transitive graph Trials on record Completed
glm 5.3 OpenCode · Z.AI · highNot recorded 3/3 (43.9% to 100%) Not recorded 1 100%
glm 5.3 flash OpenCode · Z.AI · high6/7 (48.7% to 97.4%) Not recorded Not recorded 1 100%
kimi k3 OpenCode · Moonshot AI · high5/7 (35.9% to 91.8%) Not recorded Not recorded 1 100%
deepseek v4.1 flash OpenCode · DeepSeek · highNot recorded Not recorded 2/3 (20.8% to 93.9%) 1 100%
deepseek v4 flash OpenCode · DeepSeek · highNot recorded 0/3 (0% to 56.1%) Not recorded 1 100%
How these models were measured

Which harness a model ran under is a variable held beside it, not the thing this page compares. These sections read the harness: what it costs to find a blocking defect with one, how each has moved campaign by campaign, and what every configuration of one was priced at.

Cost of a blocking defect, by harness

Cheapest first · 17 of 75 defects are merge-blocking · newest of 9 campaigns, sweep-opencode-glm-5.3-xhigh
OpenCode Z.AI · glm 5.3 · effort xhigh · 4 runs
$0.0376 $0.12 per must-fix bug found per review
Must-fix bugs found
76.5% (52.7% to 90.4%)
All bugs found
50.7% (39.6% to 61.7%)
Time per review
2m 24s
Extra findings
0.0% of findings
Completed
100% (51%–100%)
Spent
$0.4885

The same reviewers, campaign by campaign

Oldest first · the panels above are the last column · recall carries its 95% interval over the blocking defects; a cost carries the range between repeats
Measure
Blocking recall for every reviewer in every campaign, oldest campaign first
Reviewerci-glm-5.3-flash-high 4 per armsweep-opencode-glm-5.3-high 4 per armsweep-opencode-glm-5.3-flash-high 4 per armsweep-opencode-kimi-k3-high 4 per armsweep-opencode-deepseek-v4-flash-high 4 per armsweep-opencode-deepseek-v4.1-flash-high 4 per armsweep-opencode-glm-5.3-low 4 per armsweep-opencode-glm-5.3-medium 4 per armsweep-opencode-glm-5.3-xhigh 4 per arm
OpenCode76.5% 52.7% to 90.4% 100% completed76.5% 52.7% to 90.4% 100% completed76.5% 51.2% to 91% 100% completed70.6% 46.9% to 86.7% 100% completed23.5% 8% to 52% 100% completed70.6% 46.9% to 86.7% 100% completed70.6% 46.9% to 86.7% 100% completed58.8% 36% to 78.4% 100% completed76.5% 52.7% to 90.4% 100% completed

Across 9 campaigns OpenCode moved furthest on this measure, 76.5% to 76.5%.

Every configuration we could price

Colour is reviewer · size is effort · filled markers are on the frontier

What a configuration costs against what it catches

What a review is priced in
Which bugs count

One marker per configuration, each priced over every pull request it was run on. Filled markers are on the efficient frontier: nothing in the set is both cheaper and better. 3 of 4 configurations sit on it

64%73%82%91%100%$0.05$0.10$0.20Z.AI123Cost per review (USD, log scale)Must-fix bugs found
The efficient frontier for this measure, best first. Highlighted rows are efficient on cost and on time, and are the ones badged in the chart above
#ReviewerModelEffortPer reviewTimeMust-fixMatchedCompleted
1 OpenCodeZ.AI · glm 5.3xhigh$0.05791m 3s71%100%100%
2 OpenCodeZ.AI · glm 5.3xhigh$0.11234m 39s75%100%100%
3 OpenCodeZ.AI · glm 5.3xhigh$0.15141m 59s100%100%100%

What kind of defect it finds

Severity comes from the answer keys, not the reviewer

By severity

How much each bug matters

0%25%50%75%100%76.5%MUST-FIX17 bugsmust fix before merge47.8%ISSUE46 bugsshould fix25%SUGGESTION12 bugsoptional

By difficulty tier

How hard each one was to see

0%25%50%75%100%TIER 111TIER 230TIER 325TIER 49Tier 1 floor: every bug in the diff

All one do best on the defects that block a merge and worst on the optional ones, which is the order you want, though no reviewer’s blocking and optional intervals are far enough apart for that order to be a measured difference. Tier 1 is the floor, the 11 defects visible in the changed lines, and 1 of 1 clear all of it, from 100% to 100%. From tier 3 to tier 4 recall falls 20 points, which is the ordering the tiers were built to produce

What a review costs you, in money and in time

Both panels start at 52% · bars are 95% intervals over the blocking defects

Cost

Cheaper is left, better is up. Dashed rays are equal cost per must-fix bug, priced at the label

What is written beside each mark
60%70%80%90%$0.12$0.12$0.13$0.13OpenCodeCost of one reviewbars are 95% intervals over the blocking defects, 1 trials per pull requestMust-fix bugs found

Time

Faster is left

60%70%80%90%2m 15s2m 20s2m 25s2m 30s2m 35sOpenCodeTime for one reviewbars are 95% intervals over the blocking defects, 1 trials per pull requestMust-fix bugs found

Hollow rings are the 4 pull requests behind each solid marker, each one against its own bug count. 8 readings sit outside these frames and are not drawn

What more effort buys

Same reviewer, same role stack, same pull request · only the effort knob moved
No reviewer ran two effort settings on the same pull request, so there is no move to show
benchee benchee-dashboard-2 built from d7f63f39 Static benchmark evidence ·