Benchmark

Which model reviews best, and what it costs

Kind of work

9 models on 1 harness against the same 4 pull requests, scored on 75 defects written down before any of them ran · measured since 2026-08-15, each model at the measurement that read best

Which model reviewed best

Initially ordered by blocking recall · 34 of 36 pairs overlap
Value of F1 0 → 1
Every model on record; sortable by every column, initially ordered by blocking recall
sonnet 5.5 7/7 (64.6% to 100%) 65.8% (42.6% to 83.3%) 92.6% (65.2% to 98.8%) 76.9% (51.5% to 90.4%) 100% (34.2%–100%) 7.7 V=$5.00 · review cost only$3.47 10m 45s OpenCode · Anthropic · max best of 5 efforts2 over 2 PRs
deepseek v4.1 flash tied not separated 88.2% (60.8% to 97.3%) 56% (38.2% to 72.4%) 95.5% (84.9% to 98.7%) 70.6% (52.7% to 83.5%) 100% (51%–100%) 68.3 V=$5.00 · review cost only$0.0768 4m 27s OpenCode · DeepSeek · max best of 5 efforts4 over 4 PRs
gpt 5.6 luna tied not separated 88.2% (65.7% to 96.7%) 62.7% (42.8% to 79%) 94% (80.7% to 98.3%) 75.2% (55.9% to 87.6%) 100% (51%–100%) 73.1 V=$5.00 · review cost only$0.0848 6m 4s OpenCode · OpenAI · max best of 5 efforts4 over 4 PRs
gemini 3.8 flash tied not separated 76.5% (50.1% to 91.3%) 42.7% (32.1% to 53.9%) 97% (84.7% to 99.5%) 59.3% (46.6% to 70%) 100% (51%–100%) 48.9 V=$5.00 · review cost only$0.5284 8m 53s OpenCode · Google · xhigh best of 5 efforts4 over 4 PRs
glm 5.3 tied not separated 76.5% (52.7% to 90.4%) 53.3% (35.6% to 70.3%) 88.9% (75.3% to 95.5%) 66.7% (48.3% to 81%) 100% (51%–100%) 60.8 V=$5.00 · review cost only$0.2747 4m 33s OpenCode · Z.AI · max best of 5 efforts4 over 4 PRs
glm 5.3 flash tied not separated 76.5% (51.2% to 91%) 44% (29.9% to 59.2%) 94.3% (70.9% to 99.1%) 60% (42% to 74.1%) 100% (51%–100%) 59.6 V=$5.00 · review cost only$0.0150 3m 22s OpenCode · Z.AI · high best of 5 efforts4 over 4 PRs
kimi k3 tied not separated 76.5% (52.7% to 90.4%) 45.3% (32.5% to 58.8%) 89.5% (75.9% to 95.8%) 60.2% (45.6% to 72.9%) 100% (51%–100%) 53.5 V=$5.00 · review cost only$0.3360 3m 42s OpenCode · Moonshot AI · xhigh best of 5 efforts4 over 4 PRs
sonnet 5 tied not separated 76.5% (52.7% to 90.4%) 37.3% (27.3% to 48.7%) 96.6% (82.8% to 99.4%) 53.8% (41% to 65.3%) 100% (51%–100%) 22.4 V=$5.00 · review cost only$1.57 8m 20s OpenCode · Anthropic · max best of 5 efforts4 over 4 PRs
deepseek v4 flash 35.3% (14.7% to 63.3%) 21.3% (11.5% to 36%) 84.2% (40.3% to 97.7%) 34% (17.9% to 52.7%) 100% (51%–100%) 34.0 V=$5.00 · review cost only$0.0152 4m 13s OpenCode · DeepSeek · xhigh best of 5 efforts4 over 4 PRs

8 of 8 neighbouring pairs have overlapping intervals on this column, so the order between them is not a measured difference.

The same models, one pull request at a time

4 pull requests measured · a model does not read the same on every diff
What to read each model on
higher is better
0%50%100%feat(web): shareable findings reports dashboardfeat(proxy): reclaim space from the blob store proxyfeat(registry): saved views with CSV export registryfeat(scanner): scan the transitive graph scannersonnet 5.5 sonnet 5.5deepseek v4.1 flash deepseek v4.1 flashgpt 5.6 luna gpt 5.6 lunagemini 3.8 flash gemini 3.8 flashsonnet 5 sonnet 5glm 5.3 glm 5.3kimi k3 kimi k3glm 5.3 flash glm 5.3 flashdeepseek v4 flash deepseek v4 flash
  • sonnet 5.5 OpenCode · Anthropic · max
  • deepseek v4.1 flash OpenCode · DeepSeek · max
  • gpt 5.6 luna OpenCode · OpenAI · max
  • gemini 3.8 flash OpenCode · Google · xhigh
  • glm 5.3 OpenCode · Z.AI · max
  • glm 5.3 flash OpenCode · Z.AI · high
  • kimi k3 OpenCode · Moonshot AI · xhigh
  • sonnet 5 OpenCode · Anthropic · max
  • deepseek v4 flash OpenCode · DeepSeek · xhigh
Each model's reading on each pull request
Model feat(web): shareable findings reportsfeat(proxy): reclaim space from the blob storefeat(registry): saved views with CSV exportfeat(scanner): scan the transitive graph Trials on record Completed
sonnet 5.5 OpenCode · Anthropic · maxNot recorded 4/4 (51% to 100%) 3/3 (43.9% to 100%) Not recorded 2 100%
deepseek v4.1 flash OpenCode · DeepSeek · max5/7 (35.9% to 91.8%) 4/4 (51% to 100%) 3/3 (43.9% to 100%) 3/3 (43.9% to 100%) 4 100%
gpt 5.6 luna OpenCode · OpenAI · max6/7 (48.7% to 97.4%) 3/4 (30.1% to 95.4%) 3/3 (43.9% to 100%) 3/3 (43.9% to 100%) 4 100%
gemini 3.8 flash OpenCode · Google · xhigh4/7 (25% to 84.2%) 3/4 (30.1% to 95.4%) 3/3 (43.9% to 100%) 3/3 (43.9% to 100%) 4 100%
glm 5.3 OpenCode · Z.AI · max5/7 (35.9% to 91.8%) 3/4 (30.1% to 95.4%) 3/3 (43.9% to 100%) 2/3 (20.8% to 93.9%) 4 100%
glm 5.3 flash OpenCode · Z.AI · high6/7 (48.7% to 97.4%) 3/4 (30.1% to 95.4%) 3/3 (43.9% to 100%) 1/3 (6.1% to 79.2%) 4 100%
kimi k3 OpenCode · Moonshot AI · xhigh6/7 (48.7% to 97.4%) 2/4 (15% to 85%) 3/3 (43.9% to 100%) 2/3 (20.8% to 93.9%) 4 100%
sonnet 5 OpenCode · Anthropic · max5/7 (35.9% to 91.8%) 2/4 (15% to 85%) 3/3 (43.9% to 100%) 3/3 (43.9% to 100%) 4 100%
deepseek v4 flash OpenCode · DeepSeek · xhigh4/7 (25% to 84.2%) 0/4 (0% to 49%) 1/3 (6.1% to 79.2%) 1/3 (6.1% to 79.2%) 4 100%
How these models were measured

Which harness a model ran under is a variable held beside it, not the thing this page compares. These sections read the harness: what it costs to find a blocking defect with one, how each has moved campaign by campaign, and what every configuration of one was priced at.

Cost of a blocking defect, by harness

Cheapest first · 17 of 75 defects are merge-blocking · newest of 46 campaigns, sweep-opencode-claude-sonnet-5.5-max
OpenCode Anthropic · sonnet 5.5 · effort max · 2 runs
$0.9906 $3.47 per must-fix bug found per review
Must-fix bugs found
7/7 (64.6% to 100%)
All bugs found
65.8% (49.9% to 78.8%)
Time per review
10m 45s
Extra findings
7.4% of findings
Completed
100% (34.2%–100%)
Spent
$6.93

The same reviewers, campaign by campaign

Oldest first · the panels above are the last column · recall carries its 95% interval over the blocking defects; a cost carries the range between repeats
Measure
Blocking recall for every reviewer in every campaign, oldest campaign first
Reviewerci-glm-5.3-flash-high 4 per armsweep-opencode-glm-5.3-high 4 per armsweep-opencode-glm-5.3-flash-high 4 per armsweep-opencode-kimi-k3-high 4 per armsweep-opencode-deepseek-v4-flash-high 4 per armsweep-opencode-deepseek-v4.1-flash-high 4 per armsweep-opencode-glm-5.3-low 4 per armsweep-opencode-glm-5.3-medium 4 per armsweep-opencode-glm-5.3-xhigh 4 per armsweep-opencode-glm-5.3-max 4 per armsweep-opencode-glm-5.3-flash-low 4 per armsweep-opencode-glm-5.3-flash-medium 4 per armsweep-opencode-glm-5.3-flash-xhigh 4 per armsweep-opencode-glm-5.3-flash-max 4 per armsweep-opencode-kimi-k3-low 4 per armsweep-opencode-kimi-k3-medium 4 per armsweep-opencode-kimi-k3-xhigh 4 per armsweep-opencode-kimi-k3-max 4 per armsweep-opencode-deepseek-v4-flash-low 4 per armsweep-opencode-deepseek-v4-flash-medium 4 per armsweep-opencode-deepseek-v4-flash-xhigh 4 per armsweep-opencode-deepseek-v4-flash-max 4 per armsweep-opencode-deepseek-v4.1-flash-low 4 per armsweep-opencode-deepseek-v4.1-flash-medium 4 per armsweep-opencode-deepseek-v4.1-flash-xhigh 4 per armsweep-opencode-deepseek-v4.1-flash-max 4 per armsweep-opencode-gpt-5.6-luna-high 4 per armsweep-opencode-gpt-5.6-luna-low 4 per armsweep-opencode-gpt-5.6-luna-medium 4 per armsweep-opencode-gpt-5.6-luna-xhigh 4 per armsweep-opencode-gpt-5.6-luna-max 4 per armsweep-opencode-gemini-3.8-flash-high 4 per armsweep-opencode-gemini-3.8-flash-low 4 per armsweep-opencode-gemini-3.8-flash-medium 4 per armsweep-opencode-gemini-3.8-flash-xhigh 4 per armsweep-opencode-gemini-3.8-flash-max 4 per armsweep-opencode-claude-sonnet-5-high 4 per armsweep-opencode-claude-sonnet-5-low 4 per armsweep-opencode-claude-sonnet-5-medium 4 per armsweep-opencode-claude-sonnet-5-xhigh 4 per armsweep-opencode-claude-sonnet-5-max 4 per armsweep-opencode-claude-sonnet-5.5-high 4 per armsweep-opencode-claude-sonnet-5.5-low 5 per armsweep-opencode-claude-sonnet-5.5-medium 4 per armsweep-opencode-claude-sonnet-5.5-xhigh 4 per armsweep-opencode-claude-sonnet-5.5-max 2 per arm
OpenCode76.5% 52.7% to 90.4% 100% completed70.6% 46.9% to 86.7% 100% completed76.5% 51.2% to 91% 100% completed70.6% 46.9% to 86.7% 100% completed23.5% 8% to 52% 100% completed70.6% 46.9% to 86.7% 100% completed70.6% 46.9% to 86.7% 100% completed58.8% 36% to 78.4% 100% completed70.6% 46.9% to 86.7% 100% completed76.5% 52.7% to 90.4% 100% completed41.2% 21.6% to 64% 100% completed70.6% 46.9% to 86.7% 100% completed70.6% 46.9% to 86.7% 100% completed76.5% 52.7% to 90.4% 100% completed52.9% 30.2% to 74.5% 100% completed70.6% 46.9% to 86.7% 100% completed76.5% 52.7% to 90.4% 100% completed76.5% 52.7% to 90.4% 100% completed11.8% 1.6% to 51.7% 100% completed23.5% 9.6% to 47.3% 100% completed35.3% 14.7% to 63.3% 100% completed17.6% 3.9% to 53.2% 100% completed64.7% 41.3% to 82.7% 100% completed70.6% 46.9% to 86.7% 100% completed76.5% 52.7% to 90.4% 100% completed88.2% 60.8% to 97.3% 100% completed82.4% 59% to 93.8% 100% completed35.3% 15.9% to 61.2% 100% completed70.6% 35.2% to 91.4% 100% completed76.5% 52.7% to 90.4% 100% completed88.2% 65.7% to 96.7% 100% completed70.6% 37.8% to 90.5% 100% completed47.1% 26.2% to 69% 100% completed70.6% 46.9% to 86.7% 100% completed76.5% 50.1% to 91.3% 100% completed76.5% 50.1% to 91.3% 100% completed41.2% 19% to 67.6% 100% completed35.3% 14.7% to 63.3% 100% completed41.2% 17.8% to 69.3% 100% completed58.8% 31.5% to 81.6% 100% completed76.5% 52.7% to 90.4% 100% completed82.4% 59% to 93.8% 100% completed38.2% 19.4% to 61.4% 100% completed76.5% 52.7% to 90.4% 100% completed94.1% 70.9% to 99.1% 100% completed100% 64.6% to 100% 100% completed

Across 46 campaigns OpenCode moved furthest on this measure, 76.5% to 100%.

Every configuration we could price

Colour is reviewer · size is effort · filled markers are on the frontier

What a configuration costs against what it catches

What a review is priced in
Which bugs count

One marker per configuration, each priced over every pull request it was run on. Filled markers are on the efficient frontier: nothing in the set is both cheaper and better. 1 of 1 configurations sit on it

96%97%98%99%100%Anthropic1Cost per review (USD, log scale)Must-fix bugs found
The efficient frontier for this measure, best first. Highlighted rows are efficient on cost and on time, and are the ones badged in the chart above
#ReviewerModelEffortPer reviewTimeMust-fixMatchedCompleted
1 OpenCodeAnthropic · sonnet 5.5max$3.4710m 45s100%93%100%

What kind of defect it finds

Severity comes from the answer keys, not the reviewer

By severity

How much each bug matters

0%25%50%75%100%7/7MUST-FIX17 bugsmust fix before merge68.2%ISSUE46 bugsshould fix3/9SUGGESTION12 bugsoptional

By difficulty tier

How hard each one was to see

0%25%50%75%100%TIER 111TIER 230TIER 325TIER 49Tier 1 floor: every bug in the diff

All one do best on the defects that block a merge and worst on the optional ones, which is the order you want, though no reviewer’s blocking and optional intervals are far enough apart for that order to be a measured difference. Tier 1 is the floor, the 11 defects visible in the changed lines, and 1 of 1 clear all of it, from 100% to 100%. From tier 3 to tier 4 recall falls 46 points, which is the ordering the tiers were built to produce

What a review costs you, in money and in time

Both panels start at 52% · bars are 95% intervals over the blocking defects

Cost

Cheaper is left, better is up. Dashed rays are equal cost per must-fix bug, priced at the label

What is written beside each mark
60%70%80%90%100%$3.20$3.40$3.60OpenCodeCost of one reviewbars are 95% intervals over the blocking defects, 1 trials per pull requestMust-fix bugs found

Time

Faster is left

60%70%80%90%100%10m 0s10m 50sOpenCodeTime for one reviewbars are 95% intervals over the blocking defects, 1 trials per pull requestMust-fix bugs found

Hollow rings are the 4 pull requests behind each solid marker, each one against its own bug count. 4 readings sit outside these frames and are not drawn

What more effort buys

Same reviewer, same role stack, same pull request · only the effort knob moved
No reviewer ran two effort settings on the same pull request, so there is no move to show
benchee benchee-dashboard-2 built from 6b993cc3 Static benchmark evidence ·