Benchmark

Which model reviews best, and what it costs

Kind of work

8 models on 1 harness against the same 4 pull requests, scored on 75 defects written down before any of them ran · measured since 2026-08-15, each model at the measurement that read best

Which model reviewed best

Initially ordered by blocking recall · 27 of 28 pairs overlap
Value of F1 0 → 1
Every model on record; sortable by every column, initially ordered by blocking recall
deepseek v4.1 flash 88.2% (60.8% to 97.3%) 56% (38.2% to 72.4%) 95.5% (84.9% to 98.7%) 70.6% (52.7% to 83.5%) 100% (51%–100%) 68.3 V=$5.00 · review cost only$0.0768 4m 27s OpenCode · DeepSeek · max best of 5 efforts4 over 4 PRs
gpt 5.6 luna tied not separated 88.2% (65.7% to 96.7%) 62.7% (42.8% to 79%) 94% (80.7% to 98.3%) 75.2% (55.9% to 87.6%) 100% (51%–100%) 73.1 V=$5.00 · review cost only$0.0848 6m 4s OpenCode · OpenAI · max best of 5 efforts4 over 4 PRs
gemini 3.8 flash tied not separated 76.5% (50.1% to 91.3%) 42.7% (32.1% to 53.9%) 97% (84.7% to 99.5%) 59.3% (46.6% to 70%) 100% (51%–100%) 48.9 V=$5.00 · review cost only$0.5284 8m 53s OpenCode · Google · xhigh best of 5 efforts4 over 4 PRs
glm 5.3 tied not separated 76.5% (52.7% to 90.4%) 53.3% (35.6% to 70.3%) 88.9% (75.3% to 95.5%) 66.7% (48.3% to 81%) 100% (51%–100%) 60.8 V=$5.00 · review cost only$0.2747 4m 33s OpenCode · Z.AI · max best of 5 efforts4 over 4 PRs
glm 5.3 flash tied not separated 76.5% (51.2% to 91%) 44% (29.9% to 59.2%) 94.3% (70.9% to 99.1%) 60% (42% to 74.1%) 100% (51%–100%) 59.6 V=$5.00 · review cost only$0.0150 3m 22s OpenCode · Z.AI · high best of 5 efforts4 over 4 PRs
kimi k3 tied not separated 76.5% (52.7% to 90.4%) 45.3% (32.5% to 58.8%) 89.5% (75.9% to 95.8%) 60.2% (45.6% to 72.9%) 100% (51%–100%) 53.5 V=$5.00 · review cost only$0.3360 3m 42s OpenCode · Moonshot AI · xhigh best of 5 efforts4 over 4 PRs
deepseek v4 flash tied not separated 35.3% (14.7% to 63.3%) 21.3% (11.5% to 36%) 84.2% (40.3% to 97.7%) 34% (17.9% to 52.7%) 100% (51%–100%) 34.0 V=$5.00 · review cost only$0.0152 4m 13s OpenCode · DeepSeek · xhigh best of 5 efforts4 over 4 PRs
sonnet 5.5 tied not separated 1/3 (6.1% to 79.2%) 38.9% (20.3% to 61.4%) 63.6% (35.4% to 84.8%) 48.3% (25.8% to 71.2%) 100% (20.7%–100%) 47.2 V=$5.00 · review cost only$0.0526 18s OpenCode · Anthropic · low 1 over 1 PR

7 of 7 neighbouring pairs have overlapping intervals on this column, so the order between them is not a measured difference.

The same models, one pull request at a time

4 pull requests measured · a model does not read the same on every diff
What to read each model on
higher is better
0%50%100%feat(web): shareable findings reports dashboardfeat(proxy): reclaim space from the blob store proxyfeat(registry): saved views with CSV export registryfeat(scanner): scan the transitive graph scannerdeepseek v4.1 flash deepseek v4.1 flashgpt 5.6 luna gpt 5.6 lunagemini 3.8 flash gemini 3.8 flashglm 5.3 glm 5.3kimi k3 kimi k3glm 5.3 flash glm 5.3 flashdeepseek v4 flash deepseek v4 flashsonnet 5.5 sonnet 5.5
  • deepseek v4.1 flash OpenCode · DeepSeek · max
  • gpt 5.6 luna OpenCode · OpenAI · max
  • gemini 3.8 flash OpenCode · Google · xhigh
  • glm 5.3 OpenCode · Z.AI · max
  • glm 5.3 flash OpenCode · Z.AI · high
  • kimi k3 OpenCode · Moonshot AI · xhigh
  • deepseek v4 flash OpenCode · DeepSeek · xhigh
  • sonnet 5.5 OpenCode · Anthropic · low
Each model's reading on each pull request
Model feat(web): shareable findings reportsfeat(proxy): reclaim space from the blob storefeat(registry): saved views with CSV exportfeat(scanner): scan the transitive graph Trials on record Completed
deepseek v4.1 flash OpenCode · DeepSeek · max5/7 (35.9% to 91.8%) 4/4 (51% to 100%) 3/3 (43.9% to 100%) 3/3 (43.9% to 100%) 4 100%
gpt 5.6 luna OpenCode · OpenAI · max6/7 (48.7% to 97.4%) 3/4 (30.1% to 95.4%) 3/3 (43.9% to 100%) 3/3 (43.9% to 100%) 4 100%
gemini 3.8 flash OpenCode · Google · xhigh4/7 (25% to 84.2%) 3/4 (30.1% to 95.4%) 3/3 (43.9% to 100%) 3/3 (43.9% to 100%) 4 100%
glm 5.3 OpenCode · Z.AI · max5/7 (35.9% to 91.8%) 3/4 (30.1% to 95.4%) 3/3 (43.9% to 100%) 2/3 (20.8% to 93.9%) 4 100%
glm 5.3 flash OpenCode · Z.AI · high6/7 (48.7% to 97.4%) 3/4 (30.1% to 95.4%) 3/3 (43.9% to 100%) 1/3 (6.1% to 79.2%) 4 100%
kimi k3 OpenCode · Moonshot AI · xhigh6/7 (48.7% to 97.4%) 2/4 (15% to 85%) 3/3 (43.9% to 100%) 2/3 (20.8% to 93.9%) 4 100%
deepseek v4 flash OpenCode · DeepSeek · xhigh4/7 (25% to 84.2%) 0/4 (0% to 49%) 1/3 (6.1% to 79.2%) 1/3 (6.1% to 79.2%) 4 100%
sonnet 5.5 OpenCode · Anthropic · lowNot recorded Not recorded 1/3 (6.1% to 79.2%) Not recorded 1 100%
How these models were measured

Which harness a model ran under is a variable held beside it, not the thing this page compares. These sections read the harness: what it costs to find a blocking defect with one, how each has moved campaign by campaign, and what every configuration of one was priced at.

Cost of a blocking defect, by harness

Cheapest first · 17 of 75 defects are merge-blocking · newest of 37 campaigns, sweep-opencode-claude-sonnet-5.5-low
OpenCode Anthropic · sonnet 5.5 · effort low · 1 runs
$0.0526 $0.05 per must-fix bug found per review
Must-fix bugs found
1/3 (6.1% to 79.2%)
All bugs found
38.9% (20.3% to 61.4%)
Time per review
18s
Extra findings
36.4% of findings
Completed
100% (20.7%–100%)
Spent
$0.0526

The same reviewers, campaign by campaign

Oldest first · the panels above are the last column · recall carries its 95% interval over the blocking defects; a cost carries the range between repeats
Measure
Blocking recall for every reviewer in every campaign, oldest campaign first
Reviewerci-glm-5.3-flash-high 4 per armsweep-opencode-glm-5.3-high 4 per armsweep-opencode-glm-5.3-flash-high 4 per armsweep-opencode-kimi-k3-high 4 per armsweep-opencode-deepseek-v4-flash-high 4 per armsweep-opencode-deepseek-v4.1-flash-high 4 per armsweep-opencode-glm-5.3-low 4 per armsweep-opencode-glm-5.3-medium 4 per armsweep-opencode-glm-5.3-xhigh 4 per armsweep-opencode-glm-5.3-max 4 per armsweep-opencode-glm-5.3-flash-low 4 per armsweep-opencode-glm-5.3-flash-medium 4 per armsweep-opencode-glm-5.3-flash-xhigh 4 per armsweep-opencode-glm-5.3-flash-max 4 per armsweep-opencode-kimi-k3-low 4 per armsweep-opencode-kimi-k3-medium 4 per armsweep-opencode-kimi-k3-xhigh 4 per armsweep-opencode-kimi-k3-max 4 per armsweep-opencode-deepseek-v4-flash-low 4 per armsweep-opencode-deepseek-v4-flash-medium 4 per armsweep-opencode-deepseek-v4-flash-xhigh 4 per armsweep-opencode-deepseek-v4-flash-max 4 per armsweep-opencode-deepseek-v4.1-flash-low 4 per armsweep-opencode-deepseek-v4.1-flash-medium 4 per armsweep-opencode-deepseek-v4.1-flash-xhigh 4 per armsweep-opencode-deepseek-v4.1-flash-max 4 per armsweep-opencode-gpt-5.6-luna-high 4 per armsweep-opencode-gpt-5.6-luna-low 4 per armsweep-opencode-gpt-5.6-luna-medium 4 per armsweep-opencode-gpt-5.6-luna-xhigh 4 per armsweep-opencode-gpt-5.6-luna-max 4 per armsweep-opencode-gemini-3.8-flash-high 4 per armsweep-opencode-gemini-3.8-flash-low 4 per armsweep-opencode-gemini-3.8-flash-medium 4 per armsweep-opencode-gemini-3.8-flash-xhigh 4 per armsweep-opencode-gemini-3.8-flash-max 4 per armsweep-opencode-claude-sonnet-5.5-low 1 per arm
OpenCode76.5% 52.7% to 90.4% 100% completed70.6% 46.9% to 86.7% 100% completed76.5% 51.2% to 91% 100% completed70.6% 46.9% to 86.7% 100% completed23.5% 8% to 52% 100% completed70.6% 46.9% to 86.7% 100% completed70.6% 46.9% to 86.7% 100% completed58.8% 36% to 78.4% 100% completed70.6% 46.9% to 86.7% 100% completed76.5% 52.7% to 90.4% 100% completed41.2% 21.6% to 64% 100% completed70.6% 46.9% to 86.7% 100% completed70.6% 46.9% to 86.7% 100% completed76.5% 52.7% to 90.4% 100% completed52.9% 30.2% to 74.5% 100% completed70.6% 46.9% to 86.7% 100% completed76.5% 52.7% to 90.4% 100% completed76.5% 52.7% to 90.4% 100% completed11.8% 1.6% to 51.7% 100% completed23.5% 9.6% to 47.3% 100% completed35.3% 14.7% to 63.3% 100% completed17.6% 3.9% to 53.2% 100% completed64.7% 41.3% to 82.7% 100% completed70.6% 46.9% to 86.7% 100% completed76.5% 52.7% to 90.4% 100% completed88.2% 60.8% to 97.3% 100% completed82.4% 59% to 93.8% 100% completed35.3% 15.9% to 61.2% 100% completed70.6% 35.2% to 91.4% 100% completed76.5% 52.7% to 90.4% 100% completed88.2% 65.7% to 96.7% 100% completed70.6% 37.8% to 90.5% 100% completed47.1% 26.2% to 69% 100% completed70.6% 46.9% to 86.7% 100% completed76.5% 50.1% to 91.3% 100% completed76.5% 50.1% to 91.3% 100% completed33.3% 6.1% to 79.2% 100% completed

Across 37 campaigns OpenCode moved furthest on this measure, 76.5% to 33.3%.

Every configuration we could price

Colour is reviewer · size is effort · filled markers are on the frontier

What a configuration costs against what it catches

What a review is priced in
Which bugs count

One marker per configuration, each priced over every pull request it was run on. Filled markers are on the efficient frontier: nothing in the set is both cheaper and better. 1 of 1 configurations sit on it

32%33%33%34%34%$0.05Anthropic1Cost per review (USD, log scale)Must-fix bugs found
The efficient frontier for this measure, best first. Highlighted rows are efficient on cost and on time, and are the ones badged in the chart above
#ReviewerModelEffortPer reviewTimeMust-fixMatchedCompleted
1 OpenCodeAnthropic · sonnet 5.5low$0.052618s33%64%100%

What kind of defect it finds

Severity comes from the answer keys, not the reviewer

By severity

How much each bug matters

0%25%50%75%100%1/3MUST-FIX17 bugsmust fix before merge45.5%ISSUE46 bugsshould fix1/4SUGGESTION12 bugsoptional

By difficulty tier

How hard each one was to see

0%25%50%75%100%TIER 111TIER 230TIER 325TIER 49Tier 1 floor: every bug in the diff

0 of 1 do best on the defects that block a merge and worst on the optional ones, though no reviewer’s blocking and optional intervals are far enough apart for that order to be a measured difference. Tier 1 is the floor, the 11 defects visible in the changed lines, and nobody clears all of it: 50% at best, 50% at worst. From tier 3 to tier 4 recall falls 25 points, which is the ordering the tiers were built to produce

What a review costs you, in money and in time

Both panels start at 52% · bars are 95% intervals over the blocking defects

Cost

Cheaper is left, better is up. Dashed rays are equal cost per must-fix bug, priced at the label

What is written beside each mark
40%50%60%70%80%90%$0.05$0.05$0.05$0.06OpenCodeCost of one reviewbars are 95% intervals over the blocking defects, 1 trials per pull requestMust-fix bugs found

Time

Faster is left

40%50%60%70%80%90%17s18s19sOpenCodeTime for one reviewbars are 95% intervals over the blocking defects, 1 trials per pull requestMust-fix bugs found

Hollow rings are the 4 pull requests behind each solid marker, each one against its own bug count

What more effort buys

Same reviewer, same role stack, same pull request · only the effort knob moved
No reviewer ran two effort settings on the same pull request, so there is no move to show
benchee benchee-dashboard-2 built from 5351a11f Static benchmark evidence ·