Benchmark

Which model reviews best, and what it costs

Kind of work

6 models on 1 harness against the same 4 pull requests, scored on 75 defects written down before any of them ran · measured since 2026-08-15, each model at the measurement that read best

Which model reviewed best

Initially ordered by blocking recall · 14 of 15 pairs overlap
Value of F1 0 → 1
Every model on record; sortable by every column, initially ordered by blocking recall
deepseek v4.1 flash 88.2% (60.8% to 97.3%) 56% (38.2% to 72.4%) 95.5% (84.9% to 98.7%) 70.6% (52.7% to 83.5%) 100% (51%–100%) 68.3 V=$5.00 · review cost only$0.0768 4m 27s OpenCode · DeepSeek · max best of 5 efforts4 over 4 PRs
gpt 5.6 luna tied not separated 88.2% (65.7% to 96.7%) 62.7% (42.8% to 79%) 94% (80.7% to 98.3%) 75.2% (55.9% to 87.6%) 100% (51%–100%) 73.1 V=$5.00 · review cost only$0.0848 6m 4s OpenCode · OpenAI · max best of 5 efforts4 over 4 PRs
glm 5.3 tied not separated 76.5% (52.7% to 90.4%) 53.3% (35.6% to 70.3%) 88.9% (75.3% to 95.5%) 66.7% (48.3% to 81%) 100% (51%–100%) 60.8 V=$5.00 · review cost only$0.2747 4m 33s OpenCode · Z.AI · max best of 5 efforts4 over 4 PRs
glm 5.3 flash tied not separated 76.5% (51.2% to 91%) 44% (29.9% to 59.2%) 94.3% (70.9% to 99.1%) 60% (42% to 74.1%) 100% (51%–100%) 59.6 V=$5.00 · review cost only$0.0150 3m 22s OpenCode · Z.AI · high best of 5 efforts4 over 4 PRs
kimi k3 tied not separated 76.5% (52.7% to 90.4%) 45.3% (32.5% to 58.8%) 89.5% (75.9% to 95.8%) 60.2% (45.6% to 72.9%) 100% (51%–100%) 53.5 V=$5.00 · review cost only$0.3360 3m 42s OpenCode · Moonshot AI · xhigh best of 5 efforts4 over 4 PRs
deepseek v4 flash tied not separated 35.3% (14.7% to 63.3%) 21.3% (11.5% to 36%) 84.2% (40.3% to 97.7%) 34% (17.9% to 52.7%) 100% (51%–100%) 34.0 V=$5.00 · review cost only$0.0152 4m 13s OpenCode · DeepSeek · xhigh best of 5 efforts4 over 4 PRs

5 of 5 neighbouring pairs have overlapping intervals on this column, so the order between them is not a measured difference.

How each model has moved

F1 by the day it was measured · each model at effort high on the harness and vendor of its row above, whichever rung read best there
0%50%100%2026-09-292026-09-30deepseek v4.1 flash · 2026-09-30 · 60.9% · 4 trials over 4 pull requests · opencode opencode v2.0.18 · sweep-opencode-deepseek-v4.1-flash-highgpt 5.6 luna · 2026-09-30 · 65.4% · 4 trials over 4 pull requests · opencode opencode v2.0.19 · sweep-opencode-gpt-5.6-luna-highglm 5.3 · 2026-09-30 · 63.1% · 4 trials over 4 pull requests · opencode opencode v2.0.18 · sweep-opencode-glm-5.3-highglm 5.3 flash · 2026-09-29 · 62% · 4 trials over 4 pull requests · opencode opencode v2.0.18 · ci-glm-5.3-flash-highglm 5.3 flash · 2026-09-30 · 59.9% · 4 trials over 4 pull requests · opencode opencode v2.0.18 · sweep-opencode-glm-5.3-flash-highkimi k3 · 2026-09-30 · 62.1% · 4 trials over 4 pull requests · opencode opencode v2.0.18 · sweep-opencode-kimi-k3-highdeepseek v4 flash · 2026-09-30 · 24.2% · 4 trials over 4 pull requests · opencode opencode v2.0.18 · sweep-opencode-deepseek-v4-flash-highgpt 5.6 luna gpt 5.6 lunaglm 5.3 glm 5.3kimi k3 kimi k3deepseek v4.1 flash deepseek v4.1 flashglm 5.3 flash glm 5.3 flashdeepseek v4 flash deepseek v4 flash

5 of 6 models were measured on one day, and a single mark is a reading rather than a trend.

Each model's average, first and latest F1
Model Average First Latest Days Completed Measured as
deepseek v4.1 flash 60.9% 60.9% 2026-09-30 one day only1 100% OpenCode · DeepSeek · max
gpt 5.6 luna 65.4% 65.4% 2026-09-30 one day only1 100% OpenCode · OpenAI · max
glm 5.3 63.1% 63.1% 2026-09-30 one day only1 100% OpenCode · Z.AI · max
glm 5.3 flash 61% 62% 2026-09-2959.9% 2026-09-302 100% OpenCode · Z.AI · high
kimi k3 62.1% 62.1% 2026-09-30 one day only1 100% OpenCode · Moonshot AI · xhigh
deepseek v4 flash 24.2% 24.2% 2026-09-30 one day only1 100% OpenCode · DeepSeek · xhigh

The same models, one pull request at a time

4 pull requests measured · a model does not read the same on every diff
What to read each model on
higher is better
0%50%100%feat(web): shareable findings reports dashboardfeat(proxy): reclaim space from the blob store proxyfeat(registry): saved views with CSV export registryfeat(scanner): scan the transitive graph scannerdeepseek v4.1 flash deepseek v4.1 flashgpt 5.6 luna gpt 5.6 lunaglm 5.3 glm 5.3kimi k3 kimi k3glm 5.3 flash glm 5.3 flashdeepseek v4 flash deepseek v4 flash
  • deepseek v4.1 flash OpenCode · DeepSeek · max
  • gpt 5.6 luna OpenCode · OpenAI · max
  • glm 5.3 OpenCode · Z.AI · max
  • glm 5.3 flash OpenCode · Z.AI · high
  • kimi k3 OpenCode · Moonshot AI · xhigh
  • deepseek v4 flash OpenCode · DeepSeek · xhigh
Each model's reading on each pull request
Model feat(web): shareable findings reportsfeat(proxy): reclaim space from the blob storefeat(registry): saved views with CSV exportfeat(scanner): scan the transitive graph Trials on record Completed
deepseek v4.1 flash OpenCode · DeepSeek · max5/7 (35.9% to 91.8%) 4/4 (51% to 100%) 3/3 (43.9% to 100%) 3/3 (43.9% to 100%) 4 100%
gpt 5.6 luna OpenCode · OpenAI · max6/7 (48.7% to 97.4%) 3/4 (30.1% to 95.4%) 3/3 (43.9% to 100%) 3/3 (43.9% to 100%) 4 100%
glm 5.3 OpenCode · Z.AI · max5/7 (35.9% to 91.8%) 3/4 (30.1% to 95.4%) 3/3 (43.9% to 100%) 2/3 (20.8% to 93.9%) 4 100%
glm 5.3 flash OpenCode · Z.AI · high6/7 (48.7% to 97.4%) 3/4 (30.1% to 95.4%) 3/3 (43.9% to 100%) 1/3 (6.1% to 79.2%) 4 100%
kimi k3 OpenCode · Moonshot AI · xhigh6/7 (48.7% to 97.4%) 2/4 (15% to 85%) 3/3 (43.9% to 100%) 2/3 (20.8% to 93.9%) 4 100%
deepseek v4 flash OpenCode · DeepSeek · xhigh4/7 (25% to 84.2%) 0/4 (0% to 49%) 1/3 (6.1% to 79.2%) 1/3 (6.1% to 79.2%) 4 100%
How these models were measured

Which harness a model ran under is a variable held beside it, not the thing this page compares. These sections read the harness: what it costs to find a blocking defect with one, how each has moved campaign by campaign, and what every configuration of one was priced at.

Cost of a blocking defect, by harness

Cheapest first · 17 of 75 defects are merge-blocking · newest of 31 campaigns, sweep-opencode-gpt-5.6-luna-max
OpenCode OpenAI · gpt 5.6 luna · effort max · 4 runs
$0.0226 $0.08 per must-fix bug found per review
Must-fix bugs found
88.2% (65.7% to 96.7%)
All bugs found
62.7% (51.4% to 72.7%)
Time per review
6m 4s
Extra findings
6.0% of findings
Completed
100% (51%–100%)
Spent
$0.3394

The same reviewers, campaign by campaign

Oldest first · the panels above are the last column · recall carries its 95% interval over the blocking defects; a cost carries the range between repeats
Measure
Blocking recall for every reviewer in every campaign, oldest campaign first
Reviewerci-glm-5.3-flash-high 4 per armsweep-opencode-glm-5.3-high 4 per armsweep-opencode-glm-5.3-flash-high 4 per armsweep-opencode-kimi-k3-high 4 per armsweep-opencode-deepseek-v4-flash-high 4 per armsweep-opencode-deepseek-v4.1-flash-high 4 per armsweep-opencode-glm-5.3-low 4 per armsweep-opencode-glm-5.3-medium 4 per armsweep-opencode-glm-5.3-xhigh 4 per armsweep-opencode-glm-5.3-max 4 per armsweep-opencode-glm-5.3-flash-low 4 per armsweep-opencode-glm-5.3-flash-medium 4 per armsweep-opencode-glm-5.3-flash-xhigh 4 per armsweep-opencode-glm-5.3-flash-max 4 per armsweep-opencode-kimi-k3-low 4 per armsweep-opencode-kimi-k3-medium 4 per armsweep-opencode-kimi-k3-xhigh 4 per armsweep-opencode-kimi-k3-max 4 per armsweep-opencode-deepseek-v4-flash-low 4 per armsweep-opencode-deepseek-v4-flash-medium 4 per armsweep-opencode-deepseek-v4-flash-xhigh 4 per armsweep-opencode-deepseek-v4-flash-max 4 per armsweep-opencode-deepseek-v4.1-flash-low 4 per armsweep-opencode-deepseek-v4.1-flash-medium 4 per armsweep-opencode-deepseek-v4.1-flash-xhigh 4 per armsweep-opencode-deepseek-v4.1-flash-max 4 per armsweep-opencode-gpt-5.6-luna-high 4 per armsweep-opencode-gpt-5.6-luna-low 4 per armsweep-opencode-gpt-5.6-luna-medium 4 per armsweep-opencode-gpt-5.6-luna-xhigh 4 per armsweep-opencode-gpt-5.6-luna-max 4 per arm
OpenCode76.5% 52.7% to 90.4% 100% completed70.6% 46.9% to 86.7% 100% completed76.5% 51.2% to 91% 100% completed70.6% 46.9% to 86.7% 100% completed23.5% 8% to 52% 100% completed70.6% 46.9% to 86.7% 100% completed70.6% 46.9% to 86.7% 100% completed58.8% 36% to 78.4% 100% completed70.6% 46.9% to 86.7% 100% completed76.5% 52.7% to 90.4% 100% completed41.2% 21.6% to 64% 100% completed70.6% 46.9% to 86.7% 100% completed70.6% 46.9% to 86.7% 100% completed76.5% 52.7% to 90.4% 100% completed52.9% 30.2% to 74.5% 100% completed70.6% 46.9% to 86.7% 100% completed76.5% 52.7% to 90.4% 100% completed76.5% 52.7% to 90.4% 100% completed11.8% 1.6% to 51.7% 100% completed23.5% 9.6% to 47.3% 100% completed35.3% 14.7% to 63.3% 100% completed17.6% 3.9% to 53.2% 100% completed64.7% 41.3% to 82.7% 100% completed70.6% 46.9% to 86.7% 100% completed76.5% 52.7% to 90.4% 100% completed88.2% 60.8% to 97.3% 100% completed82.4% 59% to 93.8% 100% completed35.3% 15.9% to 61.2% 100% completed70.6% 35.2% to 91.4% 100% completed76.5% 52.7% to 90.4% 100% completed88.2% 65.7% to 96.7% 100% completed

Across 31 campaigns OpenCode moved furthest on this measure, 76.5% to 88.2%.

Every configuration we could price

Colour is reviewer · size is effort · filled markers are on the frontier

What a configuration costs against what it catches

What a review is priced in
Which bugs count

One marker per configuration, each priced over every pull request it was run on. Filled markers are on the efficient frontier: nothing in the set is both cheaper and better. 1 of 1 configurations sit on it

85%86%88%89%91%$0.10OpenAI1Cost per review (USD, log scale)Must-fix bugs found
The efficient frontier for this measure, best first. Highlighted rows are efficient on cost and on time, and are the ones badged in the chart above
#ReviewerModelEffortPer reviewTimeMust-fixMatchedCompleted
1 OpenCodeOpenAI · gpt 5.6 lunamax$0.08486m 4s88%94%100%

What kind of defect it finds

Severity comes from the answer keys, not the reviewer

By severity

How much each bug matters

0%25%50%75%100%88.2%MUST-FIX17 bugsmust fix before merge58.7%ISSUE46 bugsshould fix41.7%SUGGESTION12 bugsoptional

By difficulty tier

How hard each one was to see

0%25%50%75%100%TIER 111TIER 230TIER 325TIER 49Tier 1 floor: every bug in the diff

All one do best on the defects that block a merge and worst on the optional ones, which is the order you want, though no reviewer’s blocking and optional intervals are far enough apart for that order to be a measured difference. Tier 1 is the floor, the 11 defects visible in the changed lines, and 1 of 1 clear all of it, from 100% to 100%. From tier 3 to tier 4 recall falls 25 points, which is the ordering the tiers were built to produce

What a review costs you, in money and in time

Both panels start at 52% · bars are 95% intervals over the blocking defects

Cost

Cheaper is left, better is up. Dashed rays are equal cost per must-fix bug, priced at the label

What is written beside each mark
60%70%80%90%$0.08$0.09$0.09OpenCodeCost of one reviewbars are 95% intervals over the blocking defects, 1 trials per pull requestMust-fix bugs found

Time

Faster is left

60%70%80%90%5m 40s6m 0s6m 20sOpenCodeTime for one reviewbars are 95% intervals over the blocking defects, 1 trials per pull requestMust-fix bugs found

Hollow rings are the 4 pull requests behind each solid marker, each one against its own bug count. 6 readings sit outside these frames and are not drawn

What more effort buys

Same reviewer, same role stack, same pull request · only the effort knob moved
No reviewer ran two effort settings on the same pull request, so there is no move to show
benchee benchee-dashboard-2 built from 56a02113 Static benchmark evidence ·