Benchmark

Which model reviews best, and what it costs

Kind of work

5 models on 1 harness against the same 4 pull requests, scored on 75 defects written down before any of them ran · measured since 2026-08-15, each model at the measurement that read best

Which model reviewed best

Initially ordered by blocking recall · 10 of 10 pairs overlap
Value of F1 0 → 1
Every model on record; sortable by every column, initially ordered by blocking recall
deepseek v4.1 flash 88.2% (60.8% to 97.3%) 56% (38.2% to 72.4%) 95.5% (84.9% to 98.7%) 70.6% (52.7% to 83.5%) 100% (51%–100%) 68.3 V=$5.00 · review cost only$0.0768 4m 27s OpenCode · DeepSeek · max best of 5 efforts4 over 4 PRs
glm 5.3 tied not separated 76.5% (52.7% to 90.4%) 53.3% (35.6% to 70.3%) 88.9% (75.3% to 95.5%) 66.7% (48.3% to 81%) 100% (51%–100%) 60.8 V=$5.00 · review cost only$0.2747 4m 33s OpenCode · Z.AI · max best of 5 efforts4 over 4 PRs
glm 5.3 flash tied not separated 76.5% (51.2% to 91%) 44% (29.9% to 59.2%) 94.3% (70.9% to 99.1%) 60% (42% to 74.1%) 100% (51%–100%) 59.6 V=$5.00 · review cost only$0.0150 3m 22s OpenCode · Z.AI · high best of 5 efforts4 over 4 PRs
kimi k3 tied not separated 76.5% (52.7% to 90.4%) 45.3% (32.5% to 58.8%) 89.5% (75.9% to 95.8%) 60.2% (45.6% to 72.9%) 100% (51%–100%) 53.5 V=$5.00 · review cost only$0.3360 3m 42s OpenCode · Moonshot AI · xhigh best of 5 efforts4 over 4 PRs
deepseek v4 flash tied not separated 35.3% (14.7% to 63.3%) 21.3% (11.5% to 36%) 84.2% (40.3% to 97.7%) 34% (17.9% to 52.7%) 100% (51%–100%) 34.0 V=$5.00 · review cost only$0.0152 4m 13s OpenCode · DeepSeek · xhigh best of 5 efforts4 over 4 PRs

4 of 4 neighbouring pairs have overlapping intervals on this column, so the order between them is not a measured difference.

How each model has moved

F1 by the day it was measured · each model at effort high on the harness and vendor of its row above, whichever rung read best there
0%50%100%2026-09-292026-09-30deepseek v4.1 flash · 2026-09-30 · 60.9% · 4 trials over 4 pull requests · opencode opencode v2.0.18 · sweep-opencode-deepseek-v4.1-flash-highglm 5.3 · 2026-09-30 · 63.1% · 4 trials over 4 pull requests · opencode opencode v2.0.18 · sweep-opencode-glm-5.3-highglm 5.3 flash · 2026-09-29 · 62% · 4 trials over 4 pull requests · opencode opencode v2.0.18 · ci-glm-5.3-flash-highglm 5.3 flash · 2026-09-30 · 59.9% · 4 trials over 4 pull requests · opencode opencode v2.0.18 · sweep-opencode-glm-5.3-flash-highkimi k3 · 2026-09-30 · 62.1% · 4 trials over 4 pull requests · opencode opencode v2.0.18 · sweep-opencode-kimi-k3-highdeepseek v4 flash · 2026-09-30 · 24.2% · 4 trials over 4 pull requests · opencode opencode v2.0.18 · sweep-opencode-deepseek-v4-flash-highglm 5.3 glm 5.3kimi k3 kimi k3deepseek v4.1 flash deepseek v4.1 flashglm 5.3 flash glm 5.3 flashdeepseek v4 flash deepseek v4 flash

4 of 5 models were measured on one day, and a single mark is a reading rather than a trend.

Each model's average, first and latest F1
Model Average First Latest Days Completed Measured as
deepseek v4.1 flash 60.9% 60.9% 2026-09-30 one day only1 100% OpenCode · DeepSeek · max
glm 5.3 63.1% 63.1% 2026-09-30 one day only1 100% OpenCode · Z.AI · max
glm 5.3 flash 61% 62% 2026-09-2959.9% 2026-09-302 100% OpenCode · Z.AI · high
kimi k3 62.1% 62.1% 2026-09-30 one day only1 100% OpenCode · Moonshot AI · xhigh
deepseek v4 flash 24.2% 24.2% 2026-09-30 one day only1 100% OpenCode · DeepSeek · xhigh

The same models, one pull request at a time

4 pull requests measured · a model does not read the same on every diff
What to read each model on
higher is better
0%50%100%feat(web): shareable findings reports dashboardfeat(proxy): reclaim space from the blob store proxyfeat(registry): saved views with CSV export registryfeat(scanner): scan the transitive graph scannerdeepseek v4.1 flash deepseek v4.1 flashglm 5.3 glm 5.3kimi k3 kimi k3glm 5.3 flash glm 5.3 flashdeepseek v4 flash deepseek v4 flash
  • deepseek v4.1 flash OpenCode · DeepSeek · max
  • glm 5.3 OpenCode · Z.AI · max
  • glm 5.3 flash OpenCode · Z.AI · high
  • kimi k3 OpenCode · Moonshot AI · xhigh
  • deepseek v4 flash OpenCode · DeepSeek · xhigh
Each model's reading on each pull request
Model feat(web): shareable findings reportsfeat(proxy): reclaim space from the blob storefeat(registry): saved views with CSV exportfeat(scanner): scan the transitive graph Trials on record Completed
deepseek v4.1 flash OpenCode · DeepSeek · max5/7 (35.9% to 91.8%) 4/4 (51% to 100%) 3/3 (43.9% to 100%) 3/3 (43.9% to 100%) 4 100%
glm 5.3 OpenCode · Z.AI · max5/7 (35.9% to 91.8%) 3/4 (30.1% to 95.4%) 3/3 (43.9% to 100%) 2/3 (20.8% to 93.9%) 4 100%
glm 5.3 flash OpenCode · Z.AI · high6/7 (48.7% to 97.4%) 3/4 (30.1% to 95.4%) 3/3 (43.9% to 100%) 1/3 (6.1% to 79.2%) 4 100%
kimi k3 OpenCode · Moonshot AI · xhigh6/7 (48.7% to 97.4%) 2/4 (15% to 85%) 3/3 (43.9% to 100%) 2/3 (20.8% to 93.9%) 4 100%
deepseek v4 flash OpenCode · DeepSeek · xhigh4/7 (25% to 84.2%) 0/4 (0% to 49%) 1/3 (6.1% to 79.2%) 1/3 (6.1% to 79.2%) 4 100%
How these models were measured

Which harness a model ran under is a variable held beside it, not the thing this page compares. These sections read the harness: what it costs to find a blocking defect with one, how each has moved campaign by campaign, and what every configuration of one was priced at.

Cost of a blocking defect, by harness

Cheapest first · 17 of 75 defects are merge-blocking · newest of 26 campaigns, sweep-opencode-deepseek-v4.1-flash-max
OpenCode DeepSeek · v4.1 flash · effort max · 4 runs
$0.0205 $0.08 per must-fix bug found per review
Must-fix bugs found
88.2% (60.8% to 97.3%)
All bugs found
56% (44.7% to 66.7%)
Time per review
4m 27s
Extra findings
4.5% of findings
Completed
100% (51%–100%)
Spent
$0.3074

The same reviewers, campaign by campaign

Oldest first · the panels above are the last column · recall carries its 95% interval over the blocking defects; a cost carries the range between repeats
Measure
Blocking recall for every reviewer in every campaign, oldest campaign first
Reviewerci-glm-5.3-flash-high 4 per armsweep-opencode-glm-5.3-high 4 per armsweep-opencode-glm-5.3-flash-high 4 per armsweep-opencode-kimi-k3-high 4 per armsweep-opencode-deepseek-v4-flash-high 4 per armsweep-opencode-deepseek-v4.1-flash-high 4 per armsweep-opencode-glm-5.3-low 4 per armsweep-opencode-glm-5.3-medium 4 per armsweep-opencode-glm-5.3-xhigh 4 per armsweep-opencode-glm-5.3-max 4 per armsweep-opencode-glm-5.3-flash-low 4 per armsweep-opencode-glm-5.3-flash-medium 4 per armsweep-opencode-glm-5.3-flash-xhigh 4 per armsweep-opencode-glm-5.3-flash-max 4 per armsweep-opencode-kimi-k3-low 4 per armsweep-opencode-kimi-k3-medium 4 per armsweep-opencode-kimi-k3-xhigh 4 per armsweep-opencode-kimi-k3-max 4 per armsweep-opencode-deepseek-v4-flash-low 4 per armsweep-opencode-deepseek-v4-flash-medium 4 per armsweep-opencode-deepseek-v4-flash-xhigh 4 per armsweep-opencode-deepseek-v4-flash-max 4 per armsweep-opencode-deepseek-v4.1-flash-low 4 per armsweep-opencode-deepseek-v4.1-flash-medium 4 per armsweep-opencode-deepseek-v4.1-flash-xhigh 4 per armsweep-opencode-deepseek-v4.1-flash-max 4 per arm
OpenCode76.5% 52.7% to 90.4% 100% completed70.6% 46.9% to 86.7% 100% completed76.5% 51.2% to 91% 100% completed70.6% 46.9% to 86.7% 100% completed23.5% 8% to 52% 100% completed70.6% 46.9% to 86.7% 100% completed70.6% 46.9% to 86.7% 100% completed58.8% 36% to 78.4% 100% completed70.6% 46.9% to 86.7% 100% completed76.5% 52.7% to 90.4% 100% completed41.2% 21.6% to 64% 100% completed70.6% 46.9% to 86.7% 100% completed70.6% 46.9% to 86.7% 100% completed76.5% 52.7% to 90.4% 100% completed52.9% 30.2% to 74.5% 100% completed70.6% 46.9% to 86.7% 100% completed76.5% 52.7% to 90.4% 100% completed76.5% 52.7% to 90.4% 100% completed11.8% 1.6% to 51.7% 100% completed23.5% 9.6% to 47.3% 100% completed35.3% 14.7% to 63.3% 100% completed17.6% 3.9% to 53.2% 100% completed64.7% 41.3% to 82.7% 100% completed70.6% 46.9% to 86.7% 100% completed76.5% 52.7% to 90.4% 100% completed88.2% 60.8% to 97.3% 100% completed

Across 26 campaigns OpenCode moved furthest on this measure, 76.5% to 88.2%.

Every configuration we could price

Colour is reviewer · size is effort · filled markers are on the frontier

What a configuration costs against what it catches

What a review is priced in
Which bugs count

One marker per configuration, each priced over every pull request it was run on. Filled markers are on the efficient frontier: nothing in the set is both cheaper and better. 1 of 1 configurations sit on it

85%86%88%89%91%DeepSeek1Cost per review (USD, log scale)Must-fix bugs found
The efficient frontier for this measure, best first. Highlighted rows are efficient on cost and on time, and are the ones badged in the chart above
#ReviewerModelEffortPer reviewTimeMust-fixMatchedCompleted
1 OpenCodeDeepSeek · v4.1 flashmax$0.07684m 27s88%95%100%

What kind of defect it finds

Severity comes from the answer keys, not the reviewer

By severity

How much each bug matters

0%25%50%75%100%88.2%MUST-FIX17 bugsmust fix before merge52.2%ISSUE46 bugsshould fix25%SUGGESTION12 bugsoptional

By difficulty tier

How hard each one was to see

0%25%50%75%100%TIER 111TIER 230TIER 325TIER 49Tier 1 floor: every bug in the diff

All one do best on the defects that block a merge and worst on the optional ones, which is the order you want, though no reviewer’s blocking and optional intervals are far enough apart for that order to be a measured difference. Tier 1 is the floor, the 11 defects visible in the changed lines, and 1 of 1 clear all of it, from 100% to 100%. From tier 3 to tier 4 recall falls 39 points, which is the ordering the tiers were built to produce

What a review costs you, in money and in time

Both panels start at 52% · bars are 95% intervals over the blocking defects

Cost

Cheaper is left, better is up. Dashed rays are equal cost per must-fix bug, priced at the label

What is written beside each mark
60%70%80%90%$0.07$0.08OpenCodeCost of one reviewbars are 95% intervals over the blocking defects, 1 trials per pull requestMust-fix bugs found

Time

Faster is left

60%70%80%90%4m 10s4m 20s4m 30s4m 40sOpenCodeTime for one reviewbars are 95% intervals over the blocking defects, 1 trials per pull requestMust-fix bugs found

Hollow rings are the 4 pull requests behind each solid marker, each one against its own bug count. 8 readings sit outside these frames and are not drawn

What more effort buys

Same reviewer, same role stack, same pull request · only the effort knob moved
No reviewer ran two effort settings on the same pull request, so there is no move to show
benchee benchee-dashboard-2 built from b926ed28 Static benchmark evidence ·