Benchmark

Which model reviews best, and what it costs

Kind of work

5 models on 1 harness against the same 4 pull requests, scored on 75 defects written down before any of them ran · measured since 2026-08-15, each model at the measurement that read best

Which model reviewed best

Initially ordered by blocking recall · 10 of 10 pairs overlap
Value of F1 0 → 1
Every model on record; sortable by every column, initially ordered by blocking recall
glm 5.3 76.5% (52.7% to 90.4%) 53.3% (37.4% to 68.6%) N/A N/A 100% (51%–100%) Unavailable $0.1619 2m 27s OpenCode · Z.AI · high best of 5 efforts4 over 4 PRs
glm 5.3 flash tied not separated 76.5% (51.2% to 91%) 46.7% (35.8% to 57.9%) N/A N/A 100% (51%–100%) Unavailable $0.0146 3m 22s OpenCode · Z.AI · high best of 5 efforts4 over 4 PRs
kimi k3 tied not separated 76.5% (52.7% to 90.4%) 53.3% (40.4% to 65.9%) N/A N/A 100% (51%–100%) Unavailable $0.5156 9m 53s OpenCode · Moonshot AI · max best of 5 efforts4 over 4 PRs
deepseek v4.1 flash tied not separated 70.6% (46.9% to 86.7%) 44% (32.5% to 56.2%) N/A N/A 100% (51%–100%) Unavailable $0.0212 1m 8s OpenCode · DeepSeek · high 4 over 4 PRs
deepseek v4 flash tied not separated 29.4% (9% to 63.7%) 16% (5.6% to 37.9%) Invalid Invalid 100% (51%–100%) Unavailable $0.0150 4m 13s OpenCode · DeepSeek · xhigh best of 5 efforts4 over 4 PRs

4 of 4 neighbouring pairs have overlapping intervals on this column, so the order between them is not a measured difference.

How each model has moved

F1 by the day it was measured · each model at effort high on the harness and vendor of its row above, whichever rung read best there
Nothing measured since 2026-08-15

This chart starts at 2026-08-15 and fills in as sweeps are recorded. No earlier day is on record.

The same models, one pull request at a time

4 pull requests measured · a model does not read the same on every diff
What to read each model on
higher is better
0%50%100%feat(web): shareable findings reports dashboardfeat(proxy): reclaim space from the blob store proxyfeat(registry): saved views with CSV export registryfeat(scanner): scan the transitive graph scannerglm 5.3 glm 5.3kimi k3 kimi k3deepseek v4.1 flash deepseek v4.1 flashglm 5.3 flash glm 5.3 flashdeepseek v4 flash deepseek v4 flash
  • glm 5.3 OpenCode · Z.AI · high
  • glm 5.3 flash OpenCode · Z.AI · high
  • kimi k3 OpenCode · Moonshot AI · max
  • deepseek v4.1 flash OpenCode · DeepSeek · high
  • deepseek v4 flash OpenCode · DeepSeek · xhigh
Each model's reading on each pull request
Model feat(web): shareable findings reportsfeat(proxy): reclaim space from the blob storefeat(registry): saved views with CSV exportfeat(scanner): scan the transitive graph Trials on record Completed
glm 5.3 OpenCode · Z.AI · high6/7 (48.7% to 97.4%) 2/4 (15% to 85%) 3/3 (43.9% to 100%) 2/3 (20.8% to 93.9%) 4 100%
glm 5.3 flash OpenCode · Z.AI · high6/7 (48.7% to 97.4%) 3/4 (30.1% to 95.4%) 3/3 (43.9% to 100%) 1/3 (6.1% to 79.2%) 4 100%
kimi k3 OpenCode · Moonshot AI · max5/7 (35.9% to 91.8%) 3/4 (30.1% to 95.4%) 3/3 (43.9% to 100%) 2/3 (20.8% to 93.9%) 4 100%
deepseek v4.1 flash OpenCode · DeepSeek · high4/7 (25% to 84.2%) 3/4 (30.1% to 95.4%) 3/3 (43.9% to 100%) 2/3 (20.8% to 93.9%) 4 100%
deepseek v4 flash OpenCode · DeepSeek · xhigh4/7 (25% to 84.2%) 0/4 (0% to 49%) 1/3 (6.1% to 79.2%) 0/3 (0% to 56.1%) 4 100%
How these models were measured

Which harness a model ran under is a variable held beside it, not the thing this page compares. These sections read the harness: what it costs to find a blocking defect with one, how each has moved campaign by campaign, and what every configuration of one was priced at.

Cost of a blocking defect, by harness

Cheapest first · 17 of 75 defects are merge-blocking · newest of 22 campaigns, sweep-opencode-deepseek-v4-flash-max
OpenCode DeepSeek · v4 flash · effort max · 4 runs
$0.0174 $0.01 per must-fix bug found per review
Must-fix bugs found
17.6% (3.9% to 53.2%)
All bugs found
10.7% (5.5% to 19.7%)
Time per review
5m 19s
Extra findings
11.1% of findings
Completed
100% (51%–100%)
Spent
$0.0522

The same reviewers, campaign by campaign

Oldest first · the panels above are the last column · recall carries its 95% interval over the blocking defects; a cost carries the range between repeats
Measure
Blocking recall for every reviewer in every campaign, oldest campaign first
Reviewerci-glm-5.3-flash-high 4 per armsweep-opencode-glm-5.3-high 4 per armsweep-opencode-glm-5.3-flash-high 4 per armsweep-opencode-kimi-k3-high 4 per armsweep-opencode-deepseek-v4-flash-high 4 per armsweep-opencode-deepseek-v4.1-flash-high 4 per armsweep-opencode-glm-5.3-low 4 per armsweep-opencode-glm-5.3-medium 4 per armsweep-opencode-glm-5.3-xhigh 4 per armsweep-opencode-glm-5.3-max 4 per armsweep-opencode-glm-5.3-flash-low 4 per armsweep-opencode-glm-5.3-flash-medium 4 per armsweep-opencode-glm-5.3-flash-xhigh 4 per armsweep-opencode-glm-5.3-flash-max 4 per armsweep-opencode-kimi-k3-low 4 per armsweep-opencode-kimi-k3-medium 4 per armsweep-opencode-kimi-k3-xhigh 4 per armsweep-opencode-kimi-k3-max 4 per armsweep-opencode-deepseek-v4-flash-low 4 per armsweep-opencode-deepseek-v4-flash-medium 4 per armsweep-opencode-deepseek-v4-flash-xhigh 4 per armsweep-opencode-deepseek-v4-flash-max 4 per arm
OpenCode76.5% 52.7% to 90.4% 100% completed76.5% 52.7% to 90.4% 100% completed76.5% 51.2% to 91% 100% completed70.6% 46.9% to 86.7% 100% completed23.5% 8% to 52% 100% completed70.6% 46.9% to 86.7% 100% completed70.6% 46.9% to 86.7% 100% completed58.8% 36% to 78.4% 100% completed76.5% 52.7% to 90.4% 100% completed76.5% 52.7% to 90.4% 100% completed41.2% 21.6% to 64% 100% completed64.7% 41.3% to 82.7% 100% completed76.5% 52.7% to 90.4% 100% completed76.5% 52.7% to 90.4% 100% completed52.9% 30.2% to 74.5% 100% completed70.6% 46.9% to 86.7% 100% completed76.5% 52.7% to 90.4% 100% completed76.5% 52.7% to 90.4% 100% completed11.8% 1.6% to 51.7% 100% completed23.5% 9.6% to 47.3% 100% completed29.4% 9% to 63.7% 100% completed17.6% 3.9% to 53.2% 100% completed

Across 22 campaigns OpenCode moved furthest on this measure, 76.5% to 17.6%.

Every configuration we could price

Colour is reviewer · size is effort · filled markers are on the frontier

What a configuration costs against what it catches

What a review is priced in
Which bugs count

One marker per configuration, each priced over every pull request it was run on. Filled markers are on the efficient frontier: nothing in the set is both cheaper and better. 1 of 1 configurations sit on it

17%17%18%18%18%DeepSeek1Cost per review (USD, log scale)Must-fix bugs found
The efficient frontier for this measure, best first. Highlighted rows are efficient on cost and on time, and are the ones badged in the chart above
#ReviewerModelEffortPer reviewTimeMust-fixMatchedCompleted
1 OpenCodeDeepSeek · v4 flashmax$0.01305m 19s18%89%100%

What kind of defect it finds

Severity comes from the answer keys, not the reviewer

By severity

How much each bug matters

0%25%50%75%100%17.6%MUST-FIX17 bugsmust fix before merge10.9%ISSUE46 bugsshould fix0%SUGGESTION12 bugsoptional

By difficulty tier

How hard each one was to see

0%25%50%75%100%TIER 111TIER 230TIER 325TIER 49Tier 1 floor: every bug in the diff

All one do best on the defects that block a merge and worst on the optional ones, which is the order you want, though no reviewer’s blocking and optional intervals are far enough apart for that order to be a measured difference. Tier 1 is the floor, the 11 defects visible in the changed lines, and nobody clears all of it: 25% at best, 25% at worst. From tier 3 to tier 4 recall moves 0 points, less than the one defect this key can express, so the ladder stops sorting anything once the defect leaves the diff

What a review costs you, in money and in time

Both panels start at 52% · bars are 95% intervals over the blocking defects

Cost

Cheaper is left, better is up. Dashed rays are equal cost per must-fix bug, priced at the label

What is written beside each mark
25%50%75%$0.01$0.01$0.01$0.01OpenCodeCost of one reviewbars are 95% intervals over the blocking defects, 1 trials per pull requestMust-fix bugs found

Time

Faster is left

25%50%75%5m 0s5m 20s5m 40sOpenCodeTime for one reviewbars are 95% intervals over the blocking defects, 1 trials per pull requestMust-fix bugs found

Hollow rings are the 4 pull requests behind each solid marker, each one against its own bug count. 8 readings sit outside these frames and are not drawn

What more effort buys

Same reviewer, same role stack, same pull request · only the effort knob moved
No reviewer ran two effort settings on the same pull request, so there is no move to show
benchee benchee-dashboard-2 built from 8035e7de Static benchmark evidence ·