Benchmark

Which model reviews best, and what it costs

Kind of work

5 models on 1 harness against the same 4 pull requests, scored on 75 defects written down before any of them ran · measured since 2026-08-15, each model at the measurement that read best

Which model reviewed best

Initially ordered by blocking recall · 9 of 10 pairs overlap
Value of F1 0 → 1
Every model on record; sortable by every column, initially ordered by blocking recall
deepseek v4.1 flash 94.1% (73% to 99%) 57.3% (39.4% to 73.5%) N/A N/A 100% (51%–100%) Unavailable $0.0764 4m 27s OpenCode · DeepSeek · max best of 5 efforts4 over 4 PRs
glm 5.3 tied not separated 76.5% (52.7% to 90.4%) 53.3% (37.4% to 68.6%) N/A N/A 100% (51%–100%) Unavailable $0.1619 2m 27s OpenCode · Z.AI · high best of 5 efforts4 over 4 PRs
glm 5.3 flash tied not separated 76.5% (51.2% to 91%) 46.7% (35.8% to 57.9%) N/A N/A 100% (51%–100%) Unavailable $0.0146 3m 22s OpenCode · Z.AI · high best of 5 efforts4 over 4 PRs
kimi k3 tied not separated 76.5% (52.7% to 90.4%) 48% (37.1% to 59.1%) N/A N/A 100% (51%–100%) Unavailable $0.3355 3m 42s OpenCode · Moonshot AI · xhigh best of 5 efforts4 over 4 PRs
deepseek v4 flash 29.4% (9% to 63.7%) 16% (5.6% to 37.9%) Invalid Invalid 100% (51%–100%) Unavailable $0.0150 4m 13s OpenCode · DeepSeek · xhigh best of 5 efforts4 over 4 PRs

4 of 4 neighbouring pairs have overlapping intervals on this column, so the order between them is not a measured difference.

How each model has moved

F1 by the day it was measured · each model at effort high on the harness and vendor of its row above, whichever rung read best there
Nothing measured since 2026-08-15

This chart starts at 2026-08-15 and fills in as sweeps are recorded. No earlier day is on record.

The same models, one pull request at a time

4 pull requests measured · a model does not read the same on every diff
What to read each model on
higher is better
0%50%100%feat(web): shareable findings reports dashboardfeat(proxy): reclaim space from the blob store proxyfeat(registry): saved views with CSV export registryfeat(scanner): scan the transitive graph scannerdeepseek v4.1 flash deepseek v4.1 flashglm 5.3 glm 5.3kimi k3 kimi k3glm 5.3 flash glm 5.3 flashdeepseek v4 flash deepseek v4 flash
  • deepseek v4.1 flash OpenCode · DeepSeek · max
  • glm 5.3 OpenCode · Z.AI · high
  • glm 5.3 flash OpenCode · Z.AI · high
  • kimi k3 OpenCode · Moonshot AI · xhigh
  • deepseek v4 flash OpenCode · DeepSeek · xhigh
Each model's reading on each pull request
Model feat(web): shareable findings reportsfeat(proxy): reclaim space from the blob storefeat(registry): saved views with CSV exportfeat(scanner): scan the transitive graph Trials on record Completed
deepseek v4.1 flash OpenCode · DeepSeek · max6/7 (48.7% to 97.4%) 4/4 (51% to 100%) 3/3 (43.9% to 100%) 3/3 (43.9% to 100%) 4 100%
glm 5.3 OpenCode · Z.AI · high6/7 (48.7% to 97.4%) 2/4 (15% to 85%) 3/3 (43.9% to 100%) 2/3 (20.8% to 93.9%) 4 100%
glm 5.3 flash OpenCode · Z.AI · high6/7 (48.7% to 97.4%) 3/4 (30.1% to 95.4%) 3/3 (43.9% to 100%) 1/3 (6.1% to 79.2%) 4 100%
kimi k3 OpenCode · Moonshot AI · xhigh6/7 (48.7% to 97.4%) 2/4 (15% to 85%) 3/3 (43.9% to 100%) 2/3 (20.8% to 93.9%) 4 100%
deepseek v4 flash OpenCode · DeepSeek · xhigh4/7 (25% to 84.2%) 0/4 (0% to 49%) 1/3 (6.1% to 79.2%) 0/3 (0% to 56.1%) 4 100%
How these models were measured

Which harness a model ran under is a variable held beside it, not the thing this page compares. These sections read the harness: what it costs to find a blocking defect with one, how each has moved campaign by campaign, and what every configuration of one was priced at.

Cost of a blocking defect, by harness

Cheapest first · 17 of 75 defects are merge-blocking · newest of 26 campaigns, sweep-opencode-deepseek-v4.1-flash-max
OpenCode DeepSeek · v4.1 flash · effort max · 4 runs
$0.0191 $0.08 per must-fix bug found per review
Must-fix bugs found
94.1% (73% to 99%)
All bugs found
57.3% (46.1% to 67.9%)
Time per review
4m 27s
Extra findings
2.3% of findings
Completed
100% (51%–100%)
Spent
$0.3055

The same reviewers, campaign by campaign

Oldest first · the panels above are the last column · recall carries its 95% interval over the blocking defects; a cost carries the range between repeats
Measure
Blocking recall for every reviewer in every campaign, oldest campaign first
Reviewerci-glm-5.3-flash-high 4 per armsweep-opencode-glm-5.3-high 4 per armsweep-opencode-glm-5.3-flash-high 4 per armsweep-opencode-kimi-k3-high 4 per armsweep-opencode-deepseek-v4-flash-high 4 per armsweep-opencode-deepseek-v4.1-flash-high 4 per armsweep-opencode-glm-5.3-low 4 per armsweep-opencode-glm-5.3-medium 4 per armsweep-opencode-glm-5.3-xhigh 4 per armsweep-opencode-glm-5.3-max 4 per armsweep-opencode-glm-5.3-flash-low 4 per armsweep-opencode-glm-5.3-flash-medium 4 per armsweep-opencode-glm-5.3-flash-xhigh 4 per armsweep-opencode-glm-5.3-flash-max 4 per armsweep-opencode-kimi-k3-low 4 per armsweep-opencode-kimi-k3-medium 4 per armsweep-opencode-kimi-k3-xhigh 4 per armsweep-opencode-kimi-k3-max 4 per armsweep-opencode-deepseek-v4-flash-low 4 per armsweep-opencode-deepseek-v4-flash-medium 4 per armsweep-opencode-deepseek-v4-flash-xhigh 4 per armsweep-opencode-deepseek-v4-flash-max 4 per armsweep-opencode-deepseek-v4.1-flash-low 4 per armsweep-opencode-deepseek-v4.1-flash-medium 4 per armsweep-opencode-deepseek-v4.1-flash-xhigh 4 per armsweep-opencode-deepseek-v4.1-flash-max 4 per arm
OpenCode76.5% 52.7% to 90.4% 100% completed76.5% 52.7% to 90.4% 100% completed76.5% 51.2% to 91% 100% completed70.6% 46.9% to 86.7% 100% completed23.5% 8% to 52% 100% completed70.6% 46.9% to 86.7% 100% completed70.6% 46.9% to 86.7% 100% completed58.8% 36% to 78.4% 100% completed76.5% 52.7% to 90.4% 100% completed76.5% 52.7% to 90.4% 100% completed41.2% 21.6% to 64% 100% completed64.7% 41.3% to 82.7% 100% completed76.5% 52.7% to 90.4% 100% completed76.5% 52.7% to 90.4% 100% completed52.9% 30.2% to 74.5% 100% completed70.6% 46.9% to 86.7% 100% completed76.5% 52.7% to 90.4% 100% completed76.5% 52.7% to 90.4% 100% completed11.8% 1.6% to 51.7% 100% completed23.5% 9.6% to 47.3% 100% completed29.4% 9% to 63.7% 100% completed17.6% 3.9% to 53.2% 100% completed64.7% 41.3% to 82.7% 100% completed64.7% 41.3% to 82.7% 100% completed76.5% 52.7% to 90.4% 100% completed94.1% 73% to 99% 100% completed

Across 26 campaigns OpenCode moved furthest on this measure, 76.5% to 94.1%.

Every configuration we could price

Colour is reviewer · size is effort · filled markers are on the frontier

What a configuration costs against what it catches

What a review is priced in
Which bugs count

One marker per configuration, each priced over every pull request it was run on. Filled markers are on the efficient frontier: nothing in the set is both cheaper and better. 1 of 1 configurations sit on it

90%92%94%95%97%DeepSeek1Cost per review (USD, log scale)Must-fix bugs found
The efficient frontier for this measure, best first. Highlighted rows are efficient on cost and on time, and are the ones badged in the chart above
#ReviewerModelEffortPer reviewTimeMust-fixMatchedCompleted
1 OpenCodeDeepSeek · v4.1 flashmax$0.07644m 27s94%98%100%

What kind of defect it finds

Severity comes from the answer keys, not the reviewer

By severity

How much each bug matters

0%25%50%75%100%94.1%MUST-FIX17 bugsmust fix before merge52.2%ISSUE46 bugsshould fix25%SUGGESTION12 bugsoptional

By difficulty tier

How hard each one was to see

0%25%50%75%100%TIER 111TIER 230TIER 325TIER 49Tier 1 floor: every bug in the diff

All one do best on the defects that block a merge and worst on the optional ones, which is the order you want. Tier 1 is the floor, the 11 defects visible in the changed lines, and 1 of 1 clear all of it, from 100% to 100%. From tier 3 to tier 4 recall falls 38 points, which is the ordering the tiers were built to produce

What a review costs you, in money and in time

Both panels start at 52% · bars are 95% intervals over the blocking defects

Cost

Cheaper is left, better is up. Dashed rays are equal cost per must-fix bug, priced at the label

What is written beside each mark
60%70%80%90%$0.07$0.08OpenCodeCost of one reviewbars are 95% intervals over the blocking defects, 1 trials per pull requestMust-fix bugs found

Time

Faster is left

60%70%80%90%4m 10s4m 20s4m 30s4m 40sOpenCodeTime for one reviewbars are 95% intervals over the blocking defects, 1 trials per pull requestMust-fix bugs found

Hollow rings are the 4 pull requests behind each solid marker, each one against its own bug count. 8 readings sit outside these frames and are not drawn

What more effort buys

Same reviewer, same role stack, same pull request · only the effort knob moved
No reviewer ran two effort settings on the same pull request, so there is no move to show
benchee benchee-dashboard-2 built from f444508f Static benchmark evidence ·