- Must-fix bugs found
- 94.1% (73% to 99%)
- All bugs found
- 57.3% (46.1% to 67.9%)
- Time per review
- 4m 27s
- Extra findings
- 2.3% of findings
- Completed
- 100% (51%–100%)
- Spent
- $0.3055
Benchmark
Which model reviews best, and what it costs
5 models on 1 harness against the same 4 pull requests, scored on 75 defects written down before any of them ran · measured since 2026-08-15, each model at the measurement that read best
Which model reviewed best
Initially ordered by blocking recall · 9 of 10 pairs overlap| deepseek v4.1 flash | 94.1% (73% to 99%) | 57.3% (39.4% to 73.5%) | N/A | N/A | 100% (51%–100%) | Unavailable | $0.0764 | 4m 27s | OpenCode · DeepSeek · max best of 5 efforts | 4 over 4 PRs |
|---|---|---|---|---|---|---|---|---|---|---|
| glm 5.3 tied not separated | 76.5% (52.7% to 90.4%) | 53.3% (37.4% to 68.6%) | N/A | N/A | 100% (51%–100%) | Unavailable | $0.1619 | 2m 27s | OpenCode · Z.AI · high best of 5 efforts | 4 over 4 PRs |
| glm 5.3 flash tied not separated | 76.5% (51.2% to 91%) | 46.7% (35.8% to 57.9%) | N/A | N/A | 100% (51%–100%) | Unavailable | $0.0146 | 3m 22s | OpenCode · Z.AI · high best of 5 efforts | 4 over 4 PRs |
| kimi k3 tied not separated | 76.5% (52.7% to 90.4%) | 48% (37.1% to 59.1%) | N/A | N/A | 100% (51%–100%) | Unavailable | $0.3355 | 3m 42s | OpenCode · Moonshot AI · xhigh best of 5 efforts | 4 over 4 PRs |
| deepseek v4 flash | 29.4% (9% to 63.7%) | 16% (5.6% to 37.9%) | Invalid | Invalid | 100% (51%–100%) | Unavailable | $0.0150 | 4m 13s | OpenCode · DeepSeek · xhigh best of 5 efforts | 4 over 4 PRs |
4 of 4 neighbouring pairs have overlapping intervals on this column, so the order between them is not a measured difference.
How each model has moved
F1 by the day it was measured · each model at effort high on the harness and vendor of its row above, whichever rung read best thereThis chart starts at 2026-08-15 and fills in as sweeps are recorded. No earlier day is on record.
The same models, one pull request at a time
4 pull requests measured · a model does not read the same on every diff- deepseek v4.1 flash OpenCode · DeepSeek · max
- glm 5.3 OpenCode · Z.AI · high
- glm 5.3 flash OpenCode · Z.AI · high
- kimi k3 OpenCode · Moonshot AI · xhigh
- deepseek v4 flash OpenCode · DeepSeek · xhigh
| Model | feat(web): shareable findings reports | feat(proxy): reclaim space from the blob store | feat(registry): saved views with CSV export | feat(scanner): scan the transitive graph | Trials on record | Completed |
|---|---|---|---|---|---|---|
| deepseek v4.1 flash OpenCode · DeepSeek · max | 6/7 (48.7% to 97.4%) | 4/4 (51% to 100%) | 3/3 (43.9% to 100%) | 3/3 (43.9% to 100%) | 4 | 100% |
| glm 5.3 OpenCode · Z.AI · high | 6/7 (48.7% to 97.4%) | 2/4 (15% to 85%) | 3/3 (43.9% to 100%) | 2/3 (20.8% to 93.9%) | 4 | 100% |
| glm 5.3 flash OpenCode · Z.AI · high | 6/7 (48.7% to 97.4%) | 3/4 (30.1% to 95.4%) | 3/3 (43.9% to 100%) | 1/3 (6.1% to 79.2%) | 4 | 100% |
| kimi k3 OpenCode · Moonshot AI · xhigh | 6/7 (48.7% to 97.4%) | 2/4 (15% to 85%) | 3/3 (43.9% to 100%) | 2/3 (20.8% to 93.9%) | 4 | 100% |
| deepseek v4 flash OpenCode · DeepSeek · xhigh | 4/7 (25% to 84.2%) | 0/4 (0% to 49%) | 1/3 (6.1% to 79.2%) | 0/3 (0% to 56.1%) | 4 | 100% |
How these models were measured
Which harness a model ran under is a variable held beside it, not the thing this page compares. These sections read the harness: what it costs to find a blocking defect with one, how each has moved campaign by campaign, and what every configuration of one was priced at.
Cost of a blocking defect, by harness
Cheapest first · 17 of 75 defects are merge-blocking · newest of 26 campaigns, sweep-opencode-deepseek-v4.1-flash-maxThe same reviewers, campaign by campaign
Oldest first · the panels above are the last column · recall carries its 95% interval over the blocking defects; a cost carries the range between repeats| Reviewer | ci-glm-5.3-flash-high 4 per arm | sweep-opencode-glm-5.3-high 4 per arm | sweep-opencode-glm-5.3-flash-high 4 per arm | sweep-opencode-kimi-k3-high 4 per arm | sweep-opencode-deepseek-v4-flash-high 4 per arm | sweep-opencode-deepseek-v4.1-flash-high 4 per arm | sweep-opencode-glm-5.3-low 4 per arm | sweep-opencode-glm-5.3-medium 4 per arm | sweep-opencode-glm-5.3-xhigh 4 per arm | sweep-opencode-glm-5.3-max 4 per arm | sweep-opencode-glm-5.3-flash-low 4 per arm | sweep-opencode-glm-5.3-flash-medium 4 per arm | sweep-opencode-glm-5.3-flash-xhigh 4 per arm | sweep-opencode-glm-5.3-flash-max 4 per arm | sweep-opencode-kimi-k3-low 4 per arm | sweep-opencode-kimi-k3-medium 4 per arm | sweep-opencode-kimi-k3-xhigh 4 per arm | sweep-opencode-kimi-k3-max 4 per arm | sweep-opencode-deepseek-v4-flash-low 4 per arm | sweep-opencode-deepseek-v4-flash-medium 4 per arm | sweep-opencode-deepseek-v4-flash-xhigh 4 per arm | sweep-opencode-deepseek-v4-flash-max 4 per arm | sweep-opencode-deepseek-v4.1-flash-low 4 per arm | sweep-opencode-deepseek-v4.1-flash-medium 4 per arm | sweep-opencode-deepseek-v4.1-flash-xhigh 4 per arm | sweep-opencode-deepseek-v4.1-flash-max 4 per arm |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| OpenCode | 76.5% 52.7% to 90.4% 100% completed | 76.5% 52.7% to 90.4% 100% completed | 76.5% 51.2% to 91% 100% completed | 70.6% 46.9% to 86.7% 100% completed | 23.5% 8% to 52% 100% completed | 70.6% 46.9% to 86.7% 100% completed | 70.6% 46.9% to 86.7% 100% completed | 58.8% 36% to 78.4% 100% completed | 76.5% 52.7% to 90.4% 100% completed | 76.5% 52.7% to 90.4% 100% completed | 41.2% 21.6% to 64% 100% completed | 64.7% 41.3% to 82.7% 100% completed | 76.5% 52.7% to 90.4% 100% completed | 76.5% 52.7% to 90.4% 100% completed | 52.9% 30.2% to 74.5% 100% completed | 70.6% 46.9% to 86.7% 100% completed | 76.5% 52.7% to 90.4% 100% completed | 76.5% 52.7% to 90.4% 100% completed | 11.8% 1.6% to 51.7% 100% completed | 23.5% 9.6% to 47.3% 100% completed | 29.4% 9% to 63.7% 100% completed | 17.6% 3.9% to 53.2% 100% completed | 64.7% 41.3% to 82.7% 100% completed | 64.7% 41.3% to 82.7% 100% completed | 76.5% 52.7% to 90.4% 100% completed | 94.1% 73% to 99% 100% completed |
Across 26 campaigns OpenCode moved furthest on this measure, 76.5% to 94.1%.
Every configuration we could price
Colour is reviewer · size is effort · filled markers are on the frontierWhat a configuration costs against what it catches
One marker per configuration, each priced over every pull request it was run on. Filled markers are on the efficient frontier: nothing in the set is both cheaper and better. 1 of 1 configurations sit on it
| # | Reviewer | Model | Effort | Per review | Time | Must-fix | Matched | Completed |
|---|---|---|---|---|---|---|---|---|
| 1 | OpenCode | DeepSeek · v4.1 flash | max | $0.0764 | 4m 27s | 94% | 98% | 100% |
What kind of defect it finds
Severity comes from the answer keys, not the reviewerBy severity
How much each bug matters
By difficulty tier
How hard each one was to see
All one do best on the defects that block a merge and worst on the optional ones, which is the order you want. Tier 1 is the floor, the 11 defects visible in the changed lines, and 1 of 1 clear all of it, from 100% to 100%. From tier 3 to tier 4 recall falls 38 points, which is the ordering the tiers were built to produce
What a review costs you, in money and in time
Both panels start at 52% · bars are 95% intervals over the blocking defectsCost
Cheaper is left, better is up. Dashed rays are equal cost per must-fix bug, priced at the label
Time
Faster is left
Hollow rings are the 4 pull requests behind each solid marker, each one against its own bug count. 8 readings sit outside these frames and are not drawn