- Must-fix bugs found
- 76.5% (52.7% to 90.4%)
- All bugs found
- 46.7% (35.8% to 57.8%)
- Time per review
- 3m 44s
- Extra findings
- 7.9% of findings
- Completed
- 100% (51%–100%)
- Spent
- $0.0638
Benchmark
Which model reviews best, and what it costs
1 model on 1 harness against the same 4 pull requests, scored on 75 defects written down before any of them ran · measured since 2026-08-15, each model at the measurement that read best
Which model reviewed best
Initially ordered by blocking recall · 1 of 1 model carries a ranked recall| glm 5.3 flash | 5/7 (35.9% to 91.8%) | 52.9% (31% to 73.8%) | N/A | N/A | 100% (20.7%–100%) | Unavailable | $0.0131 | 5m 26s | OpenCode · Z.AI · high | 1 over 1 PR |
|---|
How each model has moved
F1 by the day it was measured · each model at effort high on the harness and vendor of its row above, whichever rung read best thereThis chart starts at 2026-08-15 and fills in as sweeps are recorded. No earlier day is on record.
The same models, one pull request at a time
1 pull request measured · a model does not read the same on every diff- glm 5.3 flash OpenCode · Z.AI · high
| Model | feat(web): shareable findings reports | Trials on record | Completed |
|---|---|---|---|
| glm 5.3 flash OpenCode · Z.AI · high | 5/7 (35.9% to 91.8%) | 1 | 100% |
How these models were measured
Which harness a model ran under is a variable held beside it, not the thing this page compares. These sections read the harness: what it costs to find a blocking defect with one, how each has moved campaign by campaign, and what every configuration of one was priced at.
Cost of a blocking defect, by harness
Cheapest first · 17 of 75 defects are merge-blocking · campaign ci-glm-5.3-flash-highEvery configuration we could price
Colour is reviewer · size is effort · filled markers are on the frontierWhat a configuration costs against what it catches
One marker per configuration, each priced over every pull request it was run on. Filled markers are on the efficient frontier: nothing in the set is both cheaper and better. 3 of 4 configurations sit on it
| # | Reviewer | Model | Effort | Per review | Time | Must-fix | Matched | Completed |
|---|---|---|---|---|---|---|---|---|
| 1 | OpenCode | Z.AI · glm 5.3 flash | high | $0.0131 | 5m 26s | 71% | 92% | 100% |
| 2 | OpenCode | Z.AI · glm 5.3 flash | high | $0.0163 | 4m 11s | 75% | 92% | 100% |
| 3 | OpenCode | Z.AI · glm 5.3 flash | high | $0.0201 | 2m 9s | 100% | 92% | 100% |
What kind of defect it finds
Severity comes from the answer keys, not the reviewerBy severity
How much each bug matters
By difficulty tier
How hard each one was to see
All one do best on the defects that block a merge and worst on the optional ones, which is the order you want, though no reviewer’s blocking and optional intervals are far enough apart for that order to be a measured difference. Tier 1 is the floor, the 11 defects visible in the changed lines, and 1 of 1 clear all of it, from 100% to 100%. From tier 3 to tier 4 recall falls 25 points, which is the ordering the tiers were built to produce
What a review costs you, in money and in time
Both panels start at 52% · bars are 95% intervals over the blocking defectsCost
Cheaper is left, better is up. Dashed rays are equal cost per must-fix bug, priced at the label
Time
Faster is left
Hollow rings are the 4 pull requests behind each solid marker, each one against its own bug count. 7 readings sit outside these frames and are not drawn