- Must-fix bugs found
- 75% (53.1% to 88.8%)
- All bugs found
- 45.8% (35.5% to 56.4%)
- Time per review
- 2m 58s
- Extra findings
- 13.6% of findings
- Completed
- 100% (64.6%–100%)
- Spent
- $0.0823
Benchmark
Which model reviews best, and what it costs
1 model on 1 harness against the same 7 pull requests, scored on 83 defects written down before any of them ran · measured since 2026-08-15, each model at the measurement that read best
Which model reviewed best
Initially ordered by blocking recall · 1 of 1 model carries a ranked recall| glm 5.3 flash | 5/7 (35.9% to 91.8%) | 52.9% (31% to 73.8%) | N/A | N/A | 100% (20.7%–100%) | Unavailable | $0.0131 | 5m 26s | OpenCode · Z.AI · high | 1 over 1 PR |
|---|
How each model has moved
F1 by the day it was measured · each model at effort high on the harness and vendor of its row above, whichever rung read best thereThis chart starts at 2026-08-15 and fills in as sweeps are recorded. No earlier day is on record.
The same models, one pull request at a time
1 pull request measured · a model does not read the same on every diff- glm 5.3 flash OpenCode · Z.AI · high
| Model | feat(web): shareable findings reports | Trials on record | Completed |
|---|---|---|---|
| glm 5.3 flash OpenCode · Z.AI · high | 5/7 (35.9% to 91.8%) | 1 | 100% |
How these models were measured
Which harness a model ran under is a variable held beside it, not the thing this page compares. These sections read the harness: what it costs to find a blocking defect with one, how each has moved campaign by campaign, and what every configuration of one was priced at.
Cost of a blocking defect, by harness
Cheapest first · 20 of 83 defects are merge-blocking · campaign ci-glm-5.3-flash-highEvery configuration we could price
Colour is reviewer · size is effort · filled markers are on the frontierWhat a configuration costs against what it catches
One marker per configuration, each priced over every pull request it was run on. Filled markers are on the efficient frontier: nothing in the set is both cheaper and better. 4 of 5 configurations sit on it
| # | Reviewer | Model | Effort | Per review | Time | Must-fix | Matched | Completed |
|---|---|---|---|---|---|---|---|---|
| 1 | OpenCode | Z.AI · glm 5.3 flash | high | $0.004836 | 1m 27s | 67% | 86% | 100% |
| 2 | OpenCode | Z.AI · glm 5.3 flash | high | $0.0131 | 5m 26s | 71% | 86% | 100% |
| 3 | OpenCode | Z.AI · glm 5.3 flash | high | $0.0163 | 4m 11s | 75% | 86% | 100% |
| 4 | OpenCode | Z.AI · glm 5.3 flash | high | $0.0201 | 2m 9s | 100% | 86% | 100% |
What kind of defect it finds
Severity comes from the answer keys, not the reviewerBy severity
How much each bug matters
By difficulty tier
How hard each one was to see
What a review costs you, in money and in time
Both panels start at 52% · bars are 95% intervals over the blocking defectsCost
Cheaper is left, better is up. Dashed rays are equal cost per must-fix bug, priced at the label
Time
Faster is left
Hollow rings are the 7 pull requests behind each solid marker, each one against its own bug count. 9 readings sit outside these frames and are not drawn