- Must-fix bugs found
- 86.1%
- All bugs found
- 38.3%
- Time per review
- 3m 48s
- Extra findings
- 9.2% of findings
Benchmark
Model benchmark
23 models · 1 PRs · 20 keyed bugs
What one must-fix bug costs to find
Cheapest first · 4 of 20 planted bugs are must-fix- Must-fix bugs found
- 50%
- All bugs found
- 41.7%
- Time per review
- 3m 45s
- Extra findings
- 0.0% of findings
- Must-fix bugs found
- 25%
- All bugs found
- 15%
- Time per review
- 3m 9s
- Extra findings
- 10.0% of findings
- Must-fix bugs found
- 100%
- All bugs found
- 56.7%
- Time per review
- 3m 14s
- Extra findings
- 0.0% of findings
Showing the 4 cheapest per must-fix bug · 19 more models are plotted below and ranked in full on Compare
Every configuration priced
Colour is model · size is effort · filled markers are on the frontierWhat a configuration costs against what it catches
One marker per configuration, each priced over every pull request it was run on. Filled markers are on the efficient frontier: nothing in the set is both cheaper and better. 2 of 19 configurations sit on it, and 3 fell below the 5% floor and are not drawn
| # | Reviewer | Model | Effort | Per review | Time | Must-fix | Matched |
|---|---|---|---|---|---|---|---|
| 1 | Z.AI · glm 5.3 flash | Z.AI · glm 5.3 flash | high | $0.0207 | 3m 48s | 86% | 91% |
| 2 | OpenAI · gpt 5.6 sol | OpenAI · gpt 5.6 sol | high | $0.2556 | 3m 14s | 100% | 100% |
What kind of bug it finds
How serious each bug is comes from the answer key, not the reviewerBy severity
How much each bug matters
By difficulty tier
How hard each one was to see
3 reviewers are not drawn, reading under the 5% floor
15 of 23 do best on the bugs that block a merge and worst on the optional ones. Tier 1 is the floor, the 3 bugs visible in the changed lines, and 1 of 19 clear all of it, from 0% to 100%. From tier 3 to tier 4 recall moves 9 points, less than the one bug this key can express
What a review costs, in money and in time
Both panels start at 52% · bars are 95% intervalsCost
Cheaper is left, better is up. Dashed rays are equal cost per must-fix bug, priced at the label
Time
Faster is left
Hollow rings are the 1 pull requests behind each solid marker, each one against its own bug count. 3 configurations are not plotted, having found too little to price
On must-fix bugs, vertical bars across these 19 overlap in 92 of the 171 pairs. Price runs $0.01 to $2.06 a review and the wait runs 42s to 17m 51s, and on those price bars overlap in 54 of the 171 pairs while wait bars overlap in 116 of the 171 pairs
Bugs found across effort settings
Same model stack, same role stack, same pull request · only the effort knob movedDetail
Repeat the same review and you get a different answer
Same model stack, same pull request, same settings, run 2 to 9 times. The bars below span what those repeats actually returned, and the median cell moves 5% between its best and worst run while the widest moves 55%
Two model stacks on one pull request are comparable. Two pull requests are not: defect density differs across these 1
Trial by trial: the spread behind one number
Z.AI · glm 5.3 flash on proxy,
9 scored trials. Each dot is one review
Recall by tier, per pull request
The same numbers per pull request, where the variation lives. Each model stack opens with its pooled row
| Model stack and pull request | Tier 1 | Tier 2 | Tier 3 | Tier 4 |
|---|---|---|---|---|
| Z.AI · glm 5.3 flash all 1 pull requests | 33.3% | 42.9% | 50% | 50% |
| feat(proxy): reclaim space from the blob store | 29.6% | 33.3% | 48.1% | 38.9% |
| OpenAI · gpt 5.6 luna all 1 pull requests | 66.7% | 28.6% | 33.3% | 50% |
| feat(proxy): reclaim space from the blob store | 66.7% | 33.3% | 38.9% | 41.7% |
| Minimax · m3 all 1 pull requests | 33.3% | 0% | 16.7% | 0% |
| feat(proxy): reclaim space from the blob store | 22.2% | 9.5% | 22.2% | 8.3% |
| OpenAI · gpt 5.6 sol all 1 pull requests | 100% | 42.9% | 50% | 50% |
| feat(proxy): reclaim space from the blob store | 100% | 47.6% | 50% | 50% |
| xAI · grok 4.5 all 1 pull requests | 66.7% | 28.6% | 50% | 25% |
| feat(proxy): reclaim space from the blob store | 66.7% | 33.3% | 50% | 25% |
| OpenAI · gpt 5.6 terra all 1 pull requests | 33.3% | 28.6% | 16.7% | 50% |
| feat(proxy): reclaim space from the blob store | 33.3% | 23.8% | 27.8% | 50% |
| Moonshot AI · kimi k3 all 1 pull requests | 33.3% | 28.6% | 50% | 25% |
| feat(proxy): reclaim space from the blob store | 44.4% | 28.6% | 44.4% | 25% |
| Google · gemini 3.1 pro preview all 1 pull requests | 33.3% | 42.9% | 33.3% | 25% |
| feat(proxy): reclaim space from the blob store | 33.3% | 42.9% | 27.8% | 33.3% |
| Z.AI · glm 5.2 all 1 pull requests | 0% | 14.3% | 33.3% | 0% |
| feat(proxy): reclaim space from the blob store | 0% | 19% | 27.8% | 8.3% |
| Moonshot AI · kimi k2.7 code all 1 pull requests | 33.3% | 28.6% | 33.3% | 25% |
| feat(proxy): reclaim space from the blob store | 44.4% | 28.6% | 38.9% | 33.3% |
| Anthropic · opus 5 all 1 pull requests | 83.3% | 28.6% | 50% | 50% |
| feat(proxy): reclaim space from the blob store | 83.3% | 28.6% | 50% | 50% |
| DeepSeek · v4 pro all 1 pull requests | 66.7% | 0% | 33.3% | 0% |
| feat(proxy): reclaim space from the blob store | 55.6% | 19% | 38.9% | 16.7% |
| Z.AI · glm 5.3 all 1 pull requests | 0% | 0% | 16.7% | 0% |
| feat(proxy): reclaim space from the blob store | 11.1% | 19% | 16.7% | 8.3% |
| Google · gemini 3.7 flash all 1 pull requests | 33.3% | 28.6% | 33.3% | 0% |
| feat(proxy): reclaim space from the blob store | 33.3% | 23.8% | 27.8% | 16.7% |
| Anthropic · fable 5 all 1 pull requests | 33.3% | 14.3% | 50% | 50% |
| feat(proxy): reclaim space from the blob store | 44.4% | 23.8% | 50% | 50% |
| Anthropic · sonnet 5 all 1 pull requests | 33.3% | 14.3% | 25% | 0% |
| feat(proxy): reclaim space from the blob store | 33.3% | 14.3% | 25% | 0% |
| DeepSeek · v4 flash all 1 pull requests | 0% | 0% | 0% | 0% |
| feat(proxy): reclaim space from the blob store | 0% | 4.8% | 0% | 0% |
| Meta · llama 4 maverick all 0 pull requests | N/A | N/A | N/A | N/A |
| Mistral AI · devstral 2512 all 0 pull requests | N/A | N/A | N/A | N/A |
| OpenAI · gpt oss 120b all 1 pull requests | 0% | 0% | 0% | 0% |
| feat(proxy): reclaim space from the blob store | 0% | 0% | 0% | 0% |
| Qwen · qwen3 coder 30b a3b instruct all 0 pull requests | N/A | N/A | N/A | N/A |
| Qwen · qwen3.7 flash all 0 pull requests | N/A | N/A | N/A | N/A |
| Z.AI · glm 4.7 flash all 1 pull requests | 0% | 0% | 0% | 0% |
| feat(proxy): reclaim space from the blob store | 0% | 0% | 0% | 0% |
Weighted composite, if you want one number
A stated weighting, not a measured one: the answer keys rank severity and attach no number to it. Blocking 5 · issue 2 · suggestion 1, so the score is what a model stack earned over what the keys put on the table
| Model stack | Composite | Cost per weighted point |
|---|---|---|
| Z.AI · glm 5.3 flash | 54.4% | $0.000808 |
| OpenAI · gpt 5.6 luna | 46.1% | $0.001823 |
| Minimax · m3 | 17.7% | $0.007538 |
| OpenAI · gpt 5.6 sol | 70.9% | $0.007669 |
| xAI · grok 4.5 | 49.6% | $0.007958 |
| OpenAI · gpt 5.6 terra | 41.8% | $0.0122 |
| Moonshot AI · kimi k3 | 41.1% | $0.0167 |
| Google · gemini 3.1 pro preview | 46.1% | $0.0221 |
| Z.AI · glm 5.2 | 20.6% | $0.0192 |
| Moonshot AI · kimi k2.7 code | 39.7% | $0.0203 |
| Anthropic · opus 5 | 63.8% | $0.0258 |
| DeepSeek · v4 pro | 27.7% | $0.0103 |
| Z.AI · glm 5.3 | 15.6% | $0.0287 |
| Google · gemini 3.7 flash | 29.8% | $0.0367 |
| Anthropic · fable 5 | 54.6% | $0.0802 |
| Anthropic · sonnet 5 | 18.1% | $0.0855 |
| DeepSeek · v4 flash | 1.4% | $0.0212 |
| OpenAI · gpt oss 120b | 0% | N/A |
| Z.AI · glm 4.7 flash | 0% | N/A |
How much of what it reports is real
Precision needs a judge qualified to decide the pairs locality cannot, and that pass is not approved. Until it runs, the only noise signal is the share of findings that matched no keyed defect, which is not the same claim
| Model stack | Judged precision | Reviews judged | Unmatched share |
|---|---|---|---|
| Z.AI · glm 5.3 flash | N/A | 0 of 9 | 9.2% |
| OpenAI · gpt 5.6 luna | N/A | 0 of 3 | 0% |
| Minimax · m3 | N/A | 0 of 3 | 10% |
| OpenAI · gpt 5.6 sol | N/A | 0 of 3 | 0% |
| xAI · grok 4.5 | N/A | 0 of 3 | 0% |
| OpenAI · gpt 5.6 terra | N/A | 0 of 3 | 0% |
| Moonshot AI · kimi k3 | N/A | 0 of 3 | 4.5% |
| Google · gemini 3.1 pro preview | N/A | 0 of 3 | 4.5% |
| Z.AI · glm 5.2 | N/A | 0 of 3 | 0% |
| Moonshot AI · kimi k2.7 code | N/A | 0 of 3 | 15.4% |
| Anthropic · opus 5 | N/A | 0 of 2 | 0% |
| DeepSeek · v4 pro | N/A | 0 of 3 | 34.4% |
| Z.AI · glm 5.3 | N/A | 0 of 3 | 83.6% |
| Google · gemini 3.7 flash | N/A | 0 of 3 | 0% |
| Anthropic · fable 5 | N/A | 0 of 3 | 0% |
| Anthropic · sonnet 5 | N/A | 0 of 2 | 0% |
| DeepSeek · v4 flash | N/A | 0 of 3 | 75% |
| Meta · llama 4 maverick | N/A | 0 of 0 | N/A |
| Mistral AI · devstral 2512 | N/A | 0 of 0 | N/A |
| OpenAI · gpt oss 120b | N/A | 0 of 3 | 100% |
| Qwen · qwen3 coder 30b a3b instruct | N/A | 0 of 0 | N/A |
| Qwen · qwen3.7 flash | N/A | 0 of 0 | N/A |
| Z.AI · glm 4.7 flash | N/A | 0 of 2 | 100% |
Every configuration on record
| Model stack | Model | Effort | Roles | Scored runs |
|---|---|---|---|---|
| Z.AI · glm 5.3 flash | Z.AI · glm 5.3 flash | high | agent | 9 |
| OpenAI · gpt 5.6 luna | OpenAI · gpt 5.6 luna | high | agent | 3 |
| Minimax · m3 | Minimax · m3 | high | agent | 3 |
| OpenAI · gpt 5.6 sol | OpenAI · gpt 5.6 sol | high | agent | 3 |
| xAI · grok 4.5 | xAI · grok 4.5 | high | agent | 3 |
| OpenAI · gpt 5.6 terra | OpenAI · gpt 5.6 terra | high | agent | 3 |
| Moonshot AI · kimi k3 | Moonshot AI · kimi k3 | high | agent | 3 |
| Google · gemini 3.1 pro preview | Google · gemini 3.1 pro preview | high | agent | 3 |
| Z.AI · glm 5.2 | Z.AI · glm 5.2 | high | agent | 3 |
| Moonshot AI · kimi k2.7 code | Moonshot AI · kimi k2.7 code | high | agent | 3 |
| Anthropic · opus 5 | Anthropic · opus 5 | high | agent | 2 |
| DeepSeek · v4 pro | DeepSeek · v4 pro | high | agent | 3 |
| Z.AI · glm 5.3 | Z.AI · glm 5.3 | high | agent | 3 |
| Google · gemini 3.7 flash | Google · gemini 3.7 flash | high | agent | 3 |
| Anthropic · fable 5 | Anthropic · fable 5 | high | agent | 3 |
| Anthropic · sonnet 5 | Anthropic · sonnet 5 | high | agent | 2 |
| DeepSeek · v4 flash | DeepSeek · v4 flash | high | agent | 3 |
| Meta · llama 4 maverick | Meta · llama 4 maverick | high | agent | 0 |
| Mistral AI · devstral 2512 | Mistral AI · devstral 2512 | high | agent | 0 |
| OpenAI · gpt oss 120b | OpenAI · gpt oss 120b | high | agent | 3 |
| Qwen · qwen3 coder 30b a3b instruct | Qwen · qwen3 coder 30b a3b instruct | high | agent | 0 |
| Qwen · qwen3.7 flash | Qwen · qwen3.7 flash | high | agent | 0 |
| Z.AI · glm 4.7 flash | Z.AI · glm 4.7 flash | high | agent | 2 |
Provenance
What this cannot tell you yet
- Precision is unavailable Locality decided every match it could and no judge is qualified to decide the rest. That pass is paid and not approved
- Token totals are not comparable across model stacks Model tokenizers differ, and one model stack may run more than one role. Cost is provider-reported and does compare
- Cross-subject scores are not comparable Patch size and fault density differ across the 1 pull requests, which is why every suite figure here counts each one once
- 20 defects is the whole measurement Written down before any review ran, and independently confirmed. A model stack cannot be credited for anything outside them