Benchmark

Model benchmark

23 models · 1 PRs · 20 keyed bugs

What one must-fix bug costs to find

Cheapest first · 4 of 20 planted bugs are must-fix
Z.AI · glm 5.3 flash 9 runs
$0.005997 $0.02 per must-fix bug found per review
Must-fix bugs found
86.1%
All bugs found
38.3%
Time per review
3m 48s
Extra findings
9.2% of findings
OpenAI · gpt 5.6 luna 3 runs
$0.0197 $0.04 per must-fix bug found per review
Must-fix bugs found
50%
All bugs found
41.7%
Time per review
3m 45s
Extra findings
0.0% of findings
Minimax · m3 3 runs
$0.0628 $0.06 per must-fix bug found per review
Must-fix bugs found
25%
All bugs found
15%
Time per review
3m 9s
Extra findings
10.0% of findings
OpenAI · gpt 5.6 sol 3 runs
$0.0639 $0.26 per must-fix bug found per review
Must-fix bugs found
100%
All bugs found
56.7%
Time per review
3m 14s
Extra findings
0.0% of findings

Showing the 4 cheapest per must-fix bug · 19 more models are plotted below and ranked in full on Compare

Every configuration priced

Colour is model · size is effort · filled markers are on the frontier

What a configuration costs against what it catches

What a review is priced in
Which bugs count

One marker per configuration, each priced over every pull request it was run on. Filled markers are on the efficient frontier: nothing in the set is both cheaper and better. 2 of 19 configurations sit on it, and 3 fell below the 5% floor and are not drawn

0%25%50%75%100%$0.02$0.05$0.10$0.20$0.50$1$2AnthropicxAIOpenAIMoonshot AIGoogleZ.AIMinimaxDeepSeek12Cost per review (USD, log scale)Must-fix bugs found
The efficient frontier for this measure, best first. Highlighted rows are efficient on cost and on time, and are the ones badged in the chart above
#ReviewerModelEffortPer reviewTimeMust-fixMatched
1 Z.AI · glm 5.3 flashZ.AI · glm 5.3 flashhigh$0.02073m 48s86%91%
2 OpenAI · gpt 5.6 solOpenAI · gpt 5.6 solhigh$0.25563m 14s100%100%

What kind of bug it finds

How serious each bug is comes from the answer key, not the reviewer

By severity

How much each bug matters

0%25%50%75%100%MUST-FIX4 bugsmust fix before mergeISSUE11 bugsshould fixSUGGESTION5 bugsoptional

By difficulty tier

How hard each one was to see

0%25%50%75%100%TIER 13TIER 27TIER 36TIER 44Tier 1 floor: every bug in the diff

3 reviewers are not drawn, reading under the 5% floor

15 of 23 do best on the bugs that block a merge and worst on the optional ones. Tier 1 is the floor, the 3 bugs visible in the changed lines, and 1 of 19 clear all of it, from 0% to 100%. From tier 3 to tier 4 recall moves 9 points, less than the one bug this key can express

What a review costs, in money and in time

Both panels start at 52% · bars are 95% intervals

Cost

Cheaper is left, better is up. Dashed rays are equal cost per must-fix bug, priced at the label

What is written beside each mark
0%25%50%75%100%$0.00$1.00$2.00$0.12$0.16$0.20$0.26glm 5.3 flashgpt 5.6 lunam3gpt 5.6 solgrok 4.5gpt 5.6 terrakimi k3gemini 3.1 pro previewglm 5.2kimi k2.7 codeopus 5v4 proglm 5.3gemini 3.7 flashfable 5sonnet 5Cost of one reviewbars are 95% intervals over 2 to 9 trials per pull requestMust-fix bugs found

Time

Faster is left

0%25%50%75%100%0.0s3m 20s6m 40s10m 0sglm 5.3 flashgpt 5.6 lunam3gpt 5.6 solgrok 4.5gpt 5.6 terrakimi k3gemini 3.1 pro previewglm 5.2kimi k2.7 codeopus 5v4 proglm 5.3gemini 3.7 flashfable 5sonnet 5Time for one reviewbars are 95% intervals over 2 to 9 trials per pull requestMust-fix bugs found

Hollow rings are the 1 pull requests behind each solid marker, each one against its own bug count. 3 configurations are not plotted, having found too little to price

On must-fix bugs, vertical bars across these 19 overlap in 92 of the 171 pairs. Price runs $0.01 to $2.06 a review and the wait runs 42s to 17m 51s, and on those price bars overlap in 54 of the 171 pairs while wait bars overlap in 116 of the 171 pairs

Bugs found across effort settings

Same model stack, same role stack, same pull request · only the effort knob moved
No reviewer ran two effort settings on the same pull request, so there is no move to show

Detail

Repeat the same review and you get a different answer

Same model stack, same pull request, same settings, run 2 to 9 times. The bars below span what those repeats actually returned, and the median cell moves 5% between its best and worst run while the widest moves 55%

proxy 20 defects
0%50%100%38.3%9 trials41.7%3 trials15%3 trials56.7%3 trials41.7%3 trials31.7%3 trials35%3 trials35%3 trials16.7%3 trials35%3 trials47.5%2 trials30%3 trials15%3 trials25%3 trials40%3 trials17.5%2 trials1.7%3 trials0%3 trials0%2 trials

Two model stacks on one pull request are comparable. Two pull requests are not: defect density differs across these 1

Trial by trial: the spread behind one number

Z.AI · glm 5.3 flash on proxy, 9 scored trials. Each dot is one review

Defects found 6 to 9 · of 20
Cost per review $0.0120 to $0.0317 · US dollars
Time per review 1m 42s to 7m 3s · wall clock
Recall by tier, per pull request

The same numbers per pull request, where the variation lives. Each model stack opens with its pooled row

Recall by discovery tier for each model stack and pull request
Model stack and pull requestTier 1Tier 2Tier 3Tier 4
Z.AI · glm 5.3 flash all 1 pull requests33.3%42.9%50%50%
feat(proxy): reclaim space from the blob store29.6%33.3%48.1%38.9%
OpenAI · gpt 5.6 luna all 1 pull requests66.7%28.6%33.3%50%
feat(proxy): reclaim space from the blob store66.7%33.3%38.9%41.7%
Minimax · m3 all 1 pull requests33.3%0%16.7%0%
feat(proxy): reclaim space from the blob store22.2%9.5%22.2%8.3%
OpenAI · gpt 5.6 sol all 1 pull requests100%42.9%50%50%
feat(proxy): reclaim space from the blob store100%47.6%50%50%
xAI · grok 4.5 all 1 pull requests66.7%28.6%50%25%
feat(proxy): reclaim space from the blob store66.7%33.3%50%25%
OpenAI · gpt 5.6 terra all 1 pull requests33.3%28.6%16.7%50%
feat(proxy): reclaim space from the blob store33.3%23.8%27.8%50%
Moonshot AI · kimi k3 all 1 pull requests33.3%28.6%50%25%
feat(proxy): reclaim space from the blob store44.4%28.6%44.4%25%
Google · gemini 3.1 pro preview all 1 pull requests33.3%42.9%33.3%25%
feat(proxy): reclaim space from the blob store33.3%42.9%27.8%33.3%
Z.AI · glm 5.2 all 1 pull requests0%14.3%33.3%0%
feat(proxy): reclaim space from the blob store0%19%27.8%8.3%
Moonshot AI · kimi k2.7 code all 1 pull requests33.3%28.6%33.3%25%
feat(proxy): reclaim space from the blob store44.4%28.6%38.9%33.3%
Anthropic · opus 5 all 1 pull requests83.3%28.6%50%50%
feat(proxy): reclaim space from the blob store83.3%28.6%50%50%
DeepSeek · v4 pro all 1 pull requests66.7%0%33.3%0%
feat(proxy): reclaim space from the blob store55.6%19%38.9%16.7%
Z.AI · glm 5.3 all 1 pull requests0%0%16.7%0%
feat(proxy): reclaim space from the blob store11.1%19%16.7%8.3%
Google · gemini 3.7 flash all 1 pull requests33.3%28.6%33.3%0%
feat(proxy): reclaim space from the blob store33.3%23.8%27.8%16.7%
Anthropic · fable 5 all 1 pull requests33.3%14.3%50%50%
feat(proxy): reclaim space from the blob store44.4%23.8%50%50%
Anthropic · sonnet 5 all 1 pull requests33.3%14.3%25%0%
feat(proxy): reclaim space from the blob store33.3%14.3%25%0%
DeepSeek · v4 flash all 1 pull requests0%0%0%0%
feat(proxy): reclaim space from the blob store0%4.8%0%0%
Meta · llama 4 maverick all 0 pull requestsN/AN/AN/AN/A
Mistral AI · devstral 2512 all 0 pull requestsN/AN/AN/AN/A
OpenAI · gpt oss 120b all 1 pull requests0%0%0%0%
feat(proxy): reclaim space from the blob store0%0%0%0%
Qwen · qwen3 coder 30b a3b instruct all 0 pull requestsN/AN/AN/AN/A
Qwen · qwen3.7 flash all 0 pull requestsN/AN/AN/AN/A
Z.AI · glm 4.7 flash all 1 pull requests0%0%0%0%
feat(proxy): reclaim space from the blob store0%0%0%0%
Weighted composite, if you want one number

A stated weighting, not a measured one: the answer keys rank severity and attach no number to it. Blocking 5 · issue 2 · suggestion 1, so the score is what a model stack earned over what the keys put on the table

Weighted composite score by model stack
Model stackCompositeCost per weighted point
Z.AI · glm 5.3 flash54.4%$0.000808
OpenAI · gpt 5.6 luna46.1%$0.001823
Minimax · m317.7%$0.007538
OpenAI · gpt 5.6 sol70.9%$0.007669
xAI · grok 4.549.6%$0.007958
OpenAI · gpt 5.6 terra41.8%$0.0122
Moonshot AI · kimi k341.1%$0.0167
Google · gemini 3.1 pro preview46.1%$0.0221
Z.AI · glm 5.220.6%$0.0192
Moonshot AI · kimi k2.7 code39.7%$0.0203
Anthropic · opus 563.8%$0.0258
DeepSeek · v4 pro27.7%$0.0103
Z.AI · glm 5.315.6%$0.0287
Google · gemini 3.7 flash29.8%$0.0367
Anthropic · fable 554.6%$0.0802
Anthropic · sonnet 518.1%$0.0855
DeepSeek · v4 flash1.4%$0.0212
OpenAI · gpt oss 120b0%N/A
Z.AI · glm 4.7 flash0%N/A
How much of what it reports is real

Precision needs a judge qualified to decide the pairs locality cannot, and that pass is not approved. Until it runs, the only noise signal is the share of findings that matched no keyed defect, which is not the same claim

Judged precision and unmatched share by model stack
Model stackJudged precisionReviews judgedUnmatched share
Z.AI · glm 5.3 flashN/A0 of 99.2%
OpenAI · gpt 5.6 lunaN/A0 of 30%
Minimax · m3N/A0 of 310%
OpenAI · gpt 5.6 solN/A0 of 30%
xAI · grok 4.5N/A0 of 30%
OpenAI · gpt 5.6 terraN/A0 of 30%
Moonshot AI · kimi k3N/A0 of 34.5%
Google · gemini 3.1 pro previewN/A0 of 34.5%
Z.AI · glm 5.2N/A0 of 30%
Moonshot AI · kimi k2.7 codeN/A0 of 315.4%
Anthropic · opus 5N/A0 of 20%
DeepSeek · v4 proN/A0 of 334.4%
Z.AI · glm 5.3N/A0 of 383.6%
Google · gemini 3.7 flashN/A0 of 30%
Anthropic · fable 5N/A0 of 30%
Anthropic · sonnet 5N/A0 of 20%
DeepSeek · v4 flashN/A0 of 375%
Meta · llama 4 maverickN/A0 of 0N/A
Mistral AI · devstral 2512N/A0 of 0N/A
OpenAI · gpt oss 120bN/A0 of 3100%
Qwen · qwen3 coder 30b a3b instructN/A0 of 0N/A
Qwen · qwen3.7 flashN/A0 of 0N/A
Z.AI · glm 4.7 flashN/A0 of 2100%
Every configuration on record
Every model stack, model and effort setting the registry holds
Model stackModelEffortRolesScored runs
Z.AI · glm 5.3 flashZ.AI · glm 5.3 flashhighagent9
OpenAI · gpt 5.6 lunaOpenAI · gpt 5.6 lunahighagent3
Minimax · m3Minimax · m3highagent3
OpenAI · gpt 5.6 solOpenAI · gpt 5.6 solhighagent3
xAI · grok 4.5xAI · grok 4.5highagent3
OpenAI · gpt 5.6 terraOpenAI · gpt 5.6 terrahighagent3
Moonshot AI · kimi k3Moonshot AI · kimi k3highagent3
Google · gemini 3.1 pro previewGoogle · gemini 3.1 pro previewhighagent3
Z.AI · glm 5.2Z.AI · glm 5.2highagent3
Moonshot AI · kimi k2.7 codeMoonshot AI · kimi k2.7 codehighagent3
Anthropic · opus 5Anthropic · opus 5highagent2
DeepSeek · v4 proDeepSeek · v4 prohighagent3
Z.AI · glm 5.3Z.AI · glm 5.3highagent3
Google · gemini 3.7 flashGoogle · gemini 3.7 flashhighagent3
Anthropic · fable 5Anthropic · fable 5highagent3
Anthropic · sonnet 5Anthropic · sonnet 5highagent2
DeepSeek · v4 flashDeepSeek · v4 flashhighagent3
Meta · llama 4 maverickMeta · llama 4 maverickhighagent0
Mistral AI · devstral 2512Mistral AI · devstral 2512highagent0
OpenAI · gpt oss 120bOpenAI · gpt oss 120bhighagent3
Qwen · qwen3 coder 30b a3b instructQwen · qwen3 coder 30b a3b instructhighagent0
Qwen · qwen3.7 flashQwen · qwen3.7 flashhighagent0
Z.AI · glm 4.7 flashZ.AI · glm 4.7 flashhighagent2

Provenance

What this cannot tell you yet
  • Precision is unavailable Locality decided every match it could and no judge is qualified to decide the rest. That pass is paid and not approved
  • Token totals are not comparable across model stacks Model tokenizers differ, and one model stack may run more than one role. Cost is provider-reported and does compare
  • Cross-subject scores are not comparable Patch size and fault density differ across the 1 pull requests, which is why every suite figure here counts each one once
  • 20 defects is the whole measurement Written down before any review ran, and independently confirmed. A model stack cannot be credited for anything outside them
benchee benchee-dashboard-1 built from 10f4ec58 Static benchmark evidence ·