Comparison builder
default
afi / afi · cli / proxy / Default profile / OpenAI · gpt 5.6 luna high / afi 0.30.0
P50 across 3 scored trials
3 recorded; malformed and unscored trials remain inspectable but do not participate in aggregate metrics
Recall40%
Model-judged precisionN/A
Total bill$0.0387
Duration3m 12s
Per-metric distributions
Exact Type-7 percentiles over complete scored samples; partial scored samples remain unavailable; non-scored trials are disclosed below
Recall 3/3 scored trials measured
P50 40% P95 44.5%
Model-judged precision 0/3 scored trials measured
N/A no qualified judge decided these findings, so only locality was measured
N/A · 0 measured
Total bill 3/3 scored trials measured
P50 $0.0387 P95 $0.0560
Duration 3/3 scored trials measured
P50 3m 12s P95 6m 0s
| Metric | Availability | Measured | Min | P50 | P95 | Max |
|---|---|---|---|---|---|---|
Findingsfinding_count | Measured | 3/3 | 8 | 8 | 8.9 | 9 |
Defects foundfound | Measured | 3/3 | 8 | 8 | 8.9 | 9 |
Defects missedmissed | Measured | 3/3 | 11 | 12 | 12 | 12 |
Unkeyed findingsunkeyed | Measured | 3/3 | 0 | 0 | 0 | 0 |
Intended findingsintended | Measured | 3/3 | 0 | 0 | 0 | 0 |
Recallrecall | Measured | 3/3 | 40% | 40% | 44.5% | 45% |
Model-judged precisionprecision | N/A no qualified judge decided these findings, so only locality was measured | 0/3 | - | - | - | - |
F1f1 | N/A no qualified judge decided these findings, so only locality was measured | 0/3 | - | - | - | - |
Tier 1 recalltier_1 | Measured | 3/3 | 33.3% | 66.7% | 96.7% | 100% |
Tier 2 recalltier_2 | Measured | 3/3 | 28.6% | 28.6% | 41.4% | 42.9% |
Tier 3 recalltier_3 | Measured | 3/3 | 33.3% | 33.3% | 48.3% | 50% |
Tier 4 recalltier_4 | Measured | 3/3 | 25% | 50% | 50% | 50% |
Security recallcategory_security | Measured | 3/3 | 0% | 0% | 0% | 0% |
Defect recallcategory_defect | Measured | 3/3 | 46.2% | 53.8% | 60.8% | 61.5% |
Maintainability recallcategory_maintainability | Measured | 3/3 | 0% | 0% | 0% | 0% |
Performance recallcategory_performance | Measured | 3/3 | 33.3% | 33.3% | 63.3% | 66.7% |
Median anchor distanceanchor_median | Measured | 3/3 | 0 | 0 | 0 | 0 |
Worst anchor distanceanchor_max | Measured | 3/3 | 0 | 0 | 2.6999999999999997 | 3 |
Refusals a judge overturnedanchor_missed | Not recorded no judging pass has looked beyond the scored slack | 0/3 | - | - | - | - |
Review billreview_bill | Measured | 3/3 | $0.0218 | $0.0387 | $0.0560 | $0.0580 |
Judge billjudge_bill | N/A no qualified judge decided these findings, so only locality was measured | 0/3 | - | - | - | - |
Total billtotal_bill | Measured | 3/3 | $0.0218 | $0.0387 | $0.0560 | $0.0580 |
Tokenstokens | Measured | 3/3 | 1099 | 1427 | 1673.6 | 1701 |
Durationseconds | Measured | 3/3 | 1m 43s | 3m 12s | 6m 0s | 6m 19s |
Carried findingscarried | Not recorded the reviewer reported no reuse figure | 0/3 | - | - | - | - |
Compatibility identity
Any change to these inputs creates another group
- Reviewer
- afi
- Reviewer tool
- afi · cli
- Subject
- proxy
- Reviewed SHA
9b51f95ef609a219e211e37b082cd2e6913190e0- Key fingerprint
bd331c0b144e- Settings fingerprint
8c8c79c46163- Configuration ID
config-657388b88e14616d- Build
- afi 0.30.0 /
b0f313c0b59a66ecc7612396dc8db0ea5da13a7a - Harness
- bench 1 /
631c4107752a97104719622800c9815e030613cfdirty - Adapter
- afi 1 /
sha256:a1a935298956020ee6d767a756847abb4c00694c88c2eb80d454886ebc4acf8c - Build ID
build-33906dcd701a0660- Cohort ID
cohort-fdab2adfb8a4bd39- Comparison ID
comparison-2f7717ed5785716d
Trials
Select View evidence to inspect one run; the static table is paginated at 100 trials
| Details | ||||||||
|---|---|---|---|---|---|---|---|---|
| afi afi · cli | default | proxy | 45% | N/A | 9 | scored | View evidence | |
| afi afi · cli | default | proxy | 40% | N/A | 8 | scored | View evidence | |
| afi afi · cli | default | proxy | 40% | N/A | 8 | scored | View evidence |