Comparison builder
default
afi / afi · cli / proxy / Default profile / Z.AI · glm 4.7 flash high / afi 0.30.0
Observed values from 1 scored trial
1 recorded; malformed and unscored trials remain inspectable but do not participate in aggregate metrics
Recall0%
Model-judged precisionN/A
Total bill$0.0142
Duration15m 14s
Per-metric observations
Exact values from the one complete scored sample; unavailable metrics and non-scored trials remain explicit
Recall 1/1 scored trials measured
Observed 0% One complete scored trial
Model-judged precision 0/1 scored trials measured
N/A no qualified judge decided these findings, so only locality was measured
N/A · 0 measured
Total bill 1/1 scored trials measured
Observed $0.0142 One complete scored trial
Duration 1/1 scored trials measured
Observed 15m 14s One complete scored trial
| Metric | Availability | Measured | Min | P50 | P95 | Max |
|---|---|---|---|---|---|---|
Findingsfinding_count | Measured | 1/1 | 2 | 2 | 2 | 2 |
Defects foundfound | Measured | 1/1 | 0 | 0 | 0 | 0 |
Defects missedmissed | Measured | 1/1 | 20 | 20 | 20 | 20 |
Unkeyed findingsunkeyed | Measured | 1/1 | 2 | 2 | 2 | 2 |
Intended findingsintended | Measured | 1/1 | 0 | 0 | 0 | 0 |
Recallrecall | Measured | 1/1 | 0% | 0% | 0% | 0% |
Model-judged precisionprecision | N/A no qualified judge decided these findings, so only locality was measured | 0/1 | - | - | - | - |
F1f1 | N/A no qualified judge decided these findings, so only locality was measured | 0/1 | - | - | - | - |
Tier 1 recalltier_1 | Measured | 1/1 | 0% | 0% | 0% | 0% |
Tier 2 recalltier_2 | Measured | 1/1 | 0% | 0% | 0% | 0% |
Tier 3 recalltier_3 | Measured | 1/1 | 0% | 0% | 0% | 0% |
Tier 4 recalltier_4 | Measured | 1/1 | 0% | 0% | 0% | 0% |
Security recallcategory_security | Measured | 1/1 | 0% | 0% | 0% | 0% |
Defect recallcategory_defect | Measured | 1/1 | 0% | 0% | 0% | 0% |
Maintainability recallcategory_maintainability | Measured | 1/1 | 0% | 0% | 0% | 0% |
Performance recallcategory_performance | Measured | 1/1 | 0% | 0% | 0% | 0% |
Median anchor distanceanchor_median | N/A no finding matched a defect, so there is no anchor distance | 0/1 | - | - | - | - |
Worst anchor distanceanchor_max | N/A no finding matched a defect, so there is no anchor distance | 0/1 | - | - | - | - |
Refusals a judge overturnedanchor_missed | Not recorded no judging pass has looked beyond the scored slack | 0/1 | - | - | - | - |
Review billreview_bill | Measured | 1/1 | $0.0142 | $0.0142 | $0.0142 | $0.0142 |
Judge billjudge_bill | N/A no qualified judge decided these findings, so only locality was measured | 0/1 | - | - | - | - |
Total billtotal_bill | Measured | 1/1 | $0.0142 | $0.0142 | $0.0142 | $0.0142 |
Tokenstokens | Measured | 1/1 | 167260 | 167260 | 167260 | 167260 |
Durationseconds | Measured | 1/1 | 15m 14s | 15m 14s | 15m 14s | 15m 14s |
Carried findingscarried | Not recorded the reviewer reported no reuse figure | 0/1 | - | - | - | - |
Compatibility identity
Any change to these inputs creates another group
- Reviewer
- afi
- Reviewer tool
- afi · cli
- Subject
- proxy
- Reviewed SHA
9b51f95ef609a219e211e37b082cd2e6913190e0- Key fingerprint
bd331c0b144e- Settings fingerprint
91bb17f752bc- Configuration ID
config-5e7d6a6b490e6059- Build
- afi 0.30.0 /
b0f313c0b59a66ecc7612396dc8db0ea5da13a7a - Harness
- bench 1 /
10b3b2068100b7ba429855e062500eaf83b90120dirty - Adapter
- afi 1 /
sha256:a1a935298956020ee6d767a756847abb4c00694c88c2eb80d454886ebc4acf8c - Build ID
build-4ceb7cbd44e9619b- Cohort ID
cohort-3b7bd95c9526de2e- Comparison ID
comparison-cdc6b4c79792fb9b
Trials
Select View evidence to inspect one run; the static table is paginated at 100 trials
| Details | ||||||||
|---|---|---|---|---|---|---|---|---|
| afi afi · cli | default | proxy | 0% | N/A | 2 | scored | View evidence |