THE COST OF USEFUL AI / LAB 06
What does the score hide?
Fictional teaching examples. Calculations stay in your browser. No provider calls or input uploads.
Fictional examples. Human reference labels are resolved judgements under the rubric, not infallible truth. Intervals assume independent representative Bernoulli observations from a stable population; they do not correct sampling bias, clustered cases or grading errors. The zero-event bound applies only when zero failures were observed.
Inspect the grader
Enter the four counts from a reference-labelled audit. Accept is the positive decision.
Keep reviewer disagreement
Agreement can look high while the disputed cases contain the most useful rubric feedback.
Open the average by slice
The sample has more routine work than risk-sensitive work. Production weights let you ask whether that mix matches the real queue.
Relative whole-number weight
Relative whole-number weight
Put a range beside a rate
This is a separate sample of task acceptance, not a confidence interval for the grader matrix above.
When no failures appeared
A third, separate sample with exactly zero failures of the specified kind.
YOUR GRADER AUDIT
- Overall agreement
- 85%
- All 200 audited outputs
- False-accept rate
- 40%
- Among 50 reference failures
- Release precision
- 87.5%
- Among 160 outputs the grader accepts
- False-reject rate
- 6.67%
- Among 150 reference passes
Reviewer calibration
84% agreement
Keep all 8 disagreements with both labels and the resolution reason.
Slice view
Routine
98%
Risk-sensitive
70%
Raw sample average: 92.4%
Production-weighted average: 95.2%
Weights are normalized automatically. They currently total 100. A weighted average still cannot excuse a failed mandatory slice.
Separate task sample
90% accepted
Two-sided 95% Wilson interval: 82.56% to 94.48%
The 95% level describes this method's long-run coverage under its assumptions. It is not a guarantee that 95% of future tasks will pass.
Separate zero-failure sample
0.99%
Exact one-sided 95% upper bound on the failure probability. Zero observed failures does not establish zero risk.
Try the printed exercise
The default grader agrees on 85% of cases while accepting 40% of reference failures. The slice view shows why 92.4% across the sample does not describe the 70% risk-sensitive result.
Interval methods: NIST/SEMATECH e-Handbook. Neither this lab nor a single score supplies a universal release threshold.