AI Evaluation Label Agreement Calculator
A pass rate is only useful when people apply its rubric consistently. Check raw agreement and Cohen's kappa before using two reviewers' binary labels to compare an AI baseline with a candidate.
Measure whether two reviewers read the rubric alike
Enter counts from the same labeled cases. The tool reports raw agreement and Cohen's kappa, a measure that accounts for agreement expected from each reviewer's pass/fail tendency.
Use one shared, frozen rubric and independently assigned labels. This calculator does not send or store your inputs.
A shared spreadsheet is not a shared standard
Two reviewers can agree often simply because almost every case passes. Cohen's kappa subtracts the agreement expected from each person's own pass/fail pattern, making disagreement visible when a simple percentage can be flattering.
Use this check after reviewers label the same calibration set independently. It is not a score for model quality; it is evidence about whether the measurement process is stable enough to support a model comparison.
Calibrate before scaling the evaluation
- 1. Freeze the rubric. Define pass and fail examples before either reviewer starts.
- 2. Double-label a representative set. Include difficult task families, not only easy cases.
- 3. Inspect disagreement cells. Re-read the rubric and discuss examples rather than just chasing a higher number.
- 4. Recalibrate, then measure candidate performance. Keep calibration cases separate from the final holdout set.
When a single kappa is not enough
This calculator supports exactly two reviewers and a binary pass/fail label. It does not handle multi-class categories, ranked judgments, repeated conversation turns, or rare severe harms well enough to replace specialist analysis.
Very imbalanced labels can also make kappa hard to interpret. Always inspect the four raw counts, sample composition, and concrete disagreements. A strong agreement score does not prove that the rubric measures the right business or safety outcome.
Turn a stable rubric into a measurable pilot
After calibration, estimate the independent cases needed to detect a meaningful pass-rate improvement, then check whether the wider workflow has an owner, fallback, and success metric.
