Skip to content
FREE EVALUATION QUALITY TOOL

AI Evaluation Label Agreement Calculator

A pass rate is only useful when people apply its rubric consistently. Check raw agreement and Cohen's kappa before using two reviewers' binary labels to compare an AI baseline with a candidate.

No signupBrowser-only inputsRubric check, not certification
BINARY LABEL QUALITY CHECK

Measure whether two reviewers read the rubric alike

Enter counts from the same labeled cases. The tool reports raw agreement and Cohen's kappa, a measure that accounts for agreement expected from each reviewer's pass/fail tendency.

Use one shared, frozen rubric and independently assigned labels. This calculator does not send or store your inputs.

WHY CHECK THIS FIRST

A shared spreadsheet is not a shared standard

Two reviewers can agree often simply because almost every case passes. Cohen's kappa subtracts the agreement expected from each person's own pass/fail pattern, making disagreement visible when a simple percentage can be flattering.

Use this check after reviewers label the same calibration set independently. It is not a score for model quality; it is evidence about whether the measurement process is stable enough to support a model comparison.

PRACTICAL FLOW

Calibrate before scaling the evaluation

  1. 1. Freeze the rubric. Define pass and fail examples before either reviewer starts.
  2. 2. Double-label a representative set. Include difficult task families, not only easy cases.
  3. 3. Inspect disagreement cells. Re-read the rubric and discuss examples rather than just chasing a higher number.
  4. 4. Recalibrate, then measure candidate performance. Keep calibration cases separate from the final holdout set.
LIMITS

When a single kappa is not enough

This calculator supports exactly two reviewers and a binary pass/fail label. It does not handle multi-class categories, ranked judgments, repeated conversation turns, or rare severe harms well enough to replace specialist analysis.

Very imbalanced labels can also make kappa hard to interpret. Always inspect the four raw counts, sample composition, and concrete disagreements. A strong agreement score does not prove that the rubric measures the right business or safety outcome.

Turn a stable rubric into a measurable pilot

After calibration, estimate the independent cases needed to detect a meaningful pass-rate improvement, then check whether the wider workflow has an owner, fallback, and success metric.