Skip to content
FREE EVALUATION PLANNING TOOL

AI Evaluation Sample Size Calculator

A demo can make an AI workflow look ready. A small, unrepresentative test set can do the same. Use this transparent estimate to plan enough independently labeled cases before deciding whether a candidate actually improves a baseline.

No signupBrowser-only inputsPlanning estimate, not certification
TWO-VARIANT PLANNING ESTIMATE

Plan an AI evaluation before collecting examples

Enter the current pass rate and the smallest improvement worth detecting. The calculator estimates independently labeled cases for a baseline and a candidate variant.

The estimate uses a two-sided independent-proportions approximation at 95% confidence and 80% power. It does not send or store your inputs.

WHAT THE NUMBER MEANS

Cases are a budget for uncertainty

The calculator compares two independent pass rates. If a baseline passes 80% of cases and a candidate must reach 90% before it is worth adopting, it estimates the labeled examples needed in each group to detect that gap under its stated assumptions.

The output is not a target for collecting random prompts. A hundred near-duplicates tell a team far less than a smaller set that represents important task types, failure modes, languages, and high-impact edge cases.

PRACTICAL FLOW

Use the estimate without gaming it

  1. 1. Define pass before testing. Write a rubric and identify who labels ambiguous cases.
  2. 2. Set the smallest worthwhile gain. Ten percentage points may matter for a support workflow; it may be too small for a high-impact decision.
  3. 3. Build representative splits. Keep scenario families and repeated customer records from leaking across baseline and candidate sets.
  4. 4. Review failures, not only averages. A higher pass rate does not excuse a concentrated safety, privacy, or workflow failure.
SCOPE AND LIMITS

When this calculator is the wrong tool

This estimate assumes independent binary outcomes and a two-variant comparison. It is not appropriate for correlated conversations, multi-label quality rubrics, rare severe failures, sequential testing, or regulated validation without a statistician and domain review.

It also does not prove safety, fairness, legal compliance, or financial value. NIST's AI Risk Management Framework describes broader governance work across design, development, use, and evaluation. AQ Score is not affiliated with NIST, and this page is not a NIST assessment.

Make the pilot measurable after the test plan

Case count is only one evidence requirement. Before launch, assign an owner, capture a baseline, define a fallback, and decide how people affected by the workflow can flag harm or errors.