AI Evaluation Sample Size Calculator
A demo can make an AI workflow look ready. A small, unrepresentative test set can do the same. Use this transparent estimate to plan enough independently labeled cases before deciding whether a candidate actually improves a baseline.
Plan an AI evaluation before collecting examples
Enter the current pass rate and the smallest improvement worth detecting. The calculator estimates independently labeled cases for a baseline and a candidate variant.
The estimate uses a two-sided independent-proportions approximation at 95% confidence and 80% power. It does not send or store your inputs.
Cases are a budget for uncertainty
The calculator compares two independent pass rates. If a baseline passes 80% of cases and a candidate must reach 90% before it is worth adopting, it estimates the labeled examples needed in each group to detect that gap under its stated assumptions.
The output is not a target for collecting random prompts. A hundred near-duplicates tell a team far less than a smaller set that represents important task types, failure modes, languages, and high-impact edge cases.
Use the estimate without gaming it
- 1. Define pass before testing. Write a rubric and identify who labels ambiguous cases.
- 2. Set the smallest worthwhile gain. Ten percentage points may matter for a support workflow; it may be too small for a high-impact decision.
- 3. Build representative splits. Keep scenario families and repeated customer records from leaking across baseline and candidate sets.
- 4. Review failures, not only averages. A higher pass rate does not excuse a concentrated safety, privacy, or workflow failure.
When this calculator is the wrong tool
This estimate assumes independent binary outcomes and a two-variant comparison. It is not appropriate for correlated conversations, multi-label quality rubrics, rare severe failures, sequential testing, or regulated validation without a statistician and domain review.
It also does not prove safety, fairness, legal compliance, or financial value. NIST's AI Risk Management Framework describes broader governance work across design, development, use, and evaluation. AQ Score is not affiliated with NIST, and this page is not a NIST assessment.
Make the pilot measurable after the test plan
Case count is only one evidence requirement. Before launch, assign an owner, capture a baseline, define a fallback, and decide how people affected by the workflow can flag harm or errors.
