Skip to content
FREE AI COST PLANNING TOOL

AI Inference Cost Calculator

A model can look cheap until retrieval context, output length, and fixed platform costs meet real request volume. Enter your own current rate card and observed token profile to estimate a cost per request and a budget worth reviewing.

No signupBrowser-only inputsTransparent formula
YOUR ESTIMATE

Model a request before scaling it

BROWSER ONLY
PLANNING OUTPUT

What each request costs

Use actual token telemetry and the current provider price card. A low per-request cost can still become a large line item at volume.

Token cost per request
$0.0040

Input-token cost plus output-token cost

Total cost per request
$0.0240

Includes fixed monthly cost allocated across requests

Monthly token cost
$40.00

Usage-only cost before fixed charges

Monthly total cost
$240.00

Token cost plus fixed monthly cost

Annualized total
$2,880.00

Monthly total multiplied by 12; not a price guarantee

Copies the displayed assumptions and estimates. Answers stay in your browser.

Privacy: calculations happen entirely in your browser. The calculator does not send or save values or calculated costs. Analytics may count only a generic calculation event, without parameters.

FORMULA

Separate a request price from a platform budget

The calculator multiplies input tokens by the input price per million, then does the same for output tokens. It adds those two values for token cost per request, multiplies by monthly requests, and finally adds a fixed monthly cost.

That separation matters. A model change can reduce token cost while a monitoring minimum, vector database, or workflow platform keeps the real monthly bill from moving much.

USE A REAL TOKEN PROFILE

Measure a representative week before trusting the estimate

  1. 1. Export actual usage. Start with provider logs, not a hand-picked demo prompt.
  2. 2. Split traffic by job. Retrieval-heavy, long-form, and tool-using requests often have different token shapes.
  3. 3. Use the right price tier. Confirm standard, batch, cached-input, long-context, and regional pricing.
  4. 4. Recalculate after a prompt change. More context or hidden reasoning can change output costs before users notice.
LIMITS

What this estimate deliberately leaves out

This is not a vendor quote. It does not automatically fetch model prices and cannot infer cached-token discounts, batch pricing, multimodal tokens, web-search calls, storage, egress, taxes, committed-use discounts, or failed/retried requests.

Check the current OpenAI pricing documentation and Gemini API pricing documentation before using any rate. Provider rate cards and billing rules change.

Cost is only one launch gate

A cheaper request is not a better workflow if it fails task quality, safety, or ownership checks. Pair the budget with an evaluation plan and a bounded pilot.