AI Inference Cost Calculator
A model can look cheap until retrieval context, output length, and fixed platform costs meet real request volume. Enter your own current rate card and observed token profile to estimate a cost per request and a budget worth reviewing.
What each request costs
Use actual token telemetry and the current provider price card. A low per-request cost can still become a large line item at volume.
- Token cost per request
- $0.0040
- Total cost per request
- $0.0240
- Monthly token cost
- $40.00
- Monthly total cost
- $240.00
- Annualized total
- $2,880.00
Input-token cost plus output-token cost
Includes fixed monthly cost allocated across requests
Usage-only cost before fixed charges
Token cost plus fixed monthly cost
Monthly total multiplied by 12; not a price guarantee
Copies the displayed assumptions and estimates. Answers stay in your browser.
Privacy: calculations happen entirely in your browser. The calculator does not send or save values or calculated costs. Analytics may count only a generic calculation event, without parameters.
Separate a request price from a platform budget
The calculator multiplies input tokens by the input price per million, then does the same for output tokens. It adds those two values for token cost per request, multiplies by monthly requests, and finally adds a fixed monthly cost.
That separation matters. A model change can reduce token cost while a monitoring minimum, vector database, or workflow platform keeps the real monthly bill from moving much.
Measure a representative week before trusting the estimate
- 1. Export actual usage. Start with provider logs, not a hand-picked demo prompt.
- 2. Split traffic by job. Retrieval-heavy, long-form, and tool-using requests often have different token shapes.
- 3. Use the right price tier. Confirm standard, batch, cached-input, long-context, and regional pricing.
- 4. Recalculate after a prompt change. More context or hidden reasoning can change output costs before users notice.
What this estimate deliberately leaves out
This is not a vendor quote. It does not automatically fetch model prices and cannot infer cached-token discounts, batch pricing, multimodal tokens, web-search calls, storage, egress, taxes, committed-use discounts, or failed/retried requests.
Check the current OpenAI pricing documentation and Gemini API pricing documentation before using any rate. Provider rate cards and billing rules change.
Cost is only one launch gate
A cheaper request is not a better workflow if it fails task quality, safety, or ownership checks. Pair the budget with an evaluation plan and a bounded pilot.
