Free AI API cost forecasting tool
AI API & Token Cost Forecast Calculator
Forecast AI API spend from request volume, input and output tokens, cached-input pricing, retries, platform costs and future growth without hardcoding vendor prices.
Simple token-cost model
User-entered API ratesForecast one AI workload
Request volume
API pricing
Caching, retries & budget
Simple methodMonthly tokens = requests × tokens/request × retry/tool-call overhead. Cache hits move eligible input tokens from the normal input rate to the cached-input rate. Output tokens stay at the output rate. Fixed monthly platform cost is added after variable token spend.
Forecast table
Compounded annuallyAnnual API spend as volume and prices change
| Year | Monthly requests | Monthly token cost | Fixed monthly cost | Monthly total | Annual total | Budget status |
|---|
Advanced token portfolio
Multi-model cost mixForecast six AI workloads or model routes
Portfolio assumptions
AI workloads / model routes
Workload / modelMonthly requestsInput tok / reqOutput tok / reqInput $ / 1MCached $ / 1MOutput $ / 1MCacheableCache hit
Advanced methodEach workload calculates uncached input, cached input and output token cost separately. Global overhead increases all token volumes for retries, agent loops or tool calls. Forecast years compound request growth and API-price changes independently. Budget contingency is for planning and is not treated as actual provider spend.
Workload cost table
Current volume and pricesMonthly token economics by workload
| Workload | Requests | Input tokens | Cached tokens | Output tokens | Input cost | Output cost | Variable cost | Cost / request |
|---|
Multi-year forecast
Contingency shown separatelyPortfolio budget as traffic and API prices change
| Year | Monthly requests | Variable cost | Monthly total | Planning cost | Annual total | Budget status |
|---|
Traffic sensitivity
Cost as request volume changes
| Traffic scale | Monthly requests | Variable token cost | Monthly total | Budget headroom |
|---|
Output-token sensitivity
Cost impact of longer responses
| Output-token scale | Output tokens / month | Output cost | Monthly total | Change vs. baseline |
|---|
Cache-hit sensitivity
Value of improving prompt/context caching
| Cache hit rate | Cached input tokens | Cache savings | Monthly total | Change vs. baseline |
|---|
Scope
Use current provider pricing for the rates you enter
This calculator intentionally does not hardcode AI model prices because API rates, caching discounts and billing rules change. Confirm whether reasoning tokens, tool calls, images, audio, embeddings, storage, batch discounts or committed-use discounts are billed separately. Token estimates are planning inputs, not guaranteed provider invoices.
