Free tool

AI API cost calculator for your monthly bill.

Enter how many tokens you send and receive each month and see what it costs on each Claudech model, next to the published list prices of comparable Claude, GPT and Gemini models.
  • Input, output and cached tokens
  • Published list prices
  • No sign-up

Vendor prices, sources and check dates are listed on the price comparison page. Claudech rates are on pricing.

Method

How the estimate is calculated

Every provider prices input and output tokens separately, per million tokens. Cached input, the part of a prompt repeated from an earlier request, has its own lower rate. A month of usage therefore costs:

uncached input × input rate + cached input × cached rate + output × output rate

Each term uses millions of tokens and the rate per million. To find your own volumes, take the token counts from a typical day in your provider’s usage export and multiply by 30, or estimate from text length: a token is about four characters of English.

Worked example

A chat product sends 100 million input tokens a month, 40% of them from the cache, and receives 25 million output tokens. On Claude Opus 5.5 ($0.50 input, $0.010 cached input and $1.90 output per million):

  • 60M uncached input × $0.50 = $30.00
  • 40M cached input × $0.010 = $0.40
  • 25M output × $1.90 = $47.50

Total: $77.90 a month. The calculator above starts with these numbers; change them to match your workload.

FAQ

Frequently asked questions

How many tokens is a word?

For English text a token is roughly four characters, or about three quarters of a word, so 1,000 words is around 1,300 tokens. Code, other languages and unusual formatting use more tokens per word. The usage field of every API response gives the exact count.

Why is output priced higher than input?

Output tokens are generated one at a time, while input tokens are read in a single pass, so output costs the provider more to serve. Most providers, including Claudech, charge several times more per output token.

What is the cache hit rate?

When the start of a prompt repeats between requests, such as a long system prompt or a coding agent’s file context, the repeated part can be served from the prompt cache at a much lower rate. The hit rate is the share of your input tokens billed that way.

How accurate is the estimate?

It uses each provider’s published per-token list prices. It does not include batch discounts, long-context surcharges, enterprise contracts or taxes. Claudech figures are exactly what the API bills.

Pay only for the tokens you use

Buy any token package to start. Packages never expire, and every request shows exactly what it cost.