AI Integration & Development

xAI Grok API Pricing (2026): Every Model, Real Costs, and How to Estimate Your Bill

Grok 4.5 just reset xAI's pricing. Every model's real cost, the tool fees nobody mentions, and a formula to estimate your monthly bill.

xAI shipped Grok 4.5 on July 8, and it reset the top of the lineup. So if you priced out Grok for a project even a month ago, your numbers are stale. Here is the whole pricing picture as it stands today, plus the part the pricing page will never give you: a way to turn these per-token rates into an actual monthly bill before you commit.

https://grizzlypeaksoftware.com/articles/api/image/416

USB-C for AI Is Already Here. Are You Building With It?

USB-C for AI Is Already Here. Are You Building With It?

MCP is how Claude, Cursor, and VS Code connect AI to the real world. This is the complete technical guide.

Build Your First MCP Server — $3.99

I am going to be straight about where Grok is a genuine bargain and where the sticker price hides the real cost, because those are two different conversations.

The current model lineup and what each costs

Every price below is per million tokens. Grok bills input and output separately, and as with every provider, output is the expensive side.

Model Input Cached input Output Context What it is for
Grok 4.5 $2.00 $0.50 $6.00 500K Newest flagship. Coding and agentic work.
Grok 4.20 $2.00 n/a $6.00 2M Prior flagship, largest context.
Grok 4.3 $1.25 n/a $2.50 1M The value flagship. Most work fits here.
Grok Build 0.1 $1.00 n/a $2.00 256K App-building and scaffolding tier.
Grok Code Fast 1 $0.20 $0.02 $1.50 256K Fast code completion. Cheap in, pricey out.
Grok 4.1 Fast $0.20 n/a $0.50 2M High-volume, cost-sensitive work.
Grok 3 Mini $0.30 n/a $0.50 128K The floor. Cheap classification and routing.

Two things on this table that will bite you if you skim:

The flagship context surcharge is real. On the flagship tiers, prompts above 200K tokens cost more. Grok 4.3, for example, roughly doubles to $2.50 input and $5.00 output once you cross that line. If your workload lives in long-context territory, the headline rate is not your rate.

Grok 4.5 costs more than the model it sits above. This is the counterintuitive part. Grok 4.3 is $1.25 in and $2.50 out. Grok 4.5 is $2.00 in and $6.00 out. The newest model more than doubled the output price of the value flagship. That is not a mistake on xAI's part, it is positioning: 4.5 is the coding and agent model, 4.3 is the one you run for everything else. Reaching for the newest model by reflex is how you overpay.

The second meter: server-side tool fees

Tokens are only half the bill the moment your app starts using tools. These are billed separately, per call, on top of tokens:

  • Web Search, X Search, Code Execution: $5.00 per 1,000 calls
  • File Attachments: $10.00 per 1,000 calls
  • Collections Search: $2.50 per 1,000 calls

Half a cent per web search sounds like nothing until an agent loop fires eight of them to answer one question, and every step also re-sends the growing context on the token meter. If you are building agents, model the tool calls explicitly. They are not a rounding error.

Prompt caching: the discount most people leave on the table

Grok 4.5 reads cached input at $0.50 against a $2.00 base. That is a 75% discount on any part of your prompt that repeats. If you send the same long system prompt or the same reference document on every call and pay full freight each time, you are lighting money on fire.

The catch is that caching only pays off when your stable content sits at the front of the prompt and gets reused before the entry expires. Put the variable part first and you break the cache. This is a structural decision about how you build your prompts, and it is worth more than any model you pick.

How to estimate your actual bill

Here is the part a pricing page cannot do for you. Per-token rates are useless until you multiply them by your real traffic shape. The formula is simple:

monthly cost  =  (avg input tokens  x  input rate)
              +  (avg output tokens x  output rate)
              ,  per request, times requests per month
              +  tool calls per month
              x  your retry factor

Let me make it concrete with a clearly hypothetical workload: an app doing 1 million requests a month, each sending 2,000 input tokens and getting back 500 output tokens. No tools, no caching yet.

On Grok 4.5:

  • Input: 2,000 tokens x 1M requests = 2,000M tokens x $2.00 = $4,000
  • Output: 500 tokens x 1M requests = 500M tokens x $6.00 = $3,000
  • Base: $7,000 / month

Now the same workload on the value flagship, Grok 4.3:

  • Input: 2,000M x $1.25 = $2,500
  • Output: 500M x $2.50 = $1,250
  • Base: $3,750 / month

And on Grok 4.1 Fast:

  • Input: 2,000M x $0.20 = $400
  • Output: 500M x $0.50 = $250
  • Base: $650 / month

Same traffic, and the bill swings by more than 10x depending on which model you point it at. That spread is the entire budgeting decision. If the job does not need the coding and agent muscle of 4.5, running it there is a five-figure annual tax for nothing.

Now add caching to the 4.5 case. Say 1,500 of those 2,000 input tokens are a stable prefix you can cache:

  • Cached input: 1,500M x $0.50 = $750
  • Fresh input: 500M x $2.00 = $1,000
  • Output: unchanged at $3,000
  • New total: $4,750 / month, down from $7,000

Caching alone cut the flagship bill by a third, and I did not change a single thing about the model or the quality. This is why I care more about how you structure a workload than which logo is on the model.

One last multiplier that never appears in anyone's estimate: your retry rate. Timeouts, truncated responses, and validation failures all bill in full and then bill again for the retry. If one call in ten fails and re-runs, your real rate is 10% higher than every number above. I wrote a whole piece on that hidden math in the hidden cost of LLM APIs, because it is the part that turns a clean estimate into a surprising invoice.

Is Grok worth it in 2026?

On raw price, Grok is aggressive and has been for a while. Grok 4.5 is the interesting case, because xAI is not selling it as the cheap option anymore. They are pitching it as an Opus-class coding and agent model that undercuts the top tier from Anthropic and OpenAI by a wide margin on price while landing near them on capability. Elon's framing was "comparable to Opus, but much faster and lower cost."

Treat that framing the way you would treat any vendor benchmarking its own model: as a claim to verify, not a fact to plan around. The honest version is that Grok 4.5 gives you a credible flagship for coding and agents at a fraction of flagship pricing, and Grok 4.3 gives you a genuinely cheap workhorse for everything that is not that. The strategy that actually saves money is not picking one model. It is routing: send the hard agentic work to 4.5, the bulk work to 4.3, and the trivial classification to Mini, and cache the stable parts of every prompt.

Cheap only helps you if the model does the job. But when the cheap model does do the job, and Grok's often does, the savings are not marginal. They are the difference between a hobby project you can afford to run and one you cannot.

Price your own traffic through the formula above before you commit to anything. The sticker rate is where the conversation starts, not where your bill lands.

Prices current as of July 23, 2026. xAI adjusts these more often than they announce, so re-check the rates before you build a budget on them.