API versus Subscription: How You Actually Pay for AI (and Which to Choose)

Two ways to pay, two different products, and one popular mistake. Tokens, rate limits, caching, batch discounts, and a decision guide for students, builders and teams.

Sahi Padhai · 2026-09-30 · 6 min read

Every big AI company sells the same models twice. Once as a subscription to their app. Once as an API for your code. They look like two prices for the same thing. They are not, and people regularly pay for one while expecting the other.

Snapshot

Checked on 30 September 2026. Every number here is an example of the shape of pricing. Use the official links at the bottom for current prices.

Door one: the subscription

A subscription is a flat monthly fee for a product: the Claude app, ChatGPT, the Gemini app, Cursor.

  • You pay: a fixed amount per month (or a discounted amount per year).
  • You get: access to the company's app, plus an allowance of usage that resets on a schedule (for example every five hours, or monthly).
  • When you run out: you wait for the reset, or upgrade to a bigger plan. You do not get a surprise bill.
  • Who it is for: people who use AI: students, writers, analysts, and developers using the company's own coding agent.

Plans usually come in ladders. For example, Claude's ladder runs Free, Pro (around $20 a month), Max (from around $100, with 5 times or 20 times the usage), then Team and Enterprise seats. ChatGPT's and Google's ladders look similar, with names like Plus, Pro, Ultra. Each step up gives you more usage, more powerful models and earlier access to features, not a different kind of product.

Door two: the API

An API (application programming interface) gives programs direct access to a model. There is no app. You write code that sends a request and receives a response.

  • You pay: per token. A token is a small chunk of text, roughly three quarters of an English word. You are charged separately for input tokens (what you send) and output tokens (what comes back), usually quoted per million tokens.
  • You get: exact control. You pick the model, write the prompt, set limits and handle the answer.
  • When you run out: you do not run out. You pay for what you use, up to a spending limit you set.
  • Who it is for: people who build: a startup adding AI to a product, a script that classifies thousands of emails, a researcher running experiments.

Example of how different the scale is: Anthropic's API lists its small model Haiku 4.5 at about $1 per million input tokens and $5 per million output tokens, and its top Fable 5.1 at $10 and $50. A million tokens is roughly 750,000 words, so a short question and answer costs a tiny fraction of a cent on the small model.

The five differences that matter

SubscriptionAPI
Billingflat monthlymetered per token
What you accessthe company's app and its agentsthe raw model, from your code
Cost predictabilityhigh (fixed)depends on your usage
Limitsusage allowance that resetsrate limits and a spend cap
Customisationwhat the app offersanything you can code

And the one that surprises people:

They are separate accounts

A ChatGPT subscription does not include OpenAI API credits. A Claude subscription does not include API credits, and an API key does not unlock the consumer app. Two logins, two bills. Always check which one you are paying for.

How token billing works

Some API billing details that save real money:

  • Input is cheaper than output. Output tokens typically cost several times more, because generating text is more work than reading it.
  • Prompt caching. If you send the same long prefix again and again (a big system prompt, a document), the provider can cache it, and re-reading the cached part costs a small fraction of the normal input price. Anthropic's documentation puts cache reads at a tenth of the base input price or less.
  • Batch mode. If you do not need an instant answer, batch requests run later at a discount; Anthropic lists 50% off.
  • Tiers. Small models are many times cheaper than large ones. Using the largest model for a simple task is like hiring a surgeon to apply a plaster.
  • Free tiers. Google's Gemini API has a free tier for its Flash models, with lower rate limits and, importantly, your data may be used to improve the product. Paid tiers do not.
  • Thinking costs tokens. Reasoning models "think" before answering, and that thinking is billed as output. Higher effort settings mean better answers on hard tasks and a bigger bill.

Coding agents: where the two doors meet

Coding agents like Claude Code, OpenAI's Codex CLI and Gemini CLI let you choose how to pay:

AgentSign in with......or use
Claude Codea Claude subscriptionan API key from the Console
Codex CLIa ChatGPT plan (recommended)an OpenAI API key
Gemini CLIa Google account (free allowance)a Gemini API key or Vertex AI
Cursora Cursor plan with pooled model usageextra usage billed after the allowance

The trade-off is the same everywhere. Logging in with a subscription gives you a predictable monthly price and a usage window, which suits a developer working a normal day. An API key gives you no window and a bill that scales with the work, which suits automation, CI pipelines and heavy parallel agents that would exhaust any plan.

A simple decision guide

Choose a subscription if...

  • you mostly chat, write, study or research,
  • you use the company's own coding agent a few hours a day,
  • you want a fixed bill and do not want to think about tokens.

Choose the API if...

  • you are building software that calls a model,
  • you need to process large volumes automatically,
  • you need exact control over the model, version and cost,
  • you want a workflow that can run without you at the keyboard.

Use both if...

  • you code with an agent on a subscription and also ship an AI feature in a product.

Three traps to avoid

  1. Paying for the wrong door. Buying a subscription and expecting API credits (or the reverse).
  2. Using the flagship for everything. Start with the balanced tier, move up when the answers fall short, move down for high-volume simple tasks.
  3. No spending cap on an API key. Set a limit in the provider's dashboard, and never paste a key into public code. A leaked key is a bill.

Learn what a token really is

Tokens, context windows and what a model is actually doing when it predicts the next word are covered in our free Deep Learning course, and you can watch it happen in the simulator's LLM inference tab.

CompanySubscriptionAPI
Anthropic (Claude, Claude Code)claude.com/pricingplatform.claude.com/docs/en/about-claude/pricing
OpenAI (ChatGPT, Codex)chatgpt.com/pricingopenai.com/api/pricing
Google (Gemini, Antigravity, Gemini CLI)one.google.com/about/google-ai-plansai.google.dev/gemini-api/docs/pricing
Cursorcursor.com/pricing(plan with pooled usage)
Meta (Muse Spark)free in the Meta AI app at meta.aiprivate preview, no public pricing yet
Antigravityantigravity.google (free for individuals)via Gemini API