You type a sentence.
The AI service tells you it used 37 tokens.
Then you send a translation of the same idea in another language and the count changes.
Or you try another model and the count changes again.
So what exactly is a token?
The simplest answer is useful but incomplete:
A token is a model-specific unit used to represent and generate data. It is not the same thing as a word.
The more useful answer is that “token” appears in three different conversations at once:
- how information is represented,
- how usage is measured and billed,
- how work moves through an inference system.
Those three are connected.
They are not identical.
Quick Answer
For text models, a tokenizer breaks text into units from a fixed vocabulary—the set of token patterns the model knows how to represent.
A token can correspond to a whole word, part of a word, punctuation, whitespace, a character, or another learned unit depending on the tokenizer.
Google gives a rough Gemini rule of thumb of about four characters per token, or roughly 60–80 English words per 100 tokens. It explicitly treats that as an approximation.[1]
The important part is what comes next:
Token count is useful for context limits and billing, but one token is not a universal unit of meaning, compute, latency, or electricity.
After this article, a “1 million token context window” or “$2 per million input tokens” headline should be easier to decode: what is being counted, what it tells you, and what it does not.
Original Asset 1: A Token Has Three Different Roles
The word token becomes confusing because people use one number to discuss three layers.
1. Representation
The tokenizer converts text into a sequence of discrete model-readable units.
This is why:
human text → tokens → token IDs → model representations
The exact split depends on the tokenizer.
2. Billing and limits
Providers count tokens to define context windows and usage-based pricing.
They may distinguish:
- uncached input,
- cached input,
- reasoning or thinking usage,
- output,
- and modality- or tool-specific usage.
3. Compute
Tokens trigger computation and memory movement inside the model.
But one token on one model does not require the same physical work as one token on another model.
Different model sizes, architectures, reasoning settings, hardware, batching and context lengths can all change the real cost.
So:
Token as representation
≠
Token as billing meter
≠
Fixed unit of physical compute
Why the Same Sentence Can Produce Different Token Counts
There is no universal tokenizer shared by every model.
Each tokenizer has a vocabulary built from patterns it was designed or trained to represent efficiently.
A common word or character sequence may be one token in one vocabulary and several tokens in another.
This gives us a useful term:
Token fertility is the number of tokens needed to represent a unit of human text.
You can measure fertility as:
- tokens per word,
- tokens per character,
- or tokens for the same translated sentence or task.
Lower fertility means the text is represented with fewer tokens.
The Language Question: The Same Meaning Can Occupy Different Amounts of Context
This is where tokenization stops being a small technical detail.
Academic work has repeatedly found that the same information can require substantially different token counts across languages.
Ahia and colleagues compared commercial tokenization across 22 languages and found that API token counts and therefore token-based costs were not uniform across languages.[2]
A 2026 AI & Society paper reviews the continuing issue and notes that many non-Latin scripts can be fragmented less efficiently by tokenizers whose training distributions and vocabularies favor other patterns.[3]
A separate 2026 controlled study found about a 3.4× difference in token fertility between two script forms in the specific mBERT and XLM-R setups it tested, with a large inference-speed difference on the same hardware.[4]
That last number should not be turned into a rule such as “Korean costs 3.4× English.”
The correct lesson is broader:
Same meaning does not guarantee the same sequence length.
For a multilingual application, this can change:
- how much text fits in the context window,
- how many billable tokens are used,
- how much input the model must process,
- and sometimes latency and quality.
Original Asset 2: The Language Cost Chain
A useful causal chain is:
same meaning
→ different tokenization
→ different sequence length
→ different context occupancy / billable tokens
→ potentially different latency and cost
The word potentially matters.
Token count is one driver, not the only driver.
Why $1 per Million Tokens Is Not Automatically an Apples-to-Apples Comparison
Suppose Provider A and Provider B both advertise the same price per million input tokens.
You still do not know whether the same job costs the same.
You need to ask:
- Do the models tokenize the same text into the same number of tokens?
- Do they need the same amount of reasoning?
- Do they produce the same output length?
- Do they need the same retries?
- Are long-context requests priced differently?
- Does one model succeed more often?
The token price is a rate.
The task determines the quantity.
Current Example: Context and Pricing Are Model Rules, Not Token Definitions
As of October 4, 2026, OpenAI lists GPT-6.1 Sol with a 1.05-million-token context window and a maximum output of 128,000 tokens.[5]
For standard text usage up to the listed long-context threshold, its published rates are:
- $2 per million input tokens,
- $0.10 per million cached input tokens,
- $2.50 per million cache-write tokens,
- $10 per million output tokens.
Prompts above 272,000 input tokens use higher rates for the full request.[5]
These numbers will change.
The reusable lesson is:
A token tells you the unit being counted. The model page tells you what that unit costs and how many the model can handle.
Input Speed and Output Speed Are Different Problems
“More tokens make AI slower” is directionally useful but too simple.
It helps to split inference into two stages.
Prefill
Prefill is the stage where the model processes the input context before it begins generating the answer.
A larger prompt means more input must be processed and more state must be prepared.
This can increase time to first token—the delay before visible output starts.
Decode
Decode is the stage where the model generates new output tokens.
Autoregressive language models generate output sequentially: the next token depends on the sequence already generated.
That is why a long answer can take time even after the input has already been processed.
So a better speed model is:
Input length → prefill work → time to first token
Output length → decode steps → generation time
Tokenization Time Is Not the Same as Token Processing Time
Another subtle distinction matters.
The tokenizer itself must convert characters into token IDs.
But the expensive AI work usually discussed in model inference is what happens after that conversion, when the neural network processes those token sequences.
Recent systems discussions about faster tokenizers show that tokenizer implementation can matter for very large inputs or specific workloads. But that does not make tokenizer CPU time equivalent to transformer inference cost.[6]
So when someone says “more tokens require more compute,” ask which part of the pipeline they mean.
Prompt Caching Changes the Economics of Repeated Tokens
Applications often reuse long instructions, tool definitions or conversation prefixes.
OpenAI's current prompt-caching system can preserve the model's reusable key-value state for matching prompt prefixes. A later request can reuse that state instead of recomputing the entire prefix at the uncached rate.[7]
Cached input therefore means repeated context processed under a caching path and billing rule.
It does not mean:
the token disappeared or became free.
It remains part of the model's context and a separate usage category.
Context Window: Think Capacity, Not Pages
A context window is the maximum tokenized working sequence a model can accept under its rules.
That means “a 1-million-token context window” is not the same as “one million words.”
How many pages, words or source documents fit depends on:
- the tokenizer,
- the language,
- formatting,
- code and structured data,
- tool definitions,
- and output space reserved by the model/API.
So the useful model is:
instructions + conversation + documents + tools + generated work ≤ context budget
Why Tokens Matter More in the Agent Era
A chatbot may process one prompt and generate one answer.
An agent can call the model repeatedly.
Each new step can send some of the previous context again.
That means the same information can contribute to input usage many times across one user task.
This is the deeper topic behind Why Cheap Tokens Still Make AI Agents Expensive to Run.
For this article, remember only the bridge:
A token is counted per model interaction, not per unique idea in the user's task.
Original Asset 3: The Token Interpretation Checklist
Whenever you see a token number in an AI article, product page or invoice, ask six questions.
- Which model and tokenizer? The same text may not produce the same count elsewhere.
- Which token category? Input, cached input, reasoning, output or a modality-specific category?
- Which language or modality? Text representation efficiency is not universal.
- Which context and pricing tier? Long-context rules can change the rate.
- What are we trying to measure? Context capacity, price, latency or physical compute?
- Is token cost the right business metric? A cheaper token does not guarantee a cheaper successful task.
This checklist prevents most common token misunderstandings.
What to Watch Next
- Tokenizer efficiency across languages. Newer models may reduce—but not automatically eliminate—cross-language disparities.
- Token-free or byte/patch architectures. Research continues on models that reduce reliance on fixed subword tokenization.
- Long-context economics. Context windows are growing faster than “cheap full-context processing” is becoming trivial.
- Cache-aware applications. Stable context may increasingly be designed around reusable model state.
- Agent accounting. As one task creates many calls, token accounting will move from message-level to workflow-level economics.
The Simple Idea to Remember
A token is small, but the concept sits at the intersection of language, software and infrastructure.
It tells the model how information is represented.
It gives providers a practical unit for context and billing.
And it influences how much work an inference system has to do.
But it is not all three things in exactly the same way.
A token is a useful meter for AI activity—not a universal physical unit of AI work.
Key Vocabulary
token
A tokenizer-dependent unit used to represent model input or output.
tokenizer
The component that converts raw input such as text into a sequence of model-readable tokens.
vocabulary
The set of token patterns a tokenizer can represent directly.
token fertility
How many tokens are required to represent a unit of human text, such as a word, character or translated sentence.
prefill
The stage where a model processes the input context before generating output.
decode
The stage where an autoregressive model generates output tokens sequentially.
cached input
Repeated prompt content processed through a provider's prompt-caching mechanism and billed under its applicable cached-input rules.
context window
The model's token-based limit for the information and generated work it can keep active under a request or interaction.
Related Articles
- Why Cheap Tokens Still Make AI Agents Expensive to Run
- GPU vs. HBM: Why AI Needs Both Compute and Fast Memory
- Training vs. Inference: Why AI Needs Different Infrastructure for Each
- What Is the AI Full Stack? From Power and Chips to Models and Robots
Sources
- Google AI for Developers — Understand and count tokens, checked October 4, 2026.
- Ahia et al. — Do All Languages Cost the Same? Tokenization in the Era of Commercial Language Models, EMNLP 2023.
- Caffoni — The cost of language: tokenization as a metric of labor, AI & Society, 2026.
- Dixit & Dixit — The Script Tax: Measuring Tokenization-Driven Efficiency and Latency Disparities in Multilingual Language Models, 2026.
- OpenAI API — GPT-6.1 Sol model, context and pricing, checked October 4, 2026.
- Hacker News discussion — GigaToken: faster language-model tokenization, 2026. Used as a systems-question signal, not a primary performance benchmark.
- OpenAI API — Prompt caching, checked October 4, 2026.
- Sennrich, Haddow & Birch — Neural Machine Translation of Rare Words with Subword Units, ACL 2016.
- Kudo & Richardson — SentencePiece, EMNLP 2018.
Update History
October 4, 2026 — Updated with current model context/pricing, multilingual tokenization evidence, and a new representation-billing-compute framework.
Sources checked through October 4, 2026. Tokenization and billing rules vary by model and provider. Cross-language ratios cited from research are model- and dataset-specific and should not be treated as fixed ratios for current proprietary models.