Token
The small chunk of text, roughly a word or part of a word, that language models read and write. Providers usually price and limit usage in tokens, but token cost is not the cost of a result.
A token is the unit a language model works with. Text is split into tokens, which can be whole short words, parts of longer words or punctuation, and the model reads and generates them one at a time. A rough rule of thumb in English is that a token is a bit less than a word.
Providers usually meter usage in tokens, both for input (the prompt and context) and output (the response), and they set limits on how many tokens can be handled at once. That is why "usage" in AI is often discussed as a token bill.
A misleading denominator
Token cost is easy to measure and easy to optimize, but it is not the cost of a useful result. Cheaper tokens can still produce a more expensive workflow if the output needs more retries, review or recovery. The better denominator is the cost per accepted outcome, and the broader frame is total cost of ownership.
Tokens are consumed during inference by an LLM.
Read more in A cheaper AI model can move the cost instead of removing it.
Related terms
LLM
Large language model: an AI model trained on very large amounts of text to predict and generate language, which can be used to answer, summarise, write, reason about and act on text.
Inference
Running a trained AI model to produce an output for a request, as opposed to training it. Inference cost is the model-side spend that is easy to see but only one part of total AI cost.
Cost per accepted outcome
The total cost of getting one AI-assisted result that is actually accepted and used, counting model spend, retries, human review, exceptions and recovery — a truer measure than cost per call.
Used in these essays
The wrong questions about AI right now
Many of the questions that helped us orient ourselves around generative AI are now too blunt to be useful. The harder work is no longer asking what AI is in the abstract, but specifying where it works, where it fails, what authority it should have, and what the whole system costs.
A cheaper AI model can move the cost instead of removing it
A lower model bill can hide a higher workflow bill. The useful AI TCO question is not only what got cheaper, but where the cost moved.
Keep the AI ideas you rejected
A new model release should not restart your AI roadmap. It should reopen only the ideas that were rejected for a constraint the release actually changed.