Berk Bayri

Token

The small chunk of text, roughly a word or part of a word, that language models read and write. Providers usually price and limit usage in tokens, but token cost is not the cost of a result.

A token is the unit a language model works with. Text is split into tokens, which can be whole short words, parts of longer words or punctuation, and the model reads and generates them one at a time. A rough rule of thumb in English is that a token is a bit less than a word.

Providers usually meter usage in tokens, both for input (the prompt and context) and output (the response), and they set limits on how many tokens can be handled at once. That is why "usage" in AI is often discussed as a token bill.

A misleading denominator

Token cost is easy to measure and easy to optimize, but it is not the cost of a useful result. Cheaper tokens can still produce a more expensive workflow if the output needs more retries, review or recovery. The better denominator is the cost per accepted outcome, and the broader frame is total cost of ownership.

Tokens are consumed during inference by an LLM.

Read more in A cheaper AI model can move the cost instead of removing it.