Skip to main content

glossary terms

Token

Category
LLMs & Generative AI
Difficulty
Beginner

Definition

A token is the fundamental unit of text that a Large Language Model (LLM) processes, representing a sequence of characters, a word, or a sub-word fragment.

How It Works and Context

In the architecture of modern Large Language Models, text is not processed as raw strings of characters. Instead, it undergoes a process called tokenization, where the input is converted into a sequence of numerical identifiers known as tokens. A single token can represent a common word, a part of a word, or even a single character, depending on the specific tokenizer used by the model. For example, the word 'unhappiness' might be split into 'un', 'happi', and 'ness'. This sub-word approach allows models to handle rare words, misspellings, and complex morphological structures effectively. The total number of tokens a model can process at once is defined by its 'context window,' which acts as the model's short-term memory. Understanding tokenization is critical because it directly influences model performance, computational costs, and the accuracy of output generation.

Why It Matters

Tokens are the primary metric for both cost and capacity in generative AI. API pricing is almost universally based on the number of tokens processed, and the size of a model's context window determines how much information it can 'remember' during a conversation. Developers must understand token limits to prevent data truncation and to optimize the efficiency of their prompts and applications.

Real-world Example

When you send a prompt to an AI like ChatGPT, the system first tokenizes your input. If you ask, 'How many tokens are in this sentence?', the model breaks that sentence into roughly 10-12 tokens. If you were building an application using an API, you would be billed based on the total count of these tokens for both your input and the model's generated response, making token management essential for budget control.

Common Mistakes

  • Assuming one token always equals one word; in reality, many words are split into multiple tokens.
  • Ignoring the token limit of a model, which leads to the AI 'forgetting' the beginning of a long conversation.
  • Underestimating the cost of long prompts, as input tokens are billed just like output tokens.
  • Believing that tokenization is identical across all AI models; different models use different vocabularies and tokenization strategies.

Frequently Asked Questions

Why do models use sub-word tokens instead of whole words?

Using sub-word tokens allows the model to handle an infinite vocabulary. It can construct the meaning of rare or new words by combining known sub-word fragments, whereas a word-based system would fail if it encountered a word not in its fixed dictionary.

Can I calculate the exact number of tokens in my text?

Yes, most model providers offer 'tokenizer' tools or libraries that allow you to input text and see exactly how it will be broken down into tokens by that specific model's vocabulary.

Does the number of tokens affect the speed of the AI response?

Yes, because LLMs generate text one token at a time. A longer response requires more sequential generation steps, which increases the total time taken to complete the output.