AI & DataBeginner
Tokenization
PronunciationTOH-kuh-nih-ZAY-shun
Definition
Tokenization is the process of breaking down raw text into smaller units called tokens, such as words, sub-words, or characters. These tokens serve as the fundamental building blocks that large language models use to process and understand input data.
Where you hear it
In discussions about LLM input limits, data preprocessing pipelines, and model training configurations.
Examples
The model failed because the input text exceeded the maximum tokenization limit.
Our preprocessing script handles tokenization before sending the data to the API.
Common mistake
Assuming that one token always equals one word, whereas many modern tokenizers break words into smaller sub-word units to handle complex vocabulary.