QamoosTech
AI & DataBeginner

Tokenization

PronunciationTOH-kuh-nih-ZAY-shun

Definition

Tokenization is the process of breaking down raw text into smaller units called tokens, such as words, sub-words, or characters. These tokens serve as the fundamental building blocks that large language models use to process and understand input data.

Where you hear it

In discussions about LLM input limits, data preprocessing pipelines, and model training configurations.

Examples

  • The model failed because the input text exceeded the maximum tokenization limit.
  • Our preprocessing script handles tokenization before sending the data to the API.

Common mistake

Assuming that one token always equals one word, whereas many modern tokenizers break words into smaller sub-word units to handle complex vocabulary.