Tokenizer
The component that splits text into tokens (sub-word units) before feeding it to an AI model. Different models use different tokenizers — affecting how they count input length, handle multilingual text, and process code.
Why it matters
Models don't read letters or whole words; they read tokens, the small chunks a tokenizer breaks your text into. This quietly shapes cost, since you're usually billed per token, and it explains quirks like models miscounting characters or struggling with rare words. Understanding tokens helps you write leaner prompts and makes sense of why the same sentence can cost more in one model than another.
In practice
The word "unbelievable" might be split into pieces like "un," "believ," and "able," while a common word like "the" is a single token. Emojis or unusual names can take several tokens each. If you paste a long document and hit a length limit, it's the token count, not the word count, that matters, which is why trimming filler text can let more of your content fit.
Related terms
Artificial Intelligence (AI)
The simulation of human intelligence by computer systems, including learning, reasoning, and self-correction. Modern AI is primarily powered by machine learning and neural networks.
Context Window
The maximum amount of text (measured in tokens) an AI model can consider at once. Larger context windows allow the model to reference more information in a single conversation.
Large Language Model (LLM)
An AI model trained on vast amounts of text data that can generate, summarize, translate, and analyze human language. Examples include GPT-4, Claude, Gemini, and Llama.
Natural Language Processing (NLP)
The branch of AI focused on enabling computers to understand, interpret, and generate human language. Powers chatbots, translation, sentiment analysis, and text summarization.
Prompt
The text instruction you give to an AI model to generate a response. Prompt quality directly impacts output quality — better prompts yield dramatically better results.
Token
The basic unit of text that AI models process — roughly 3/4 of a word in English. 'Unbelievable' is 3 tokens. Token limits determine how much text a model can process at once.