A token is a basic unit of text that artificial intelligence (AI) models use to process and understand language. Unlike humans, who read entire words at once, AI systems break down text into smaller, manageable parts called tokens. These can be whole words, parts of words, or even punctuation marks. This process of breaking down text is known as tokenization. For example, the word "unhappy" might be split into "un" and "happy" as separate tokens, or it might remain as one token depending on the model's training. To understand the scale of tokens, consider that in English, one token is roughly four characters or about three-quarters of a word. So, 100 words in English would be approximately 130 tokens. However, this is an average and can vary. Common words like "the" fit into one token, while less common words might take more. OpenAI, a major AI research company, provides a free online tool called the Tokenizer that allows users to input text and see exactly how many tokens it contains. This tokenization process can be less efficient in some languages compared to English. For example, the French word "incroyable" (which means "amazing") is often split into two tokens, whereas the English word "amazing" is a single token. As a result, expressing the same idea in French can require more tokens than in English. This pattern is not unique to French. Studies show that languages like German and Italian require about 50% more tokens than English, while languages that use non-Latin alphabets, such as Thai or Arabic, may need up to fifteen times more tokens. This has implications for the cost of using AI models, as token count directly affects pricing. For instance, the greeting "Hello, how are you?" in English uses six tokens, while the same greeting in Thai uses 24 tokens. This means that a message in Thai can be four times more expensive in terms of tokens than the same message in English. While French is not the worst off, it still faces higher token costs compared to English. This has led to the development of AI models specifically trained on French and other European languages, such as Mistral, which aim to improve tokenization for these languages. The concept of a "context window" is another important aspect of AI models. This refers to the amount of text an AI can process at one time, which is measured in tokens. For example, Google’s Gemini 3.5 Flash model has a context window of one million tokens, which is equivalent to around 700,000 words in French. This allows the AI to remember previous messages in a conversation and provide more coherent responses. The size of context windows has grown significantly over time, from around 4,000 tokens in early versions of ChatGPT to millions in current models. In paid AI services, users are billed based on the number of tokens, not words or characters. Input tokens (what you send to the AI) and output tokens (what the AI responds with) are both counted, with output tokens often being more expensive. For instance, using the Claude Fable 5 model, a million input tokens cost around $10, while a million output tokens cost about $50. This means that reducing the number of tokens in your prompts can lower costs and free up more space in the context window. Tips for reducing token usage include being concise, avoiding unnecessary details, and using bullet points to structure instructions clearly.