Contact, if you are interested in this website / domain name / Sponsorship / Advertisement / Partnership

How large language models work: tokens, training and sampling

Understand what happens inside a chat assistant, from tokens and training to sampling, context windows and why models hallucinate.

Large language models, usually shortened to LLMs, are the engines behind chat assistants such as ChatGPT, Claude and Gemini. They can draft a report, explain a tax form, fix a bug and translate a menu, yet at their core they do one surprisingly narrow thing: given some text, they predict what text should come next.

This guide walks through how that narrow skill becomes a useful assistant. You will see how text is broken into tokens, how models are trained in stages, how a single word gets chosen, what a context window is, why temperature changes the style of answers, why models sometimes state falsehoods with confidence, and what "reasoning" modes add.

Step one: turning text into tokens

Computers work with numbers, so the first job is converting your text into a sequence of numbers. A component called a tokenizer splits text into units called tokens and looks each one up in a fixed vocabulary, typically tens of thousands to a few hundred thousand entries long.

Tokens are not quite words. Common words are usually a single token. Rarer words are split into reusable pieces. Spaces and punctuation are often folded into tokens too. This lets the model handle any text, including typos, new product names and other languages, without needing an infinite vocabulary.

Worked example: tokenizing a sentence

Here is how a typical tokenizer might split a short sentence. Exact splits and ID numbers differ between models, so treat this as an illustration of the principle rather than the output of any specific system.

PositionToken textExample IDWhy it splits this way
1The464Very common word, one token
2 cat3797Common word; the leading space is part of the token
3 unbuck41528Rare word, so it is split into pieces
4led992A reusable ending seen in many words
5 the262Lowercase with a space is a different token from "The"
6 harness19356Reasonably common, one token
7.13Punctuation is its own token

So "The cat unbuckled the harness." becomes seven tokens, and the model actually sees a list of seven numbers. Two practical lessons follow. First, usage limits and pricing for AI services are usually counted in tokens, and a rough rule for English is that a token averages about three-quarters of a word. Second, because the model sees pieces rather than letters, tasks like counting the letters in a word or reversing a string can trip it up in ways that seem odd to a human.

How models are trained

An LLM is a very large neural network, a mathematical function with billions of adjustable numbers called parameters or weights. Training is the process of setting those numbers. It usually happens in stages.

1. Pretraining

The model is shown an enormous body of text: public web pages, books, reference works, code and other sources. Its task is simple: hide the next token, ask the model to predict it, measure how wrong it was, and nudge every parameter slightly to make it less wrong next time.

To predict the next token well across that much text, the model has to absorb grammar, facts, styles, the structure of arguments and code, and a great deal of everyday reasoning. The result is a base model: knowledgeable, but not yet an assistant. Ask it a question and it may continue with three more questions, because that is a pattern it has seen.

2. Fine-tuning

Next, developers train the base model on a much smaller, carefully prepared set of examples showing the desired behaviour: a user asks something, and a good assistant answers clearly, follows instructions and declines harmful requests. This supervised fine-tuning teaches the format and manner of a helpful assistant.

3. Learning from feedback

Finally, the model is refined using judgements about which responses are better. In the approach commonly called reinforcement learning from human feedback (RLHF), people compare pairs of answers and pick the better one. Those preferences train a scoring system, and the model is then adjusted to produce answers that score well. Variants use written principles or AI-assisted judgements alongside human ones, and some training rewards answers that can be checked automatically, such as code that passes tests or maths with a verifiable result. This stage shapes helpfulness, honesty, tone and safety.

Want this working in your business, not just on paper? Get a free, written AI starting plan.

Get my free AI plan

Next-token prediction in action

When you send a message, the model processes all the tokens in the conversation so far and outputs a score for every token in its vocabulary. Those scores are converted into probabilities that add up to 100 percent. One token is chosen, added to the text, and the whole process repeats for the next token. Your answer appears word by word on screen because it is literally being produced that way.

The heavy lifting happens in a design called the transformer. Its key mechanism, attention, lets each position in the text weigh how relevant every earlier token is. That is how the model connects "it" to the right noun three sentences back, or keeps track of which variable a line of code refers to.

Sampling, temperature and why answers vary

Once the model has probabilities for the next token, something has to pick one. Always choosing the single most likely token (called greedy decoding) tends to produce flat, repetitive text. So most systems sample: they pick randomly, weighted by the probabilities. A setting called temperature controls how sharply those probabilities are weighted.

Worked example: choosing the next word

Suppose the text so far is "Looking out the window, I could see the sky was" and the model's top four candidates are the ones below. These numbers are illustrative, calculated from simple made-up scores, but the effect is exactly how temperature works.

Next tokenLow temperature (0.5)Default (1.0)High temperature (1.5)
blue84%61%50%
grey11%22%25%
clear4%14%19%
fallingunder 1%3%7%

Read the table column by column. At low temperature the favourite dominates, so the output is predictable and focused, which suits data extraction, classification and factual summaries. At high temperature the less likely options get a real chance, so you see more variety and surprise, which suits brainstorming and fiction but also raises the odds of odd or wrong choices. If "falling" is picked, every following token must build on it, and the sentence may veer somewhere strange.

Two related settings often appear alongside temperature. Top-k sampling only considers the k most likely tokens. Top-p (or nucleus) sampling only considers the smallest set of tokens whose probabilities add up to p, such as 90 percent. Both cut off the long tail of unlikely tokens. Chat apps usually set these for you; developer platforms let you adjust them, which is one factor in choosing and configuring a model.

Context windows: the model's working memory

The context window is the maximum number of tokens the model can consider at once. It includes the system instructions set by the app, your messages, any documents you paste, the model's previous replies and the answer it is currently writing.

A model does not remember your conversation the way a person does. Each time it generates a reply, it rereads the whole conversation within the window. When a conversation grows beyond the limit, the app has to drop or summarise older material, and details from the start can quietly disappear.

Practical habits that follow from this:

  • Put the most important instructions near the start of a new chat, and restate them if a long conversation drifts.
  • Start a fresh conversation for a new task rather than piling everything into one thread.
  • When you paste long documents, ask about specific sections; even models with large windows can give less attention to material buried in the middle.

Why models hallucinate

A hallucination is a fluent, confident statement that is false. Once you understand next-token prediction, the cause is easier to see.

  1. The objective is plausibility. Training rewards text that resembles what tends to follow. A made-up citation with a realistic author, title and year looks very much like a real one.
  2. Knowledge is compressed, not stored. The model does not keep a copy of every document. Rare facts seen only a few times are stored weakly and can be blended with similar facts.
  3. There is no built-in fact check. Unless a tool such as web search or a document lookup is connected, nothing compares the output to an external source before it reaches you.
  4. Each token commits the next. An early wrong choice gets elaborated by everything that follows, producing a coherent but wrong answer.
  5. Training can reward guessing. If confident answers score better than "I am not sure", the model learns to guess. Developers increasingly train against this, but it has not disappeared.

The best defences are to ground the model in sources you supply, ask it to quote the passage supporting each claim, allow it to say it does not know, and verify any specific fact you will act on. Our guide to prompt engineering includes templates for these techniques.

What "reasoning" modes add

Many assistants now offer a reasoning or "thinking" mode. At a high level, the model is trained and allowed to generate a long stretch of intermediate working, breaking the problem into steps, trying approaches and checking them, before writing its final answer. That working may be hidden, summarised or shown to you, depending on the product.

This helps because each token gets a fixed amount of computation. Giving the model more tokens to think with effectively gives it more computation for hard problems, and training that rewards correct final answers teaches it to use that space productively. Reasoning modes tend to improve multi-step maths, logic, planning and complex coding.

The trade-offs are speed and cost: responses take longer and use more tokens. Reasoning also does not create knowledge the model lacks, so it can still reason carefully from a false premise. When a model can also call tools, search and act across multiple steps, you are entering the territory of AI agents.

Frequently asked questions

Does an LLM look up answers in a database?

Not by default. It generates answers from patterns stored in its parameters. Some products add search or document retrieval, which lets the model base its answer on fetched text, but the core model is not a database.

Why does the same prompt give different answers?

Because the next token is usually sampled from a probability distribution rather than always taking the top choice. Lower temperature, more specific instructions and examples of the desired output make responses more consistent.

Is a bigger context window always better?

It lets you include more material, which is valuable for long documents and codebases. But longer inputs cost more and take longer, and focused, relevant context usually beats dumping in everything you have.

Do models learn from my conversations?

The model does not change its parameters while chatting with you. Some providers may use conversations to train future models depending on your settings and plan, so check the provider's data controls before sharing sensitive information.

This guide is general information, not professional advice. Spotted an error? Tell us.