How Large Language Models Actually Work: Tokens, Embeddings and Attention Explained
Learn how large language models turn text into tokens, build numerical representations and use attention to generate answers, with worked examples and practical exercises.
Key takeaways
- Models process tokens rather than words, so tokenisation affects cost, context limits and some common mistakes.
- Embeddings represent tokens numerically; transformer layers turn them into context-sensitive representations.
- Attention mixes information between token positions, while other network components transform that information.
- Text generation repeatedly predicts and selects another token; fluent output is not a guarantee of truth.
- Better results come from relevant context, clear constraints, external tools and checks against reliable evidence.
On this page
- From a sentence to a stream of predictions
- Tokens: the pieces a model actually reads
- Embeddings: turning token labels into useful numbers
- Position: how the model distinguishes order
- Attention: choosing which information to combine
- What happens inside a transformer layer?
- Learning: how the model acquires its parameters
- Generation: how scores become an answer
- Context windows: temporary working material, not unlimited memory
- A complete worked example: answer from evidence
- What this mechanism explains—and what it does not
- Common mistakes and better habits
- FAQ
From a sentence to a stream of predictions
Type a question into an AI assistant and the response arrives as readable text. Behind that familiar interface, a large language model is processing numbers: identifying text fragments, transforming vectors, mixing information between positions and estimating what should come next.
Understanding that process is useful even if you never train a model. It explains why a chatbot can write a persuasive argument but invent a reference, why a long document may overwhelm it, and why a small change in your prompt can alter the answer.
The essential sequence is:
- Tokenisation: split text into pieces and map those pieces to numerical identifiers.
- Embedding: look up a learned numerical representation for each token.
- Contextual processing: pass those representations through layers that combine attention with other calculations.
- Prediction: calculate scores for possible next tokens.
- Generation: choose a token, add it to the sequence and repeat.
That sequence describes the core of many text-generating systems. The surrounding product may also search the web, retrieve documents, execute code or store preferences. Those additions matter, but they are not the same thing as the language model itself.
The kind of model we will examine
This article focuses on autoregressive, decoder-only transformers, the architecture behind many modern text assistants. “Autoregressive” means the model generates a sequence by conditioning each new token on earlier tokens.
Not every language model works this way. Some models are designed to encode text rather than generate it; others combine separate encoder and decoder components. Multimodal systems also process representations of images, audio or other inputs.
The transformer architecture was introduced in *Attention Is All You Need*. Modern implementations have evolved, but its central idea remains important: tokens can gather information from other positions through attention.
We will follow a small example throughout:
The parcel arrived on Tuesday. It was damaged.
When did the parcel arrive?
A useful answer is “Tuesday”. Producing it involves much more than searching for the nearest date. The model must represent the question, relate “arrive” to “arrived”, identify the relevant statement and generate an appropriate response.
Tokens: the pieces a model actually reads
A model does not normally receive a sentence as a list of complete words. It receives tokens, which may correspond to words, parts of words, punctuation, whitespace or other text fragments.
A possible, purely illustrative split is:
Text:
The parcel arrived on Tuesday.
Possible pieces:
["The", " parcel", " arrived", " on", " Tuesday", "."]
The actual split depends on the tokenizer. Another tokenizer might divide “Tuesday” into several pieces or handle spaces differently. Never treat a hand-written example as the exact output of a particular model.
Why use tokens rather than words?
A fixed vocabulary containing every possible word would be impractical. Language includes new names, technical terms, spelling mistakes, code, product identifiers and combinations that no vocabulary designer could anticipate.
Subword tokenisation offers a compromise. Common text sequences can receive their own tokens, while rarer sequences can be assembled from smaller pieces.
For example, a tokenizer might encode a familiar word in one token but split an unfamiliar surname into several. Many tokenizers can fall back to byte-level representations, allowing them to encode unusual text without needing a dedicated vocabulary entry for every character sequence.
Common approaches include byte-pair encoding, WordPiece and Unigram. Their details differ, but they share the goal of representing text using a manageable vocabulary. The Hugging Face tokenizer summary explains these approaches and their trade-offs.
The important distinction is that the tokenizer supplies a segmentation scheme, not an understanding of the sentence.
Token IDs are labels, not meanings
After splitting the input, the tokenizer converts each piece into an integer identifier.
Illustrative tokens:
["The", " parcel", " arrived"]
Invented token IDs:
[41, 9827, 615]
These numbers are arbitrary labels within that tokenizer’s vocabulary. Token 9827 is not more important than token 41, and nearby IDs do not necessarily represent similar concepts.
The model uses each ID to look up a vector. Meaning-related relationships appear in those learned vectors and subsequent computations, not in the numerical ordering of the IDs.
A tokenizer and a model must therefore agree. Passing IDs from an unrelated tokenizer into a model is rather like using page numbers from the wrong edition of a reference book.
Why tokenisation affects everyday use
Tokenisation has several practical consequences.
Length and cost: Many APIs measure input and output in tokens. A page count or word count is only an approximation of the computational input.
Language differences: The same idea can require different numbers of tokens in different languages. Vocabulary design and training data influence this, so English-based rules of thumb do not generalise reliably.
Character-level tasks: Counting letters, manipulating unusual strings and reversing text can be awkward because the model is not necessarily processing one character at a time.
Identifiers: A product code such as ZX-1847-Q may be split across several tokens. Exact reproduction requires preserving a sequence, not simply recalling one indivisible item.
These are contributing factors, not universal explanations. A model can sometimes perform character tasks correctly, and tokenisation alone does not explain every arithmetic or spelling error.
Exercise: inspect rather than guess
Use a tokenizer tool or library associated with the model you intend to use.
- Enter a short ordinary sentence.
- Replace one common word with a long invented name.
- Add an emoji and a product identifier.
- Compare the token counts and visible pieces.
- Repeat with the same meaning in another language, if you know one.
Record the result in a small table:
| Input type | What to inspect |
|---|---|
| Ordinary sentence | Which words are single tokens? |
| Invented name | Where does the tokenizer split it? |
| Product identifier | Are digits grouped or separated? |
| Emoji | Does one visible symbol require several tokens? |
| Another language | How does the token count change? |
Do not ask a chatbot to estimate its exact tokenisation and treat that answer as authoritative. Unless it can call the relevant tokenizer, it may simply generate a plausible-looking split.
Embeddings: turning token labels into useful numbers
A token ID tells the model which vocabulary item it has received. To compute with that item, the model needs a numerical representation.
An embedding is a vector: an ordered list of numbers. The model has a learned embedding table, with a vector associated with each token ID.
For a toy system, a lookup might look like this:
Token ID 9827
↓ embedding lookup
[0.18, -0.42, 0.77, 0.05]
Real model vectors generally have many more dimensions. The values are learned during training rather than manually assigned.
What does a dimension mean?
It is tempting to imagine an embedding with neatly labelled axes:
[animalness, friendliness, size, formality]
That is a helpful teaching picture, but usually a misleading description of a real model.
Information is distributed across dimensions, and individual dimensions need not have a simple human-readable meaning. The model learns whatever numerical arrangements help it reduce its training error.
Words used in related ways often develop related representations. However, “related” is broader than “has the same meaning”. Opposites may occur in similar contexts. “Hot” and “cold” both describe temperature and can appear in nearly identical sentence structures.
Consequently, closeness in a vector space does not automatically establish agreement, truth or interchangeability.
A worked example of vector similarity
Consider three invented two-dimensional vectors:
parcel = [1.0, 0.2]
package = [0.9, 0.3]
Tuesday = [-0.2, 1.0]
One common comparison is cosine similarity, which measures how closely two vectors point in the same direction.
cosine_similarity(a, b)
= dot_product(a, b) / (length(a) × length(b))
For parcel and package, the dot product is:
(1.0 × 0.9) + (0.2 × 0.3) = 0.96
Their lengths are approximately 1.020 and 0.949. Dividing 0.96 by their product gives a cosine similarity of about 0.992: very similar directions.
For parcel and Tuesday, the dot product is zero, so their cosine similarity is zero in this invented example.
These numbers are deliberately simple. Real embeddings do not come with such tidy boundaries, and different models create different geometries. Similarity scores only become useful in relation to the embedding model and the task.
Initial embeddings versus contextual representations
The word “bank” can refer to a financial institution or land beside a river.
She deposited the cheque at the bank.
They sat on the bank beside the river.
If the same token represents “bank” in both inputs, its initial token embedding is the same. But its representation changes as the model processes the surrounding words.
After several transformer layers, the vectors associated with those two occurrences can differ substantially. One has incorporated information about cheques and deposits; the other about rivers and sitting.
This distinction is fundamental:
- Initial token embedding: a learned starting representation for a vocabulary item.
- Contextual hidden state: the evolving representation at a particular position after processing its context.
When people casually say “the embedding understands the sentence”, they often blur these stages.
Token embeddings are not automatically search embeddings
Embedding models used for document search often produce one vector for an entire sentence, paragraph or document. They may be specifically trained so that relevant queries and passages are close together.
A generative model’s internal token vectors are not automatically suitable substitutes. Simply averaging arbitrary hidden states may produce poor search results.
For the practical search version of this idea, see Embeddings and Vector Databases: A Practical Introduction.
Position: how the model distinguishes order
Embeddings alone do not tell a model where tokens occur. Yet order changes meaning:
The dog chased the cyclist.
The cyclist chased the dog.
The same major words appear in both sentences, but their relationships differ.
Transformer models therefore need positional information. Different architectures introduce it differently: some add position vectors to token embeddings, while others modify the attention calculation according to position.
A common modern approach is rotary positional embedding, described in the RoFormer paper. It applies position-dependent rotations to parts of the query and key representations used in attention.
You do not need to calculate those rotations to grasp their purpose. They allow attention scores to depend on where tokens are relative to one another, rather than only on token content.
Position is not a perfect memory system
Providing positional information does not guarantee that the model will use every part of a long input effectively. A model can technically accept a document while still overlooking a relevant detail buried inside it.
Nor does increasing a context limit automatically teach a model to handle every possible distance or document structure equally well.
A useful distinction is between:
- Capacity: how much material can be supplied.
- Access: which earlier positions the architecture can attend to.
- Effective use: whether the model actually extracts and applies the right information.
These are related but not interchangeable. We will return to them when discussing context windows.
Attention: choosing which information to combine
Attention is the mechanism that lets one token position gather information from other positions.
In our parcel example, the representation used to answer the question needs information from the sentence containing “Tuesday”. Attention provides a trainable way to route such information through the network.
It does not work by highlighting a word and then reading it as a person would. It performs numerical comparisons and weighted combinations.
Queries, keys and values
Within an attention head, each position’s current representation is transformed into three vectors:
- A query, used to compare that position with available positions.
- A key, used in those comparisons.
- A value, containing information that can be mixed into the output.
These vectors are produced by learned transformations. They are not separately written prompts, database keys or literal dictionary entries.
A useful analogy is a library enquiry:
- The query describes what the current computation is looking for.
- Keys provide features against which that enquiry is matched.
- Values provide the information collected from matching items.
The analogy has limits. The model does not explicitly formulate a human-readable enquiry, and the vectors can represent many features at once.
The calculation, step by step
For one query position, a simplified attention calculation proceeds as follows:
- Take the dot product of its query with each available key.
- Scale the resulting scores.
- Mask any positions that must not be accessed.
- Apply softmax to turn the scores into non-negative weights that sum to one.
- Compute the weighted sum of the corresponding value vectors.
The standard compact formula is:
Attention(Q, K, V) = softmax(QKᵀ / √dₖ + mask)V
Here, dₖ is the size of a key vector. Scaling by its square root helps control score magnitudes. The mask rules out prohibited positions before the weights are calculated.
Suppose three accessible positions receive these invented scores:
Position Score
parcel 1
Tuesday 3
damaged 0
Softmax turns them into weights of approximately:
parcel 0.114
Tuesday 0.844
damaged 0.042
The attention output is then:
0.114 × value(parcel)
+ 0.844 × value(Tuesday)
+ 0.042 × value(damaged)
This is a mixture of vectors, not a copied word. The mixture becomes part of the network’s next representation.
Also note that these weights concern one head at one layer and one query position. They are not the model’s overall confidence that the answer is Tuesday.
Why there are multiple attention heads
A single weighted mixture has limited capacity. Multi-head attention lets the model perform several such calculations in parallel, using different learned projections.
One head may become sensitive to a repeated sequence; another may help connect related grammatical positions. Some heads have identifiable patterns, but it is misleading to assign every head a permanent role such as “the grammar head” or “the facts head”.
The heads’ outputs are combined through another learned transformation. Across layers, these computations build more complex representations.
Attention is therefore less like one spotlight and more like several information-routing operations running together, repeatedly.
Causal masking: no looking ahead
A decoder-only language model uses causal attention. At a given position, it can attend to that position and earlier ones, but not to later tokens.
During training, a whole sequence can be processed with this mask. The computation can be parallelised across positions while preserving the rule that predictions must not see the answers ahead of them.
During generation, later tokens do not yet exist. The model produces them one after another.
This is an important distinction from many text-encoding models, which can let a word representation use both preceding and following text.
Attention is not a transparent explanation
An attention diagram can be informative, but it does not fully explain why a model produced an answer.
Information passes through many heads, layers, residual connections and nonlinear transformations. A high attention weight does not necessarily establish that a particular source token caused the final output in the intuitive sense.
The paper *Attention is not Explanation* examines this problem in attention-based models. The practical lesson is modest: do not treat an attractive heatmap as a complete audit trail.
For architectural context, Transformers vs RNNs: Why Attention Changed Machine Learning explains how this differs from processing sequences through recurrent hidden states.
What happens inside a transformer layer?
Attention is central, but it is not the whole model.
A typical transformer block also contains a feed-forward network, normalisation and residual connections. Exact ordering and implementation vary between architectures.
A simplified block looks like this:
Current token representations
↓ normalisation and attention
Add the attention update through a residual connection
↓ normalisation and feed-forward network
Add the feed-forward update through a residual connection
↓
Updated token representations
The model repeats this pattern through many layers.
Feed-forward networks transform the information
The feed-forward component, often called an MLP, applies learned transformations and nonlinear operations to each position.
In a standard block, the same feed-forward network is applied independently at each position. Attention moves information between positions; the MLP transforms the information now present at each position.
“Independently” does not mean “without context”. The input to the MLP has already been shaped by attention and earlier layers.
Some models use a mixture of experts, routing positions through selected feed-forward sub-networks rather than one shared dense network. That changes the computation and scaling, but not the basic distinction between routing information across positions and transforming representations.
Residual connections preserve a working stream
A residual connection adds a component’s output back to its input instead of replacing the input completely.
updated representation = existing representation + computed update
This helps information and training signals pass through deep networks. A helpful picture is a working document that successive components revise, rather than a message completely rewritten at every stage.
Normalisation helps keep the numerical behaviour manageable. The details matter to model builders, but the user-facing lesson is that stable generation depends on a whole engineered system, not attention alone.
Layers do not form a neat human checklist
It is tempting to say that early layers handle spelling, middle layers handle grammar and later layers handle reasoning.
There can be broad patterns in what different layers represent, but this tidy hierarchy is not a reliable universal description. Features are distributed, tasks interact, and architecture and training affect what happens where.
Similarly, there is no single “fact database layer” that can be opened to inspect everything the model knows. Information is encoded in learned parameters and activated through context-dependent computation.
Learning: how the model acquires its parameters
The embeddings, attention transformations and feed-forward networks contain adjustable numerical parameters, often called weights.
Before useful training, these weights do not encode a capable language model. Training changes them so that the model becomes better at an objective.
For an autoregressive language model, the core pre-training objective is usually next-token prediction.
A small training example
Suppose the training text is:
The parcel arrived on Tuesday.
The model receives progressively available context and is trained to predict the next token at each position. Using word-like pieces for clarity:
Context Target
The parcel
The parcel arrived
The parcel arrived on
The parcel arrived on Tuesday
The parcel arrived on Tuesday .
The model produces a probability distribution over its vocabulary at each position. A loss function penalises assigning low probability to the actual next token.
Backpropagation calculates how changes to parameters would affect that loss. An optimiser then adjusts the parameters. Repeating this over extensive training data teaches patterns that support better prediction.
This does not mean the model merely stores a table of sentence endings. Its shared parameters learn reusable relationships across examples, although models can also memorise some training sequences.
Why prediction can produce broad capabilities
Predicting text well often requires modelling more than adjacent words.
To continue a recipe, the model benefits from representing ingredient relationships. To continue a dialogue, it benefits from tracking speakers. To complete a mathematical explanation, it benefits from learning mathematical patterns and procedures.
A sufficiently capable model can acquire useful abstractions through this training pressure. The GPT-3 paper, Language Models are Few-Shot Learners, documents how examples placed in context can elicit varied tasks without task-specific weight updates.
However, the training objective remains crucial: producing likely continuations is not identical to establishing what is true, safe or appropriate in every situation.
Pre-training is not the end of training
A model trained only to continue text is not necessarily a useful assistant. Post-training can teach more suitable behaviours.
Common approaches include:
- Supervised fine-tuning: training on examples of desired responses.
- Preference-based training: learning from comparisons between better and worse outputs.
- Reinforcement learning: adjusting behaviour using rewards, which may come from preference models or verifiable task outcomes.
The paper *Training language models to follow instructions with human feedback* describes an influential instruction-following and human-feedback approach.
Different systems use different mixtures of these methods. Post-training can improve instruction following, usefulness and caution, but it does not make every answer correct.
For a fuller account, see Pre-training, Fine-tuning and RLHF: How Chatbots Are Trained.
Learning from a prompt is usually not a weight update
When you give a model examples in a conversation, it can adapt its response to them. That is often called in-context learning.
Normally, the model’s weights remain unchanged during that interaction. The examples influence its current computation because they are part of the input.
A product may separately save memories or later use conversations for training, depending on its settings and policies. Neither should be confused with the immediate mechanism of answering your prompt.
Generation: how scores become an answer
After the final transformer layer, the model converts a hidden representation into one score for each possible next token. These scores are called logits.
A softmax operation can turn the logits into a probability distribution. A decoding method then selects a token.
The selected token is added to the context, and the process repeats.
Input: When did the parcel arrive?
Output step 1: Tuesday
Output step 2: .
Output step 3: end-of-message token
This is illustrative: the real tokenizer might divide the answer differently, and a model might choose a longer response.
Greedy decoding versus sampling
Greedy decoding selects the highest-scoring token at each step. It is straightforward, but locally choosing the most likely token does not guarantee the best overall sequence.
Sampling selects from the distribution, allowing lower-probability alternatives to appear. This supports variety but can also introduce undesirable variation.
Two common controls are:
- Temperature: rescales logits before sampling. Lower values concentrate probability on stronger candidates; higher values flatten the distribution.
- Top-p: limits sampling to a set of high-probability tokens whose cumulative probability reaches a chosen threshold.
The exact behaviour depends on the implementation. Even settings intended to be highly repeatable may not guarantee identical outputs across infrastructure changes or model updates.
Most importantly, lowering temperature does not transform false knowledge into true knowledge. It can make the same mistake more consistently.
Temperature, Top-p and Sampling: How AI Chooses Its Next Word explores these controls in more detail.
A probability is not a truth score
Suppose a model strongly favours the token sequence “Tuesday”. That means the sequence fits the model’s conditional distribution given the input.
It does not mean an independent verification process has established Tuesday as the delivery date.
In our example, the supplied passage supports the answer, so confidence and correctness may align. In an obscure factual question without evidence, a familiar-looking completion can receive high probability despite being wrong.
This distinction explains much of the danger of fluent output: grammatical confidence is visible; evidential support may be absent.
Why longer answers take more work
Generating output requires repeated decoding steps. Systems often cache previously computed attention keys and values so they do not have to recalculate everything from scratch.
This KV cache improves efficiency, but it consumes memory. Longer inputs and outputs increase resource demands.
The initial processing of a prompt is often called prefill. Subsequent token-by-token production is decoding. These phases have different performance characteristics, which helps explain why an assistant may pause before starting and then stream text at a steadier pace.
Context windows: temporary working material, not unlimited memory
A context window limits the sequence a model can process in a given operation.
Depending on the model and API, input and generated output may share a total allowance, and separate output limits may also apply. The practical budget can include system instructions, conversation history, retrieved passages and tool results—not just your latest message.
A chat interface can conceal this complexity. It may truncate older messages, summarise them or retrieve selected parts rather than send the full history every time.
Fitting does not guarantee using
Imagine supplying a long set of delivery records and asking for one parcel’s arrival date. The relevant line may fit comfortably within the advertised window, yet the model can still miss it.
The study *Lost in the Middle* found that performance on evaluated retrieval-style tasks could depend substantially on where relevant information appeared in the context. This is not a fixed law for every newer model, but it illustrates why capacity is not the same as reliable use.
Other difficulties include:
- Several records with similar identifiers.
- Conflicting dates without clear version labels.
- Irrelevant text that resembles the answer.
- Instructions separated from the material they govern.
Good document organisation helps because it makes the intended relationships easier to represent.
Attention also has a computational cost
In standard dense attention, the number of pairwise attention scores grows quadratically with sequence length. Doubling the length produces four times as many position pairs.
That does not mean every aspect of real-world runtime or memory follows exactly the same rule. Efficient kernels, caching, sparse patterns and alternative architectures change the practical costs.
Nevertheless, long context is not free. Sending an entire archive for every question can be expensive and less reliable than selecting the relevant material first.
What Is a Context Window and Why It Limits What AI Can Do covers the budgeting implications.
A complete worked example: answer from evidence
Let us turn the mechanism into a practical workflow.
Suppose you have this record:
Delivery record
Parcel ID: P-1847
Expected arrival: Monday
Actual arrival: Tuesday
Condition on arrival: Outer packaging damaged
Inspection completed: Thursday
You want a concise answer to “When did parcel P-1847 arrive?”
Step 1: make the evidence unambiguous
Use labels that distinguish the dates. Without them, a model might conflate the expected arrival, actual arrival and inspection date.
This is not special syntax the architecture requires. It is ordinary document design that reduces ambiguity.
Step 2: specify the operation
A useful prompt is:
Use only the delivery record below.
Answer the question in one sentence.
Include the parcel ID and quote the field that supports your answer.
If the actual arrival is not stated, say "Actual arrival not stated".
<record>
Parcel ID: P-1847
Expected arrival: Monday
Actual arrival: Tuesday
Condition on arrival: Outer packaging damaged
Inspection completed: Thursday
</record>
Question: When did parcel P-1847 arrive?
The delimiters separate source material from instructions. They are organisational aids, not a security boundary or a guarantee of compliance.
Step 3: understand the internal route
The tokenizer encodes the instructions, record and question.
Embedding lookups supply initial vectors. Positional information helps preserve structure and order. Across transformer layers, attention and feed-forward operations develop representations that connect the requested parcel with its actual-arrival field.
The output projection scores possible next tokens. Decoding selects a sequence that may become:
Parcel P-1847 arrived on Tuesday, supported by
"Actual arrival: Tuesday".
The model does not need a literal database row lookup inside its architecture to produce this result. But it also has no automatic guarantee that the selected text is correct.
Step 4: test a missing-data case
Remove the actual-arrival field and run the prompt again.
The desired output is:
Actual arrival not stated
This is a more revealing test than the original easy case. A weak workflow may substitute Monday, because an expected date is available, or Thursday, because it is another recorded event.
Step 5: test conflict and exactness
Now add:
Correction:
Parcel P-1847 actual arrival updated to Wednesday.
This correction supersedes the earlier arrival field.
Ask whether the model respects the correction and preserves the parcel identifier exactly.
You now have three useful test cases: ordinary evidence, missing evidence and superseded evidence. Together they reveal more than a single successful demonstration.
For high-volume processing, require structured fields and validate them programmatically. For consequential decisions, retain a route to the original record and a human review step.
What this mechanism explains—and what it does not
A good mental model should improve your decisions, not merely supply vocabulary.
Why hallucinations happen
A generative model is built to produce continuations, not to guarantee that every claim has a source.
When evidence is weak, missing or ambiguous, it may still generate a convincing answer. Learned associations, misleading context and decoding choices can all contribute.
Giving it reliable evidence and an explicit missing-information response helps, but does not eliminate the problem. Quoted text can also be copied inaccurately, and citations can be invented.
The practical remedy is to check important claims against sources rather than infer reliability from fluency. See Why AI Hallucinates: Causes, Types and How to Reduce Them.
Why retrieval helps without fixing everything
Retrieval supplies relevant documents at answer time. It can provide current or private information that is absent from the model’s learned parameters.
But retrieval can select the wrong passage, miss a correction or supply contradictory material. The model can then misread even a good passage.
Separate the checks:
- Did the system retrieve the right evidence?
- Did the answer faithfully reflect that evidence?
- Was the evidence itself reliable and current?
This separation makes troubleshooting much easier than labelling the entire system “unreliable”.
Why tools matter for precise tasks
A language model can explain arithmetic, generate a spreadsheet formula or write a database query. That does not make its unaided text generation the best place to perform exact calculations.
A calculator executes arithmetic. A database applies explicit query rules. A parser can count characters in a string.
Use the model to interpret the request and arrange the work; use appropriate tools for operations with strict correctness requirements. Then inspect the inputs and outputs, because tool selection and argument generation can also fail.
What not to infer about understanding
Calling an LLM “just autocomplete” understates the complexity of what it can learn. Calling it a human-like thinker overstates what its fluent language establishes.
The mechanism supports sophisticated pattern learning, contextual representation and useful problem-solving. It does not, by itself, establish consciousness, intentions or human-like experience.
For practical use, the better question is usually: What evidence shows that this system performs this task reliably under these conditions?
Common mistakes and better habits
Mistake: treating more context as automatically better
Large quantities of loosely related material can obscure the evidence that matters.
Better habit: provide relevant sections, label their source and date, and remove stale duplicates. Preserve enough surrounding context to avoid changing the meaning.
Mistake: assuming a prompt teaches permanent knowledge
Correcting a chatbot may influence the current conversation without changing its underlying weights.
Better habit: supply recurring requirements through supported instructions, saved settings or application design. Test whether they are actually present in later interactions.
Mistake: treating an explanation as a faithful execution trace
A model can generate a plausible explanation of its answer. That text is not necessarily a complete account of the internal computations that produced it.
Better habit: request checkable evidence, calculations, assumptions and intermediate results where useful. Verify those artefacts rather than relying on a narrative of certainty.
Mistake: confusing token similarity with correctness
Related vectors can support useful matching, but similarity does not prove that two passages agree.
Better habit: distinguish retrieval from verification. A passage about cancellations may be relevant to a cancellation question while describing the wrong policy.
Mistake: using one successful answer as an evaluation
Generation can vary, and easy examples hide important failure modes.
Better habit: build a small test set containing normal cases, missing information, conflicting evidence, unusual identifiers and formatting requirements.
Keep the expected answer or checking rule for each case. Repeat testing after changing the prompt, model or document preparation.
Exercise: build a compact reliability checklist
Choose a real, low-risk task such as extracting delivery dates or summarising public event listings.
Write down:
- What information the model must receive.
- What it should do when a required fact is absent.
- Which fields must be reproduced exactly.
- Which operations need a tool.
- How you will check the output.
- What kind of failure requires human review.
Then run five deliberately varied examples. Record errors by category rather than simply marking each response “good” or “bad”.
This turns knowledge of tokens, embeddings and attention into a useful habit: designing the task around the system’s strengths while making its failures visible.
FAQ
Does a large language model predict words or tokens?
Usually tokens. A token may represent a whole word, a word fragment, punctuation, whitespace or another text unit.
“Predicting the next word” is convenient shorthand, but it can mislead when discussing billing, character counting, unusual identifiers or context limits. Exact counts require the tokenizer associated with the model.
Are embeddings the same as the model’s knowledge?
No. Token embeddings are one set of learned parameters, but useful information is distributed throughout the model’s attention and feed-forward components as well.
Contextual representations also change as the input passes through layers. There is no single embedding table containing a clean, readable catalogue of everything the model has learned.
Does attention let the model read the whole prompt at once?
During prompt processing, many calculations can run in parallel. However, a causal model still uses a mask that prevents each position from accessing later positions.
An answer position can generally use earlier prompt positions, subject to the architecture and available context. That access does not guarantee that every relevant detail will be used correctly.
Why can a model answer questions it has never seen before?
Training can produce reusable representations and procedures rather than only memorised responses. The model can combine learned patterns with new information in the prompt.
That supports generalisation, but not unlimited generalisation. Unfamiliar formats, unusual combinations and tasks outside its learned strengths can expose brittle behaviour.
Does temperature zero stop hallucinations?
No. Selecting the highest-scoring continuation does not establish that the continuation is true.
A low-temperature configuration can be useful for consistent extraction or formatting, but factual reliability still depends on evidence, task design and verification. A model can confidently repeat the same incorrect answer.
Can a longer context window replace document search?
Sometimes a whole document fits comfortably and can be supplied directly. For larger collections, retrieval is often more efficient and easier to manage.
Even when everything fits, the model may struggle with competing records or buried details. Compare approaches using representative questions and known answers rather than choosing solely by advertised window size.
Is the model searching the internet when it answers?
Not necessarily. Core text generation uses the supplied context and learned parameters.
A surrounding application may provide a search tool, retrieve pages and add their contents to the context. Check the product’s visible tool activity and sources rather than assuming that a recent-sounding answer was researched.
What is the most useful mental model to remember?
Think of an LLM as a trained system that repeatedly transforms numerical representations of context into predictions about the next token.
Tokens define the pieces. Embeddings provide starting representations. Attention moves information between positions. Other network components transform it. Decoding turns scores into text.
That process can be remarkably useful—but evidence and verification remain separate responsibilities.
Sources
- Attention Is All You Need
- Hugging Face: Tokenizer summary
- The Illustrated Transformer
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- Language Models are Few-Shot Learners
- Training language models to follow instructions with human feedback
- Lost in the Middle: How Language Models Use Long Contexts
- Attention is not Explanation
About the author
Editorial team · Editorial team
Author identity has not been supplied yet. Replace this record with the real writer's name, background and verifiable experience before publishing anything on the public site.
Spotted an error? Report a correction.
Related reading
A Repeatable AI Research Workflow: From Question to Verified Brief
A disciplined research process that uses AI for planning and synthesis while keeping every important claim tied to evidence you have checked.
How to Summarise Long Documents with AI Without Missing What Matters
A practical, source-grounded workflow for turning long documents into reliable summaries while preserving caveats, contradictions and important detail.
AI Meeting Notes: A Safe Workflow from Transcript to Action Items
A careful end-to-end method for using AI to draft meeting notes without inventing decisions, owners or deadlines.