Skip to content

Temperature, Top-p and Sampling: How AI Chooses Its Next Word

Temperature and top-p shape how an AI samples its next token, but useful control starts with understanding what these settings can and cannot change.

By Editorial teamPublished 22 min read

Key takeaways

  • Language models generate tokens, not necessarily whole words, by repeatedly scoring possible continuations.
  • Temperature reshapes a probability distribution; top-p restricts sampling to a probability-based shortlist.
  • Lower randomness does not guarantee factual accuracy, and higher randomness does not guarantee useful creativity.
  • Change one setting at a time and compare multiple outputs using task-specific checks.
  • Prompts, source material, schemas and verification often matter more than fine-tuning sampling settings.
On this page
  1. The choice hiding inside every AI answer
  2. Before sampling: what the model produces
  3. Greedy decoding and random sampling
  4. Temperature: changing the relative weights
  5. Top-p: building a probability-based shortlist
  6. Temperature and top-p together
  7. Other controls that can change the result
  8. Exercise 1: build a miniature sampler
  9. Exercise 2: test settings on a real task
  10. Choosing settings for different jobs
  11. Common mistakes and how to avoid them
  12. A compact method for sensible tuning
  13. FAQ

The choice hiding inside every AI answer

Ask an AI assistant to finish “The best way to learn is…” and it might continue with “to practise”, “by doing”, or “to ask better questions”. All three are plausible. Which one appears depends partly on the model and the prompt, and partly on how the system selects from the model’s possible continuations.

That selection process is called decoding. When it involves random draws, it is called sampling. Temperature and top-p are two common controls over that process.

These controls are useful, but they are often described badly. Temperature is not an intelligence setting. Top-p is not a truth filter. Neither adds knowledge, retrieves missing evidence, or makes a model understand your intentions better.

They influence which of the continuations already available to the model gets chosen.

For practical purposes, keep three questions separate:

  • What can the model do? This depends on its training, architecture and available tools.
  • What have you asked it to do? This depends on the instructions, examples and information in its context.
  • Which continuation does it select? This depends partly on the decoding method and its settings.

This article explains that third question, then connects it back to the first two. You will see the arithmetic behind temperature and top-p, run a small sampler yourself, and learn how to test settings without confusing attractive wording with better results.

Before sampling: what the model produces

A “next word” is usually a next token

Although we talk about AI choosing its next word, language models usually generate tokens. A token might be a whole word, part of a word, punctuation, a space-linked word fragment, or another unit defined by the model’s tokeniser.

For example, an unfamiliar surname might need several tokens. A common short word might need only one. The boundaries depend on the tokeniser, so do not assume a particular word always corresponds to a single token.

The model also has ways to signal that generation should stop. In a chat system, some special tokens may mark message boundaries or other structural information.

The simplified generation loop is:

  1. Read the available context.
  2. Score possible next tokens.
  3. Apply the chosen decoding rules.
  4. Select one token.
  5. Add that token to the context.
  6. Repeat until a stopping condition is reached.

Consequently, temperature does not choose the “creativity level” of an entire paragraph in one operation. It acts on successive token choices, whose effects accumulate.

For the underlying machinery, see How Large Language Models Work: Tokens, Embeddings and Attention. The transformer architecture behind many current models was introduced in Attention Is All You Need.

Scores come before probabilities

At each step, a model produces numerical scores called logits. These scores are not probabilities: they do not have to sit between zero and one, or add up to one.

A mathematical operation called softmax converts them into a probability distribution. After that conversion, every candidate has a non-negative probability, and the probabilities sum to one.

Imagine this context:

At the café, I ordered a cup of

For illustration, suppose the next-token distribution contains:

Candidate tokenProbability
tea0.50
coffee0.25
cocoa0.15
water0.07
sand0.03

These are invented numbers for a tiny vocabulary. A real model typically scores a much larger vocabulary, and the token strings may include leading spaces.

The model strongly favours “tea”, but it has not assigned “tea” a probability of one. A sampler could select something else.

Token probability is not factual confidence

A probability of 0.50 for “tea” means something like “under this model and context, this is a likely continuation”. It does not mean there is a 50% chance that a real person ordered tea.

This distinction matters whenever the model answers a factual question.

A familiar but false claim can be a highly likely continuation. A correct but obscure name can be less likely. A polished sentence can have plausible token choices throughout and still describe something that never happened.

Sampling settings act on continuation probabilities, not on a separate measurement of truth. This is why a fluent, repeatable answer may still require checking.

Greedy decoding and random sampling

Greedy decoding takes the local winner

The simplest selection rule is greedy decoding: choose the highest-scoring token at every step.

In the café example, greedy decoding chooses “tea”. It then calculates a fresh distribution conditioned on the context including “tea”, chooses the next winner, and continues.

Greedy decoding is locally decisive. However, choosing the most probable token now does not necessarily produce the most probable complete sequence, much less the most useful answer.

Consider choosing a walking route by always taking the street that initially points most directly towards your destination. That local rule can work, but it does not evaluate every possible route.

Greedy output can be concise and stable. It can also be repetitive, conventional or confidently wrong. Whether those problems appear depends heavily on the model and task.

Sampling draws from the distribution

Ordinary sampling treats the probabilities as weights in a random draw.

Using the café distribution, “tea” wins more often than “coffee”, but “coffee” still wins sometimes. Over many independent draws from exactly that distribution, the observed frequencies tend towards the stated probabilities.

A single draw proves very little. Selecting “sand” once does not mean the sampler is broken; a 3% event is uncommon, not impossible.

In actual generation, the distribution changes after each token. If one run selects “tea” and another selects “coffee”, the two runs now have different contexts. Their later choices can diverge further.

That is how a small early difference can become a different example, paragraph structure or conclusion.

Random does not mean unstructured

Sampling is not equivalent to picking arbitrary words from a dictionary. It is random selection weighted by the model’s predictions.

Even at settings that encourage variety, likely continuations usually retain an advantage. But broadening selection can admit weak candidates, and weak choices can create contexts that invite further weak choices.

The challenge is therefore not simply to maximise or eliminate randomness. It is to allow useful alternatives without letting low-quality options dominate the result.

Temperature and top-p offer two different ways to manage that trade-off.

Temperature: changing the relative weights

The mathematical idea

Temperature changes how strongly the sampler favours higher-scoring tokens.

For a positive temperature, the usual calculation divides each logit by the temperature before applying softmax:

probability(token i) = exp(logit_i / T) / sum_j exp(logit_j / T)

Here, T is temperature. The denominator makes all the resulting probabilities add up to one.

The practical effects are:

  • Below 1: the distribution becomes sharper; stronger candidates gain relative weight.
  • At 1: the distribution retains its original shape.
  • Above 1: the distribution becomes flatter; weaker candidates gain relative weight.

For positive temperatures, this transformation preserves the ranking of the logits. If “tea” outranked “coffee” beforehand, it still outranks it afterwards.

Temperature changes the size of that advantage, not the identity of the original front-runner.

API ranges vary. Some systems expose only a limited interval, some offer presets, and some models do not support user-controlled temperature at all. Check the documentation for the specific model rather than assuming a setting transfers unchanged.

A worked temperature example

Starting from known probabilities at temperature 1, we can calculate a new distribution by raising each probability to the power 1/T, then normalising:

new_probability_i = old_probability_i ** (1 / T)
                    / sum(old_probability_j ** (1 / T))

This assumes the same underlying logits and no other transformations.

For our café example:

CandidateTemperature 0.5Temperature 1Temperature 2
tea73.36%50.00%34.79%
coffee18.34%25.00%24.60%
cocoa6.60%15.00%19.05%
water1.44%7.00%13.02%
sand0.26%3.00%8.52%

Rounded percentages may not add up exactly to 100%.

At temperature 0.5, we square the original probabilities. “Tea” starts at 0.50, so its unnormalised weight becomes 0.25. “Sand” starts at 0.03, so its weight becomes 0.0009.

The original tea-to-sand probability ratio was about 16.7 to one. After squaring, it is about 277.8 to one. Normalisation changes the overall scale but preserves that ratio.

At temperature 2, we take square roots. This narrows the gaps and gives less likely options a larger share.

What temperature zero means

The formula above cannot be used literally at zero, because it would require division by zero.

Interfaces that accept temperature=0 commonly interpret it as a request for greedy or near-deterministic behaviour. That is an implementation convention, not an ordinary substitution into the formula.

Even then, identical API requests are not always guaranteed to produce identical outputs. Serving infrastructure, numerical effects, routing, model updates and tie handling can matter.

Treat zero as “minimise sampling variability under this implementation”, not “create a permanent reproducibility guarantee”.

Why temperature is not a creativity dial

Higher temperature can produce more unusual wording and less obvious associations. Some of those may be useful creative departures. Others may be irrelevant, inconsistent or incoherent.

Useful creativity also requires direction: audience, purpose, constraints, examples and a way to select good ideas.

Compare these two prompts:

Give me creative names for a repair business.
Generate 12 names for a neighbourhood electronics repair shop.

Audience: local residents who value clear prices and friendly service.
Tone: memorable, trustworthy, slightly playful.
Constraints: two or three words; avoid "tech", "genius" and "solutions".
Include names built around restoration, second chances and local identity.

The second prompt creates purposeful variety even with restrained sampling. Higher temperature alone cannot supply that brief.

Top-p: building a probability-based shortlist

How nucleus sampling works

Top-p, also called nucleus sampling, limits the candidates eligible for selection.

The usual procedure is:

  1. Sort tokens from highest to lowest probability.
  2. Keep the smallest leading group whose cumulative probability reaches or exceeds p.
  3. Remove the remaining candidates.
  4. Renormalise the retained probabilities.
  5. Sample from that shortlist.

The name “nucleus” refers to the retained core of the distribution.

The method was introduced and studied in The Curious Case of Neural Text Degeneration, which examined problems with decoding open-ended text. Its findings help explain why always maximising likelihood is not automatically the best strategy for readable generation.

A worked top-p example

Return to the original café distribution:

CandidateProbabilityCumulative probability
tea0.500.50
coffee0.250.75
cocoa0.150.90
water0.070.97
sand0.031.00

At top_p=0.80, we retain “tea”, “coffee” and “cocoa”.

Why include cocoa? Because tea and coffee together reach only 0.75. We need the next candidate to meet or exceed 0.80.

The retained mass is 0.90, so the new probabilities are:

tea:    0.50 / 0.90 = 55.56%
coffee: 0.25 / 0.90 = 27.78%
cocoa:  0.15 / 0.90 = 16.67%
water:  0%
sand:   0%

Top-p does not make all retained candidates equally likely. It preserves their relative weights unless another transformation changes them.

At top_p=0.50, a standard implementation would retain only “tea” here. At top_p=1, this filter ordinarily removes no probability mass.

Exact boundary handling and minimum retained-token rules can vary between implementations.

The shortlist changes at every step

Top-p is not a fixed vocabulary size.

If one token has probability 0.96, a threshold of 0.90 may leave just that token. If probability is spread across many candidates, reaching 0.90 may require a large shortlist.

This adaptive behaviour is its defining feature. It can allow broad choice when the model distributes probability widely, while narrowing choice when one continuation dominates.

However, a broad distribution is not a reliable measurement of real-world uncertainty. Several equally valid ways to phrase a sentence can spread probability without any factual uncertainty. Conversely, a model can concentrate probability on an incorrect answer.

Top-p responds to the distribution, not to the reason the distribution has that shape.

Temperature and top-p together

Different controls, overlapping effects

Temperature and top-p both influence diversity, but they do different jobs.

ControlMain operationCan exclude candidates entirely?
TemperatureRescales relative probabilitiesNot ordinarily, for finite logits and positive temperature
Top-pTruncates the distribution by cumulative massYes
Top-kKeeps a fixed number of leading candidatesYes
Greedy decodingSelects the highest-scoring candidateAll others are unselected

Temperature changes the slopes between candidates. Top-p draws a boundary around a retained group.

An important consequence follows: the order of these operations can affect the result.

A worked interaction example

Suppose an implementation applies temperature first and then top-p.

At temperature 1 and top_p=0.80, our shortlist contains tea, coffee and cocoa.

At temperature 0.5, tea has about 73.36% and coffee about 18.34%. Together they account for roughly 91.70%. The same top-p threshold now keeps only those two candidates.

At temperature 2, the cumulative probabilities are approximately:

tea:                         34.79%
tea + coffee:                59.39%
tea + coffee + cocoa:        78.44%
tea + coffee + cocoa + water: 91.46%

Now top_p=0.80 retains four candidates.

The top-p value has not changed, but its shortlist has, because temperature reshaped the distribution first.

If a system filters before applying temperature, the interaction differs. Hosted APIs may not expose every detail of their decoding pipeline. The Hugging Face generation configuration reference is a useful guide to common controls, but individual providers can implement different rules.

Why changing one setting is a good default

When learning or tuning a workflow, leave top-p at its default while testing temperature. Alternatively, fix temperature and compare a few top-p settings.

Changing both at once makes diagnosis difficult. If one result is better, you will not know which change helped, or whether their interaction caused the improvement.

The OpenAI chat completions reference likewise recommends altering temperature or top-p rather than both as a general approach.

This is experimental discipline, not a law of mathematics. Experienced users can tune both, especially with controlled evaluation. The point is to earn the extra complexity rather than begin with it.

Other controls that can change the result

Top-k and repetition penalties

Top-k keeps a fixed number of highest-probability candidates. With top_k=3, the shortlist contains three tokens, regardless of whether they account for 40% or 99% of the probability.

That differs from top-p, whose shortlist expands or contracts with the distribution.

Some systems offer repetition, frequency or presence penalties. These modify scores based on previous token use. They can discourage loops, but excessive penalties can also interfere with necessary repetition: names, technical terms and consistent labels often need to recur.

The penalties are not interchangeable across providers. A “repetition penalty” may use a different formula from a “frequency penalty”.

Before adjusting one, identify an actual repetition problem. A repeated customer name in a support summary is not necessarily a decoding failure.

Beam search and other decoding strategies

Beam search keeps several partial sequences and extends promising candidates, rather than committing immediately to one path.

It can be useful in some generation tasks, but it is not simply a superior version of sampling. Strong preference for likely sequences can produce conventional or repetitive open-ended text.

The Hugging Face guide to generation strategies explains the distinction between greedy decoding, sampling and search-based approaches.

For ordinary hosted chat use, you may never see these controls. Knowing that they exist is still helpful: identical temperature values do not imply identical overall decoding methods.

Length limits and structured-output constraints

A maximum output-token setting controls how much can be generated. It does not directly control randomness.

An answer cut off halfway through a sentence may need a larger output allowance, not a different temperature. Some systems also account separately for internal processing tokens, so read the model-specific limits.

Structured-output features can restrict which next tokens are allowed so that the result follows a grammar or schema. Conceptually, this is different from merely favouring probable tokens.

If your application requires JSON, use supported schema constraints and validate the result. Do not rely on low temperature to enforce syntax. See Prompting for Structured Output: JSON, Tables and Schemas for practical patterns.

Exercise 1: build a miniature sampler

You can explore the mechanics without an API account. The following Python example uses only the standard library.

It starts from our invented probability distribution, applies temperature, applies top-p, and draws samples.

import random
from collections import Counter

BASE = {
    "tea": 0.50,
    "coffee": 0.25,
    "cocoa": 0.15,
    "water": 0.07,
    "sand": 0.03,
}

def temperature_scale(probabilities, temperature):
    if temperature <= 0:
        raise ValueError("Use a positive temperature.")

    weights = {
        token: probability ** (1.0 / temperature)
        for token, probability in probabilities.items()
    }
    total = sum(weights.values())

    return {
        token: weight / total
        for token, weight in weights.items()
    }

def nucleus_filter(probabilities, top_p):
    if not 0 < top_p <= 1:
        raise ValueError("top_p must be in (0, 1].")

    ranked = sorted(
        probabilities.items(),
        key=lambda item: item[1],
        reverse=True,
    )

    kept = {}
    cumulative = 0.0

    for token, probability in ranked:
        kept[token] = probability
        cumulative += probability
        if cumulative >= top_p:
            break

    total = sum(kept.values())
    return {
        token: probability / total
        for token, probability in kept.items()
    }

def experiment(temperature, top_p, draws=10000, seed=42):
    scaled = temperature_scale(BASE, temperature)
    filtered = nucleus_filter(scaled, top_p)

    rng = random.Random(seed)
    tokens = list(filtered)
    weights = list(filtered.values())
    counts = Counter(rng.choices(tokens, weights=weights, k=draws))

    print(f"\nTemperature={temperature}, top_p={top_p}")
    for token, expected in filtered.items():
        observed = counts[token] / draws
        print(
            f"{token:>6}: expected={expected:.3f}, "
            f"observed={observed:.3f}"
        )

experiment(temperature=1.0, top_p=1.0)
experiment(temperature=0.5, top_p=1.0)
experiment(temperature=2.0, top_p=1.0)
experiment(temperature=1.0, top_p=0.8)

What to do, step by step

First, run the script unchanged. The observed proportions should be close to the expected proportions, though not identical.

Second, reduce draws from 10,000 to 20. The results will fluctuate much more. This illustrates why judging a setting from one or two chatbot replies is unreliable.

Third, keep temperature at 1 and try top-p values of 0.5, 0.8 and 0.95. Watch which tokens disappear completely.

Fourth, keep top-p at 0.8 and change temperature. Notice that the shortlist itself changes, not just the weights inside it.

Finally, change the seed. A seed initialises the pseudo-random generator. With the same code and inputs in the same environment, a fixed seed makes this toy experiment repeatable.

This implementation is educational rather than production-ready. Real systems usually work directly with logits and use numerically stable operations. Also, our script repeatedly samples one unchanged distribution; a language model recalculates its distribution after every generated token.

That limitation is useful: it isolates the behaviour of the sampler before adding the complexity of full text generation.

Exercise 2: test settings on a real task

A practical comparison needs more than “this answer feels better”. It needs a task, controlled inputs and a definition of success.

Step 1: choose something you can check

Use a short extraction task with a known answer:

Extract the appointment details from the note below.

Return only a JSON object with these keys:
"name", "date_text", "time_text", "location", "needs_confirmation".

Use null for missing values.
Do not infer a year.
Set needs_confirmation to true if any required detail is missing
or explicitly uncertain.

Note:
Maya Patel is booked for 14 September at 10:30.
The location has not yet been confirmed.

Before running the model, write down the expected content:

{
  "name": "Maya Patel",
  "date_text": "14 September",
  "time_text": "10:30",
  "location": null,
  "needs_confirmation": true
}

The exact whitespace is unimportant. The field values, types and absence of invented details matter.

Step 2: establish a baseline

Use one model and keep its version fixed where possible. Keep the prompt, output limit, schema mode and all other controls unchanged.

Start with the provider’s default settings and run the prompt several times. Ten runs per condition can be a manageable exploratory start, but it is not enough to support strong reliability claims.

Record:

  • Whether the JSON parses.
  • Whether every required key appears.
  • Whether field values are correct.
  • Whether the model invents a location or year.
  • Whether extra commentary appears.

If structured outputs are enabled, keep them enabled in every condition. Otherwise, you would be comparing both sampling and format enforcement.

Step 3: vary only temperature

Try a low, default and higher supported temperature. Leave top-p unchanged.

Do not assume that “low” must mean precisely 0.2 or that “higher” must mean 1.2. Supported ranges and effective behaviour differ. Choose values appropriate to your API.

Use a simple worksheet:

ConditionValid JSONCorrect fieldsInvented detailsNotes
Low temperature
Default temperature
Higher temperature

A tiny, easy task may show no differences. That is a legitimate result. It means you have not demonstrated a benefit from changing the setting on these examples.

Step 4: expand the test cases

Now introduce ambiguous notes, missing times, multiple appointments and corrections:

Maya's appointment was originally 14 September at 10:30.
It has moved to 16 September; the new time is still being arranged.
The venue remains unconfirmed.

A useful setting must work across the kinds of inputs you actually receive, not just one neat example.

For more consistent interpretation, add carefully chosen demonstrations. Few-Shot Prompting: Designing Examples That Teach the Model explains why examples can clarify the task more directly than sampling adjustments.

Choosing settings for different jobs

Extraction, classification and administrative work

For extracting fields, assigning labels or transforming records, unnecessary variation is usually a cost. Start with low randomness or the model’s recommended default, plus clear rules and validation.

The crucial question is not whether every response looks identical. It is whether every response satisfies the contract.

Two valid JSON objects with different key order may be operationally equivalent. Two identical objects with an invented date are consistently wrong.

Where possible, compare parsed values rather than raw text. Route uncertain or incomplete cases to review rather than trying to make the sampler eliminate ambiguity.

Summaries and explanations

Summaries need both faithfulness and editorial judgement. Different outputs may emphasise different details without any obvious grammatical problem.

Use restrained or default sampling initially, then evaluate omissions, invented claims and distortion of emphasis.

A useful summary prompt should specify:

Summarise the document for a reader deciding whether to approve the project.

Include:
- the requested decision;
- the proposed cost and timeline;
- unresolved risks;
- any disagreement between contributors.

Do not add facts absent from the document.
If a requested item is missing, say "not stated".

Temperature cannot compensate for source text that is absent or truncated. What Is a Context Window and Why It Limits What AI Can Do explains why the information available to the model matters before decoding even begins.

Brainstorming and creative drafting

For idea generation, some variability is desirable. Start at defaults and increase temperature cautiously if outputs remain too similar.

Also diversify the brief explicitly:

Suggest nine workshop titles about repairing household objects.

Give three titles in each style:
1. Practical and direct.
2. Warm and community-focused.
3. Playful but understandable.

Avoid rhymes and environmental slogans.

This requests different categories of ideas instead of hoping random token variation will discover them.

Separate generation from selection. First produce options; then assess them against criteria such as clarity, relevance, originality and suitability for the audience.

A slightly unusual title that meets the brief is more useful than a highly unusual title nobody understands.

Coding, calculations and consequential answers

For code, begin with restrained or recommended settings, precise requirements and tests. Higher variability can help explore alternative implementations, but it can also introduce incompatible assumptions.

For calculations, use a calculator or executable code when appropriate. The sampler does not become an arithmetic engine at temperature zero.

For consequential factual work, source quality and verification should dominate your attention. Retrieval-Augmented Generation (RAG) Explained for Beginners describes how supplying relevant evidence can improve the information available to the model.

Retrieval still needs checking: the source may be outdated, irrelevant or misinterpreted. Sampling settings cannot repair those failures by themselves.

Common mistakes and how to avoid them

Mistaking consistency for accuracy

If the model repeats the same false claim ten times, you have demonstrated consistency, not truth.

Low temperature can make errors more repeatable because the same high-probability continuation keeps winning. Conversely, a more variable response might occasionally contain the correct answer without becoming a dependable method.

Check outputs against evidence or tests. Why AI Hallucinates: Causes, Types and How to Reduce Them covers why plausible language and factual grounding are different properties.

Treating top-p as a confidence threshold

A top-p value of 0.9 does not mean “only say things you are 90% sure about”. It means retain a leading group of tokens containing roughly at least 90% of the current probability mass.

The distinction is especially important in interfaces that label top-p vaguely. It is a sampling threshold, not an answer-level confidence score.

Similarly, lowering top-p does not force the model to admit ignorance. An explicit rule about missing evidence is more relevant, although that rule also needs evaluation.

Changing everything at once

A new prompt, new model, new temperature and new top-p may improve your result. But you cannot attribute the improvement to a particular change.

Keep a baseline. Change one variable. Save the full configuration with the outputs.

For serious use, record the model identifier, prompt version, sampling settings, schema configuration, output limit and test date. If the provider exposes a model snapshot or infrastructure fingerprint, record that too.

Without this information, later comparisons can become exercises in reconstructing what happened.

Over-interpreting a single good response

A higher-temperature run may produce a brilliant opening paragraph. The next run may ignore the brief.

Evaluate the distribution of outcomes, not your favourite example. For production work, the frequency and severity of failures often matter more than the best possible response.

This is also why generating more candidates has a cost. Someone or something must identify which candidate is actually good.

Assuming agreement proves correctness

Sampling several answers and choosing a majority can help with some tasks. The paper Self-Consistency Improves Chain of Thought Reasoning in Language Models studies a related approach of sampling multiple reasoning paths and aggregating answers.

However, outputs from one model are not independent expert witnesses. They share training, context and likely misconceptions.

Agreement is useful evidence only within an evaluated method. It does not replace checking a citation, running a test or consulting the actual source document.

Confusing sampling with model training

Temperature and top-p normally operate during generation. They do not update the model’s learned weights.

This differs from fine-tuning and other adaptation methods. For example, The Power of Scale for Parameter-Efficient Prompt Tuning studies learned soft prompts, not ordinary temperature adjustment or manually written instructions.

The practical consequence is simple: you can change a model’s output behaviour through sampling without teaching it new facts or permanently changing what it has learned.

A compact method for sensible tuning

When a response disappoints you, diagnose the failure before moving a slider.

If the answer lacks information, supply the missing source or context. If it misunderstands the task, improve the instructions or examples. If the format breaks, use schema constraints and validation. If several acceptable outputs exist but you need more or less variety, sampling settings become especially relevant.

Use this sequence:

  1. Define success. Write down what counts as correct, useful and unacceptable.
  2. Fix the inputs. Keep the model, prompt, source material and output requirements stable.
  3. Measure the default. Do not optimise away from a baseline you have never tested.
  4. Adjust one control. Usually begin with temperature while leaving top-p unchanged.
  5. Compare multiple outputs. Include representative and difficult inputs.
  6. Choose the simplest adequate configuration. Extra knobs need demonstrated benefits.
  7. Recheck after changes. Model updates and new task types can invalidate old conclusions.

The central idea is that sampling manages variation within the model’s available continuations. It does not substitute for a good brief, relevant evidence or independent checks.

Once you understand that boundary, temperature and top-p become practical tools rather than mysterious personality settings.

FAQ

What is the difference between temperature and top-p?

Temperature changes the relative weights of candidate tokens. Lower values favour leading candidates more strongly; higher values flatten the distribution. Top-p keeps a shortlist whose cumulative probability reaches a chosen threshold, then samples within that shortlist. One reshapes the distribution; the other truncates it.

Should I change temperature or top-p?

Start by changing one, usually temperature, while leaving the other at its default. This makes comparisons easier to interpret. Try top-p separately if you have a specific reason to restrict the low-probability tail. Tune both only when you can evaluate their interaction systematically.

Does temperature zero stop hallucinations?

No. It usually reduces sampling variability, but the highest-probability continuation can still be false. A deterministic error remains an error. Relevant source material, clear instructions about missing information, reliable tools and verification matter more for factual accuracy than temperature alone.

Does top-p of 0.9 keep 90% of the vocabulary?

No. It keeps enough leading tokens to cover at least roughly 90% of the probability mass, subject to implementation details. That could mean very few tokens or a large number. The retained vocabulary size changes with the distribution at each generation step.

Why do I still get different answers with the same settings?

Sampling itself produces variation unless its random state is controlled. Even with low temperature or a supported seed, hosted systems may not promise exact reproducibility. Model versions, numerical behaviour and serving infrastructure can also affect results. Check the provider’s guarantees rather than assuming them.

Can a prompt set the temperature?

Writing “use temperature 0.2” in a normal prompt does not directly set an API parameter unless the surrounding application explicitly interprets that instruction. Asking for restrained wording can influence the response through context, but it is not the same operation as changing the sampler.

Is higher temperature better for brainstorming?

Sometimes, but only within limits. It can increase variation and unusual associations, while also increasing irrelevant or incoherent output. Start with a clear creative brief, request distinct categories of ideas, and compare several results. Raise temperature only if the added variety improves your actual selection.

What settings should I use if my chat app exposes none?

Focus on what you can control: instructions, examples, evidence, requested format and review criteria. Ask for several alternatives when you need variety, and request source-grounded extraction when you need precision. Many useful workflows do not require direct access to sampling settings.

Does a higher temperature make answers longer or more expensive?

Not directly. Temperature is not a length or billing control. It may indirectly change wording, repetition or when generation ends, which can affect token use. Output limits, requested scope, model pricing and generating multiple candidates are more direct influences on cost.

What is the single most useful rule to remember?

Do not ask sampling settings to solve the wrong problem. Use them to manage variation. Use better information to address missing knowledge, clearer instructions to address ambiguity, schemas to address formatting, and tests or evidence to assess correctness.

Sources

About the author

Editorial team · Editorial team

Author identity has not been supplied yet. Replace this record with the real writer's name, background and verifiable experience before publishing anything on the public site.

Full profile

Spotted an error? Report a correction.