Skip to content

Chain-of-Thought Prompting: When Step-by-Step Reasoning Helps

Chain-of-thought prompting can help AI tackle multi-step problems, but its real value comes from useful decomposition, visible checks and knowing when a simpler prompt is better.

By Editorial teamPublished 24 min read

Key takeaways

  • Use step-by-step prompting for tasks with dependent calculations, rules or constraints, not automatically for every request.
  • Ask for concise, checkable solution steps rather than a transcript of the model’s private reasoning.
  • A convincing explanation is not evidence that an answer is correct; verify calculations, sources and assumptions separately.
  • Modern reasoning models may need clear goals and constraints more than instructions to think step by step.
  • Compare prompts on representative tasks and measure correctness, cost and review effort before adopting a technique.
On this page
  1. Chain-of-Thought Prompting: When Step-by-Step Reasoning Helps
  2. What chain-of-thought prompting means
  3. Why intermediate steps can help
  4. When step-by-step prompting is worth trying
  5. When it adds little—or makes things worse
  6. Modern reasoning models change the prompting advice
  7. A reusable pattern: answer, support, check
  8. Worked example: comparing print quotes
  9. Worked example: applying a policy without inventing exceptions
  10. Worked example: a plan that can be checked
  11. Teaching the pattern with worked examples
  12. Why a convincing explanation can still be wrong
  13. Verification that adds more than reassurance
  14. Chain of thought versus prompt chaining
  15. Three exercises to build practical judgement
  16. How to evaluate whether the technique helps
  17. Common mistakes and how to fix them
  18. A practical decision guide
  19. FAQ

Chain-of-Thought Prompting: When Step-by-Step Reasoning Helps

Ask an AI assistant a complicated question and it may jump straight to an answer. Sometimes that is efficient. Sometimes it skips a condition, combines incompatible figures or confidently solves the wrong problem.

Chain-of-thought prompting emerged as a way to improve performance on tasks that benefit from intermediate steps. Instead of requesting only a conclusion, the prompt encourages a worked solution: identify relevant information, apply rules, perform calculations and connect the results.

The useful idea is not that longer answers are smarter. It is that some problems become easier when their dependencies are made explicit.

There is also an important limit. An explanation produced by a model is not a dependable transcript of its internal computation. It can be incomplete, misleading or constructed to support an answer that is already wrong. Modern reasoning models may also perform substantial internal reasoning without needing a prompt that asks them to “think step by step”.

For everyday use, the goal is therefore practical: obtain an answer with enough visible structure to check it, without demanding pages of speculative reasoning.

This article explains when that approach helps, how to design it and how to test whether it improves your results.

What chain-of-thought prompting means

The original technique

In its influential research form, chain-of-thought prompting supplied examples that paired questions with intermediate solution steps and final answers. The model then followed that pattern on a new problem.

A simplified example looks like this:

Question: A shelf holds 8 boxes. Each box contains 6 notebooks.
How many notebooks are there?

Worked solution: 8 boxes × 6 notebooks per box = 48 notebooks.
Answer: 48 notebooks.

Question: A cupboard holds 7 packs. Each pack contains 9 folders.
How many folders are there?

The examples demonstrate more than an output format. They show a useful operation between the input and the answer.

The paper Chain-of-Thought Prompting Elicits Reasoning in Large Language Models investigated this approach on arithmetic, commonsense and symbolic reasoning tasks. Its findings established an important prompting technique, but they should not be interpreted as proof that every model or task benefits equally.

The shorter, zero-shot version

A related approach adds a short instruction without supplying worked examples:

Solve the problem step by step, then give the final answer.

This is often called zero-shot chain-of-thought prompting. The paper Large Language Models are Zero-Shot Reasoners examined how such instructions could improve performance on particular reasoning benchmarks.

In ordinary conversation, people now use “chain of thought” loosely for almost any request involving steps. That broad usage hides useful distinctions.

You might be asking for:

  • A calculation with intermediate figures.
  • A short explanation of a recommendation.
  • A plan before an action.
  • A sequence of separate tasks.
  • The model’s supposedly complete internal reasoning.

These are not interchangeable. Most users need the first three, not the last.

Prefer a checkable solution to a thought transcript

A strong practical prompt specifies the evidence you need:

Give the answer, the key calculation steps and one independent check.
State any assumption that could change the result.
Keep the explanation concise.

This asks for a useful public explanation. It does not require access to private internal reasoning, which may not be available and would not automatically provide a faithful account of how the answer was produced.

Think of the output as a worked solution prepared for a reader, not a recording of a mind.

Why intermediate steps can help

Some tasks contain dependencies: you cannot calculate the final figure correctly until you have established an earlier one.

Consider estimating an event budget. You may need to determine attendance, calculate catering quantities, apply tax to eligible items, add fixed costs and compare the result with a spending limit. An answer that skips directly to “within budget” gives you little opportunity to notice a mistake.

A structured solution offers three practical benefits.

It separates different operations

A model asked to interpret a policy and calculate a reimbursement is doing at least two jobs. It must first decide which rule applies, then apply that rule numerically.

Separating these operations can reveal where the answer goes wrong. Perhaps the arithmetic is correct but the wrong reimbursement band was selected. Without intermediate outputs, the distinction is difficult to see.

It makes intermediate information available

In text-generating models, earlier generated text becomes part of the context used to generate later text. Writing a useful intermediate result can therefore support subsequent steps.

This is a functional explanation, not a claim that the system reasons in the same way as a person. Models can also anchor on an incorrect intermediate statement and carry it forward consistently.

Visible steps help only when the steps themselves are appropriate.

It gives the user inspection points

Even when a prompt does not improve the model’s final-answer accuracy, it may make review easier.

You can check whether:

  • The quantities came from the supplied data.
  • A discount was applied before or after tax.
  • A deadline counted working days or calendar days.
  • A recommendation respects all hard constraints.

However, an explanation is useful only if the reviewer actually inspects those points. A long solution skimmed for reassuring language can increase confidence without improving reliability.

The best intermediate outputs are compact and concrete: an equation, a rule identifier, a source quotation or a constraint table.

When step-by-step prompting is worth trying

A simple test is to ask: could an early mistake change several later decisions? If so, decomposition is likely to be useful.

Multi-stage arithmetic

Budgeting, unit conversions, pricing comparisons and percentage changes often benefit from explicit operations.

For example, “increase by 20%, then decrease by 20%” does not return a quantity to its original value. Written calculations expose the changing base.

Use prompts that name the operations and units, rather than merely asking the model to be careful.

Rules with exceptions

Policies, eligibility conditions and scheduling rules often combine general provisions with exceptions.

A useful response identifies the applicable rule, explains why an exception does or does not apply and then gives the decision. This can make missing information obvious.

For legal, medical or financial decisions, that structure supports review; it does not replace qualified judgement or authoritative sources.

Constraint-heavy planning

A timetable may need to satisfy availability, room capacity, equipment requirements and travel time simultaneously.

Asking for a plan plus a constraint check is more useful than asking for an elaborate narrative about planning. The check turns broad intentions into inspectable claims.

Debugging and diagnosis

For a software fault, a helpful structure is:

  1. State the observed failure.
  2. Identify a small number of plausible causes.
  3. Propose a test that distinguishes between them.
  4. Recommend the smallest supported change.

This discourages premature fixes. It also avoids treating the first plausible explanation as proven.

Comparisons with explicit criteria

Choosing between suppliers or software packages becomes clearer when requirements, trade-offs and uncertainties are separated.

Still, a scoring table can conceal subjective assumptions. Ask the model to distinguish hard requirements from preferences, and to identify which conclusions depend on chosen weights.

Step-by-step prompting is most valuable when it exposes the decision structure, not when it decorates an unsupported recommendation.

When it adds little—or makes things worse

Not every task needs visible reasoning.

Straightforward extraction

If the task is to extract invoice numbers from supplied text, a worked explanation may introduce unnecessary text or formatting errors.

Prefer a direct request:

Extract every invoice number from the text below.
Return one number per line, preserving the original order.
Do not infer missing numbers.

The challenge here is faithful extraction, not a multi-stage argument.

Simple transformations

Changing a sentence from present to past tense, alphabetising a short list or converting a date format usually does not need a step-by-step solution.

A clear instruction and a few formatting examples may be enough.

Creative generation

For naming a newsletter or drafting dialogue, prolonged analytical output can consume space that would be better spent producing alternatives.

Constraints still help: audience, tone, length and exclusions. But those are not the same as chain-of-thought prompting.

Questions requiring missing facts

Reasoning cannot establish a company’s current prices, a recent legal change or the contents of a document the model has not received.

A longer explanation can simply organise invented facts more persuasively. Our guide to why AI hallucinates explains why fluency and factual reliability must be evaluated separately.

When the bottleneck is missing evidence, supply evidence or use an appropriate information source. Do not substitute “think harder” for “look it up”.

Time-sensitive interactions

A lengthy solution may be inappropriate when the user needs a quick label, routing decision or short operational response.

Extra generated text can increase latency, token usage and review effort. Depending on the system, internal reasoning may also affect cost and speed independently of the visible answer length.

Choose the smallest amount of structure that solves the actual problem.

Modern reasoning models change the prompting advice

Early chain-of-thought research focused on prompting models to produce reasoning patterns they might not otherwise use reliably. Many contemporary systems are explicitly trained to handle reasoning tasks, sometimes using internal reasoning processes that are not shown to users.

That changes the practical question. Instead of asking, “How do I make it reason?”, ask, “What information, constraints and checks does it need?”

Start with a clear task

For a reasoning-focused model, this may be enough:

Compare these three delivery options.

Hard requirements:
- Arrival by Friday.
- Total cost below £300.
- Tracking included.

Use only the supplied information.
Recommend a qualifying option, or say none qualifies.
Include a brief justification and flag missing information.

There is no need to prescribe every mental operation. Excessively detailed instructions can constrain the model to an unsuitable method.

Separate reasoning effort from explanation length

A short answer can follow substantial internal processing. A long answer can contain shallow or erroneous reasoning.

Where a product exposes a reasoning-effort setting, that is distinct from requesting a verbose explanation. Available controls differ between models and interfaces, so consult the documentation for the system you actually use.

Evaluate the resulting answer rather than assuming that more visible text means more computation.

Treat prompting advice as model-specific

A technique that helped one model version may be neutral or harmful on another. Fine-tuning, instruction following and reasoning training can all change how a model responds.

Our introduction to pre-training, fine-tuning and RLHF provides useful background on why behaviour differs across systems.

For a recurring workflow, retest after changing the model. A prompt is part of a working configuration, not a permanent formula.

A reusable pattern: answer, support, check

A practical alternative to unrestricted “show all your reasoning” prompts is to request four things:

  1. The answer or decision.
  2. The essential supporting steps.
  3. Important assumptions or missing inputs.
  4. A concrete verification check.

Here is a reusable template:

Task:
[Describe the result you need.]

Inputs:
[Provide the relevant facts or documents.]

Constraints:
[List hard requirements and exclusions.]

Response:
- Give the answer clearly.
- Show only the key calculations, rules or evidence needed to check it.
- State assumptions that could change the answer.
- Include one concrete verification check.
- If a required input is missing, identify it rather than inventing it.

This template works because it specifies useful output artefacts. It does not rely on a vague instruction to be intelligent, rigorous or expert.

Adapt the support section to the task:

TaskUseful visible support
Budget calculationFormula, intermediate amounts, units
Policy decisionApplicable clause and relevant facts
SchedulingTimetable and constraint checks
Document comparisonDifferences tied to quoted passages
DebuggingReproduction steps, test and observed result
RecommendationRequirements, trade-offs and uncertainties

The “verification check” should be more specific than “confirm the answer is correct”. Ask for a reverse calculation, a boundary test, a matching source passage or an executable test.

For machine-readable workflows, define fields rather than accepting a free-form essay. The guide to prompting for structured output explains how schemas can make outputs easier to validate.

A good structure cannot guarantee good content. It can, however, make weak content easier to detect.

Worked example: comparing print quotes

Suppose you need 240 booklets for an event. Two suppliers quote different pricing structures.

  • Supplier A charges £1.80 per booklet and £24 delivery.
  • Supplier B charges £1.65 per booklet and £45 setup.
  • Supplier B gives a 10% discount on booklet charges only for orders of at least 200.
  • Both totals attract 20% VAT.
  • Both suppliers meet the deadline and required quality.

A bare prompt might ask, “Which supplier is cheaper?” A better prompt makes the calculation inspectable.

Compare the two suppliers for 240 booklets using the details below.

Show each supplier's:
1. Booklet charge after any eligible discount.
2. Total before VAT.
3. Total including VAT.

Assume all listed charges attract 20% VAT.
Round final monetary totals to two decimal places.
Recommend the cheaper supplier and state the saving.

Check the saving by comparing the pre-VAT totals as well.

A correct worked solution

For Supplier A:

  • Booklet charge: 240 × £1.80 = £432.00.
  • Before VAT: £432.00 + £24.00 = £456.00.
  • Including VAT: £456.00 × 1.20 = £547.20.

For Supplier B:

  • Initial booklet charge: 240 × £1.65 = £396.00.
  • Discount: £396.00 × 0.10 = £39.60.
  • Discounted booklet charge: £356.40.
  • Before VAT: £356.40 + £45.00 = £401.40.
  • Including VAT: £401.40 × 1.20 = £481.68.

Supplier B is cheaper by £65.52.

The independent arithmetic relationship is:

Pre-VAT difference = £456.00 - £401.40 = £54.60
VAT-inclusive difference = £54.60 × 1.20 = £65.52

What the steps reveal

The important reasoning is not the multiplication alone. It is applying the discount to the correct charge and checking eligibility.

A plausible wrong solution might discount both the booklet charge and the setup fee. Another might compare one supplier’s VAT-inclusive amount with the other supplier’s pre-VAT amount.

The visible categories expose these mistakes quickly.

Notice also what the prompt does not ask for: a long discussion of purchasing strategy. Since deadline and quality are already equal, the task is a defined cost comparison.

Where tools fit

For actual purchasing, verify the arithmetic with a calculator or spreadsheet and confirm the quote terms with the suppliers. The reverse check is useful, but it could still repeat a mistaken tax assumption shared by both calculations.

Language models are valuable for translating prose into a calculation structure. Deterministic tools are usually better for executing arithmetic reliably.

Worked example: applying a policy without inventing exceptions

Consider this fictional training policy:

Permanent employees may claim up to £600 per calendar year for approved courses. Employees in their probation period need written manager approval before booking. Contractors are not eligible. Travel costs are excluded.

An employee asks:

I am a permanent employee in probation. I have already claimed £180 this year. A course costs £450, and travel will cost £35. I have not booked it yet. What can I claim?

A weak response might say “£450, because it is under £600”. That ignores previous claims, probation approval and the exclusion of travel.

Use a rule-focused prompt:

Apply only the supplied training policy.

For each relevant condition:
- Identify the policy rule.
- Match it to a stated fact.
- Flag anything not yet established.

Then give a brief conditional answer.
Do not invent exceptions or assume approval has been granted.

A careful answer

ConditionStated factImplication
Employment statusPermanent employeeEligible category
Annual allowance£600 cap; £180 already claimed£420 remains
ProbationEmployee is in probationWritten manager approval required before booking
Course approvalNot establishedApproval must be confirmed
Travel£35 travel costExcluded

The answer should distinguish eligibility from authorisation:

Subject to the required approvals, the maximum course reimbursement appears to be £420, the remaining annual allowance. Travel is excluded. Obtain written manager approval before booking and confirm that the course is approved. The policy does not establish approval from the facts provided.

Why the conditional language matters

The arithmetic is simple: £600 − £180 = £420. The harder part is avoiding a stronger conclusion than the policy supports.

The model should not say the claim “will be paid”. It does not know whether the course is approved or whether there are additional administrative requirements outside the supplied policy.

If the user asks whether partial reimbursement is administratively permitted, the excerpt may not settle that question explicitly. A careful answer can identify £420 as the cap while recommending confirmation of the booking and claims process.

Step-by-step prompting is useful here because it creates a rule-to-fact mapping. It is not useful if it encourages the model to fill policy gaps with plausible workplace customs.

Worked example: a plan that can be checked

Planning outputs often sound reasonable while violating a small but decisive constraint.

Imagine three training sessions:

  • Session A lasts 60 minutes and must finish before Session B begins.
  • Session B lasts 90 minutes and must finish by 12:30.
  • Session C lasts 60 minutes.
  • One room is available from 09:00 to 13:00.
  • A 15-minute reset is required between sessions.
  • Session C’s trainer is unavailable before 11:00.

Ask for a schedule and a verification table:

Create a feasible schedule using only these constraints.

Return:
- A timetable.
- A check against every constraint.

If no schedule is feasible, explain the conflict briefly.
Do not relax a constraint without permission.

A feasible schedule

TimeActivity
09:00–10:00Session A
10:00–10:15Room reset
10:15–11:45Session B
11:45–12:00Room reset
12:00–13:00Session C

The checks are straightforward:

  • A lasts 60 minutes and finishes before B begins.
  • B lasts 90 minutes and finishes before 12:30.
  • C lasts 60 minutes and starts after 11:00.
  • Sessions do not overlap.
  • Both gaps provide the required reset time.
  • Everything fits within the room’s availability.

Test the edge of feasibility

Now change B’s deadline to 11:30.

The earliest possible B start is 10:15: A must occupy at least 60 minutes from 09:00, followed by a 15-minute reset. B therefore cannot finish before 11:45.

That single lower-bound calculation proves the revised schedule is infeasible under the stated rules.

This is more useful than generating several attractive timetables and hoping one works. It also illustrates a strong verification habit: test whether the earliest or latest possible times already rule out a solution.

Teaching the pattern with worked examples

If a model repeatedly chooses the wrong method, one or two carefully designed demonstrations may help more than additional instructions.

This is few-shot prompting: the prompt contains examples of the behaviour you want. Chain-of-thought demonstrations include concise intermediate solution steps, rather than only input–answer pairs.

For example:

Example:
Question: A price rises from £80 to £100. What is the percentage increase?
Calculation: (£100 - £80) / £80 × 100 = 25%.
Answer: 25%.

Example:
Question: A price falls from £100 to £80. What is the percentage decrease?
Calculation: (£100 - £80) / £100 × 100 = 20%.
Answer: 20%.

Now solve:
A subscription rises from £45 to £54. What is the percentage increase?
Give the calculation and answer.

These examples teach a specific distinction: the denominator is the starting price.

Choose examples for coverage, not decoration

Useful demonstrations vary the feature most likely to cause errors:

  • A discount that applies versus one that does not.
  • A rule with an exception versus the default case.
  • Enough information versus a missing required input.
  • A feasible plan versus an impossible one.

Avoid supplying five nearly identical easy examples. They may teach a surface pattern without testing the important distinction.

Keep examples correct and consistent

A wrong worked example is especially damaging because it demonstrates both an incorrect method and an incorrect answer.

Check examples manually, use consistent rounding conventions and avoid unnecessary narrative. For a deeper design guide, see few-shot prompting: designing examples that teach the model.

The related research on least-to-most prompting explored decomposing harder problems into simpler subproblems and solving them progressively. The practical lesson is useful even without adopting a named technique: teach the model how to separate the difficult dependency, not merely how to imitate the final wording.

Why a convincing explanation can still be wrong

The central danger is confusing plausibility with verification.

A model can produce an answer and a polished explanation that fit together while both are wrong. It can omit a decisive fact, quietly introduce an assumption or select evidence that supports its initial conclusion.

Research on unfaithful explanations in chain-of-thought prompting showed that generated explanations may fail to disclose factors that influenced model answers. This is one reason not to treat visible reasoning as a transparent account of the model’s internal process.

Three things to evaluate separately

Answer correctness: Does the conclusion match the facts, calculation or reference answer?

Explanation validity: Do the stated steps logically support the conclusion?

Explanation faithfulness: Does the explanation accurately reflect the process that produced the answer?

These are different properties. A correct answer can have a bad explanation. A coherent explanation can support a false premise. A useful worked solution need not reveal the internal process that generated it.

For most practical workflows, answer correctness and explanation validity are the properties you can test most directly.

The paper Faithful Chain-of-Thought Reasoning investigated approaches that connect generated reasoning representations with deterministic execution.

The everyday equivalent is simple: ask the model to formulate a calculation, query or test, then run it in an appropriate tool.

For example:

  • Execute generated arithmetic in a calculator.
  • Run a proposed code fix against tests.
  • Validate a schedule with explicit time constraints.
  • Check policy claims against the supplied clauses.

This does not eliminate error. The model can still formulate the wrong problem. But it reduces reliance on prose that merely asserts the result is correct.

Verification that adds more than reassurance

“Check your answer” is better than nothing, but it does not specify what a check should do.

A useful check is capable of discovering a particular class of error.

Reverse the operation

For a percentage calculation, reconstruct the original quantity from the result.

If a £45 subscription rises by 20%, the new price should be £45 × 1.20 = £54. This provides a direct check against the proposed percentage increase.

Check units and bounds

If a journey covers 120 kilometres in two hours, an average speed of 240 kilometres per hour should trigger concern.

For a budget, ask whether the result can be smaller than the fixed costs. For a proportion, check whether it falls within the valid range for the definition being used.

Simple bounds often catch errors faster than rereading a long derivation.

Use a different method

Compare itemised totals with an aggregated calculation. Test code with independently written cases. For a timetable, inspect each constraint instead of asking for another narrative review.

Different wording alone does not make a check independent.

Sample alternatives cautiously

The paper Self-Consistency Improves Chain of Thought Reasoning in Language Models explored sampling multiple reasoning paths and selecting answers through agreement.

This can help on suitable tasks, especially where answers can be compared unambiguously. But repeated outputs can share the same misconception. Agreement is a signal, not proof.

Similarly, asking a model to revise its answer without new evidence is not guaranteed to improve it. Large Language Models Cannot Self-Correct Reasoning Yet examined limitations of intrinsic self-correction in the settings studied. Its lesson is not that checking never works, but that unsupported self-review should not be treated as a reliable verifier.

For additional patterns, see self-consistency and verification prompts.

Chain of thought versus prompt chaining

These terms sound similar but describe different workflow choices.

Chain-of-thought prompting encourages intermediate solution steps within a response.

Prompt chaining separates a task across multiple calls, passing selected outputs from one stage into the next.

For example, a document review workflow might use:

  1. One call to extract requirements.
  2. A validation step to confirm the extraction.
  3. Another call to compare a proposal with those requirements.
  4. A final call to draft a concise report.

The separation creates control points. You can stop if extraction fails rather than allowing the error to spread into a polished report.

When separate calls are preferable

Use prompt chaining when:

  • Different stages need different sources or tools.
  • An intermediate result requires human approval.
  • The task is too large to inspect as one response.
  • You need to retry one stage without repeating everything.
  • A downstream stage should receive only validated information.

However, a chain also introduces overhead. You must manage state, preserve relevant context and prevent errors in one stage from becoming unquestioned inputs to the next.

For a one-off calculation, a single well-structured response is usually simpler. For recurring document processing, explicit stages may be easier to maintain and audit.

Our guide to prompt chaining covers that workflow distinction in more detail.

Neither approach removes the need for verification. Breaking an unsupported claim into three calls does not make it supported.

Three exercises to build practical judgement

Exercise 1: compare direct and structured prompts

Use the print-quote problem from earlier.

First, request only the cheaper supplier and the saving. Then start a fresh conversation and request the structured calculation.

Record:

  • Whether the final answer is correct.
  • Whether the discount was applied to the right items.
  • Whether VAT treatment is consistent.
  • How long the answer takes to review.

Next, reduce the order to 180 booklets so that Supplier B no longer qualifies for the discount.

The new totals should be:

Supplier A: (180 × £1.80 + £24) × 1.20 = £417.60
Supplier B: (180 × £1.65 + £45) × 1.20 = £410.40
Saving with Supplier B: £7.20

This variation tests whether the model applies the condition rather than copying the earlier calculation.

Exercise 2: introduce missing information

Use the training-policy example, but remove the amount already claimed.

Ask the model for the exact reimbursement.

A good response should not invent a remaining allowance. It should request the previous claims total or give a conditional expression:

Maximum reimbursement is limited by:
- The approved course cost.
- The remaining annual allowance: £600 minus previous eligible claims.
- The required approvals.

An exact figure needs the amount already claimed this calendar year.

The lesson is that sometimes the correct next step is a question, not another calculation.

Exercise 3: test a hard constraint

Use the scheduling example and change Session B’s deadline to 11:30.

Ask for a timetable and a constraint check. Then inspect whether the model:

  • Recognises that no feasible schedule exists.
  • Shows the earliest possible completion time of 11:45.
  • Avoids shortening sessions or removing reset time.
  • Labels any proposed relaxation as a change requiring permission.

This exercise tests an important behaviour: refusing to manufacture feasibility.

Keep the examples and expected results in a small test file. They become reusable checks whenever you change your prompt or model.

How to evaluate whether the technique helps

One impressive answer is not enough evidence for a recurring workflow.

Build a small collection of representative tasks with answers or evaluation criteria you can defend. Include routine cases, boundary cases, exceptions, missing inputs and impossible requests.

Then compare at least two prompt versions:

  • A clear direct prompt.
  • The same prompt with specific intermediate outputs and checks.

If you also test a generic “think step by step” instruction, keep it as a separate version. Otherwise, you cannot tell which change helped.

Keep the comparison fair

Use the same model, supplied information and relevant settings. Start fresh conversations unless conversation history is deliberately part of the task.

Do not let one version see the reference answer or benefit from corrections made during another run.

For systems with variable outputs, repeat some cases. A single success or failure may not represent typical behaviour.

Score more than correctness

A practical evaluation sheet can include:

DimensionQuestion
CorrectnessIs the answer right?
Constraint adherenceDid it obey all hard requirements?
Evidence handlingAre claims supported by the supplied material?
Uncertainty handlingDid it flag missing information appropriately?
Review effortHow difficult was it to check?
EfficiencyWere latency and output length acceptable?

Do not award points merely because an answer contains steps. The structure is a means, not the outcome.

Keep some cases aside while developing the prompt, then test on them later. Otherwise, you may optimise for familiar examples rather than improve general performance.

Finally, choose the simpler approach when results are comparable. Extra reasoning text is not free, and an easy-to-maintain prompt has value.

Common mistakes and how to fix them

Asking for exhaustive reasoning

A request for every possible thought can produce excessive, repetitive material without improving the answer.

Fix: Ask for the key calculations, applicable rules, assumptions and a specific check.

Confusing task instructions with verification

“Be accurate” states a preference. It does not describe a test.

Fix: Name a check that could fail: recompute the total, compare against each constraint or quote the supporting sentence.

Giving away the desired answer

A prompt such as “Explain why Supplier A is the best choice” encourages a justification for a predetermined conclusion.

Fix: Ask which supplier meets the criteria, allowing “neither” where appropriate.

Leaving assumptions invisible

Unstated tax treatment, time zones, rounding rules and definitions can change the answer.

Fix: Supply consequential assumptions or ask the model to flag them before giving a definitive result.

Letting examples teach the wrong shortcut

If every demonstration has an eligible discount, the model may overlook an ineligible case.

Fix: Include contrastive examples that exercise the actual rule boundary.

Treating a second model as an oracle

Another model may catch errors, but it may also share the same weaknesses or accept the first answer’s framing.

Fix: Give the checker the original problem and objective criteria. Where possible, use a tool or authoritative source rather than another opinion.

Prescribing an unsuitable method

An overly rigid prompt can force the model through steps that do not fit the problem.

Fix: Specify required evidence and constraints while leaving room to choose an appropriate solution method.

Using reasoning instead of evidence

A model cannot derive the latest policy wording from general knowledge.

Fix: Provide the document or retrieve it. Retrieval-augmented generation addresses the evidence-supply problem; step-by-step prompting addresses how supplied information is used.

A practical decision guide

Before adding step-by-step instructions, identify the bottleneck.

If the task lacks information, obtain the information. If it lacks a clear goal, clarify the goal. If it requires exact arithmetic, use a calculator. If it involves dependent decisions, request useful intermediate outputs.

A sensible progression is:

  1. Write a clear direct prompt.
  2. Test it on representative cases.
  3. Add concise solution steps where failures involve dependencies.
  4. Add examples where the method is repeatedly misunderstood.
  5. Add external checks where correctness matters.
  6. Split the workflow into stages only when control points justify the overhead.

Keep the final output proportionate to the reader’s needs. A decision with three supporting bullets may be more useful than a page of analysis.

The central principle is simple: ask for what makes an answer checkable, not what makes it look thoughtful.

Chain-of-thought prompting is a useful technique when it encourages the right decomposition. It is not a guarantee of truth, a substitute for evidence or a reason to expose every possible reasoning step.

FAQ

Is chain-of-thought prompting just saying “think step by step”?

That is a common zero-shot form, but the original technique used worked examples containing intermediate steps and answers. In practice, specifying useful outputs—calculations, rule applications or constraint checks—is often more actionable than the generic phrase.

Does it always improve accuracy?

No. Results depend on the model, task, prompt and evaluation conditions. It can help with multi-stage problems, but may add little to extraction or simple transformations. Extra steps can also introduce mistakes. Test it against a clear direct-prompt baseline.

Can I see the model’s actual internal reasoning?

Not reliably. Some systems keep internal reasoning private, and a generated explanation should not be treated as a faithful transcript of internal computation. Ask instead for a concise justification, supporting evidence and checks you can inspect.

Should I use it with a modern reasoning model?

Start with a clear task, complete inputs and explicit constraints. Many reasoning-focused models do not need a generic instruction to think step by step. Request a worked solution when it helps you review the answer, and test whether additional prompting improves results.

Is a longer answer usually more reliable?

No. Length and correctness are different properties. A long answer can repeat an incorrect assumption convincingly. Prefer the shortest explanation that exposes the decisive calculation, evidence or rule, together with an appropriate verification step.

How many worked examples should I include?

Begin with one or two strong examples covering the main difficulty. Add more only when testing reveals a distinct failure pattern. Diversity matters more than volume: include exceptions, missing information and boundary cases rather than several near-duplicates.

What should I do when several AI answers agree?

Treat agreement as useful but insufficient evidence. The answers may share the same mistaken assumption. Check the underlying facts, use a different calculation method or test the result with a deterministic tool where possible.

What is the safest everyday prompt pattern?

Ask for the answer, the key supporting steps, consequential assumptions and one concrete check. Tell the model not to invent missing inputs. For high-stakes decisions, add authoritative sources and appropriate human review; prompting alone is not a sufficient safeguard.

Sources

About the author

Editorial team · Editorial team

Author identity has not been supplied yet. Replace this record with the real writer's name, background and verifiable experience before publishing anything on the public site.

Full profile

Spotted an error? Report a correction.