Tree of Thoughts and Advanced Reasoning Prompt Patterns
Tree of Thoughts turns difficult AI tasks into a controlled search through alternatives, with explicit checks, pruning and stopping rules.
Key takeaways
- Use Tree of Thoughts when a task has meaningful alternatives and decisions that may need revisiting.
- Separate candidate generation from evaluation, and reject hard-constraint failures before ranking options.
- Ask for concise decision records, evidence and checks rather than private internal reasoning.
- Bound the search by branch count, depth, cost and clear stopping rules.
- Verify important claims and calculations independently; agreement between model outputs is not proof.
On this page
- Why some tasks need alternatives, not a longer answer
- What Tree of Thoughts means
- When branching is worth the effort
- How the main reasoning patterns differ
- Design the evaluator before generating candidates
- Worked example: planning a workshop under constraints
- A reusable bounded-tree prompt
- Worked example: debugging without guessing a fix
- Verification that adds evidence
- Implementing a small search controller
- Control branching, context and cost
- Three step-by-step exercises
- Common mistakes and practical repairs
- A practical decision checklist
- FAQ
Why some tasks need alternatives, not a longer answer
Suppose you ask an AI assistant to plan a community workshop. It produces a convincing schedule, a budget and a promotional message. Only afterwards do you notice that its preferred venue is unavailable, the catering exceeds the budget, and its backup plan depends on the same venue.
The problem was not necessarily a lack of detail. The assistant committed to an approach before comparing alternatives or checking the conditions that could invalidate it.
Tree of Thoughts offers a useful way to organise this kind of work: generate distinct candidates, evaluate them, develop the promising ones, and return to an earlier decision when necessary. Instead of treating the first plausible answer as the destination, treat it as one candidate in a small search.
This article explains Tree of Thoughts and advanced reasoning prompt patterns as practical workflow designs. You will learn how to choose a pattern, define useful branches, apply checks, control costs and recognise when further prompting will not help.
The aim is not to extract an assistant’s private internal reasoning. It is to obtain inspectable work products: alternatives, assumptions, calculations, evidence references, test results and short decision summaries. Those are the things a person or a program can actually verify.
What Tree of Thoughts means
From a single answer to a search process
The research paper Tree of Thoughts: Deliberate Problem Solving with Large Language Models describes a framework in which a language model generates and evaluates intermediate candidates within a search process.
The important shift is from producing one continuation to managing alternatives. A search procedure can keep several candidates, assess their prospects, extend selected ones and abandon unpromising paths.
In everyday use, a “thought” is best understood as a task-relevant intermediate unit. Depending on the task, it might be:
- A proposed workshop format.
- A partial mathematical solution.
- A hypothesis about a software fault.
- An outline for a report.
- A proposed allocation of money or staff.
It need not be a paragraph of introspection. Often a short structured record is more useful.
The five moving parts
A practical Tree of Thoughts workflow has five components.
A state records the current partial solution. For a workshop plan, that includes the selected format, known costs, unresolved questions and remaining resources.
An expansion rule determines what choices to generate next. For example, produce three delivery formats, then two schedules for each surviving format.
An evaluator checks whether a candidate is valid and whether it deserves further work. Some checks are mechanical, such as whether costs exceed a budget. Others require judgement, such as whether beginners can follow the activities.
A selection rule decides which candidates survive. You might retain the best two valid options rather than develop every possibility.
A stopping rule limits the search. Stop when an acceptable solution passes verification, when a fixed number of rounds is complete, or when progress requires unavailable information.
Without these components, “explore a tree of thoughts” can become an invitation to write a long essay.
A prompt is not the same as a search implementation
There is a meaningful difference between asking one model response to compare three options and running a search controller that stores candidates, invokes evaluators and explicitly revisits branches.
Both can be useful. However, a single response only approximates some features of the research framework. It does not guarantee independent evaluation, systematic backtracking or exhaustive coverage.
Use precise names:
- Alternative comparison: one response compares several candidates.
- Bounded tree workflow: multiple stages expand, evaluate and prune candidates.
- Search implementation: software manages states, model calls, checks and selection.
This distinction matters when assessing reliability. A polished comparison table is not evidence that a systematic search occurred.
When branching is worth the effort
Good candidates for Tree of Thoughts
Branching is useful when three conditions hold.
First, there are genuinely different approaches. Choosing between an in-person clinic, a webinar and self-guided materials creates meaningful alternatives. Choosing three synonyms for the same headline usually does not.
Second, early decisions influence later possibilities. Selecting a venue affects capacity, travel arrangements and cost. Selecting a software architecture affects deployment, testing and maintenance.
Third, you can evaluate intermediate candidates. You need constraints, tests, evidence or at least a clear rubric. If every branch can only be labelled “interesting”, search will not become more reliable.
Common applications include constrained planning, debugging, mathematical puzzles, document architecture and selecting among implementation strategies.
Poor candidates for branching
Do not build a tree for a straightforward lookup, a simple transformation or a task whose method is already known.
For example, extracting dates from a supplied document usually needs an extraction schema and an omission check. Generating five extraction strategies may introduce unnecessary variation.
Branching also performs poorly when the central problem is missing information. If you do not know whether a venue is available, generating more elaborate venue plans will not establish availability.
Similarly, a safety-critical decision does not become trustworthy because several AI-generated branches agree. Medical, legal and financial questions may require authoritative evidence and qualified review.
A quick selection test
Before choosing a pattern, ask:
- What are the materially different choices?
- Which decisions would be expensive to reverse?
- What can eliminate a candidate early?
- What evidence would distinguish the survivors?
- What is the cheapest acceptable workflow?
If you cannot answer the first question, use a simpler prompt. If you cannot answer the third or fourth, improve the evaluation design before generating branches.
The right objective is not maximum reasoning effort. It is sufficient, verifiable work for the decision at hand.
How the main reasoning patterns differ
Chain-of-thought prompting: structure a single attempt
The original chain-of-thought prompting paper investigated how intermediate reasoning examples could improve performance on reasoning tasks.
In practice, users often benefit from requesting a clear method, relevant calculations and a concise explanation. But do not equate a long explanation with an accurate result, or assume that a displayed explanation is a faithful transcript of internal computation.
For many tasks, this is enough:
Solve the problem.
State the assumptions, show the calculations needed to check the answer,
and give a concise explanation of the result.
Our guide to Chain-of-Thought Prompting: When Step-by-Step Reasoning Helps explores where structured explanations are useful and where independent checking matters more.
Self-consistency: compare multiple attempts
Self-consistency samples multiple reasoning paths and aggregates their answers. It differs from a tree because attempts need not share partial states or branch from the same intermediate decision.
This can help when a problem has a clear final answer and individual attempts sometimes make different mistakes.
However, agreement is not proof. Several attempts may reproduce the same misconception, especially when they use the same model and source material.
For decisions with many acceptable outcomes, majority voting can be particularly misleading. The most frequent recommendation may reflect a familiar pattern rather than the user’s priorities.
Least-to-most: solve dependencies in order
Least-to-most prompting breaks a complex problem into simpler subproblems and solves them progressively.
Imagine calculating the cost of a training programme. You first establish participant numbers, then staffing requirements, then venue needs, then the total budget.
This is useful when later work depends on earlier results but the overall approach does not require competing strategies.
The main risk is error propagation. A wrong participant count can contaminate every later calculation. Check important intermediate results before passing them forward.
ReAct: obtain observations through tools
ReAct combines reasoning with actions and observations. A practical workflow might inspect a file, run a test, examine the result and choose the next action.
This matters because some uncertainty cannot be resolved through text generation alone.
A debugging assistant should often inspect logs or execute a controlled test rather than rank hypotheses indefinitely. A travel-planning assistant needs current availability, not confident estimates based on old information.
See ReAct Prompting: Combining Reasoning and Actions for the tool-use side of this approach.
Reflection: use feedback to improve a later attempt
Reflexion explores agents using feedback and stored verbal reflections to improve later attempts.
The practical lesson is narrower than “ask the model to criticise itself”. Useful reflection is connected to evidence: a failed test, a violated constraint, an evaluator’s finding or a user correction.
A productive revision instruction says:
The budget checker found that candidate B omitted equipment hire.
Revise the candidate to include that cost.
Then rerun every budget-related check.
This is much stronger than “think harder and make it better”.
Combining patterns without creating confusion
These patterns are complementary, but they serve different purposes.
Use decomposition to identify subproblems, branching to explore alternative approaches, tools to obtain observations, verification to check outputs, and reflection to repair specific failures.
Do not stack every pattern into one enormous instruction. Instead, choose a clear sequence with a defined output at each stage.
A useful default is:
Define constraints → generate alternatives → check feasibility
→ expand survivors → obtain missing evidence → verify → recommend
Each arrow should represent a genuine change in information or state.
Design the evaluator before generating candidates
Separate hard constraints from preferences
Hard constraints are pass-or-fail conditions. Preferences help rank candidates that have already passed.
For a workshop, hard constraints might include a maximum cost, a fixed date and step-free access. Preferences might include low preparation effort, opportunities for discussion and reusable teaching materials.
Never let a high preference score compensate for a hard failure.
A beautiful plan costing £900 does not become acceptable under a £600 ceiling because it scores highly for participant engagement.
Start with an explicit contract:
Hard constraints:
- Total cost must not exceed £600.
- Capacity must be at least 24.
- The activity must fit within a two-hour booking.
- Required access arrangements must be confirmed before booking.
Preferences:
- More supported hands-on practice.
- Less preparation time.
- Materials that can be reused.
Reject hard-constraint failures before applying preferences.
Mark missing evidence as unknown, not as a pass.
The final line is essential. Models often interpret “not disproven” as “acceptable”.
Prefer observable checks to impressionistic scores
“Feasibility: 8/10” gives little information unless the rubric defines feasibility.
A stronger evaluation identifies the claim and its supporting evidence:
| Check | Result | Basis |
|---|---|---|
| Capacity at least 24 | Pass | Supplied capacity: 30 |
| Total at most £600 | Pass | Listed items total £542 |
| Step-free access | Unknown | No confirmation supplied |
| Fits two-hour booking | Fail | Schedule requires 140 minutes |
This table tells you what to fix. A single overall score hides the failure.
Where judgement is unavoidable, use anchored categories. For preparation effort, “low” might mean adapting existing materials, “medium” creating one new activity, and “high” creating a complete new course.
The categories remain subjective, but at least they have operational meaning.
Define what would change the decision
Before ranking candidates, ask which uncertain facts could reverse the result.
Perhaps one option is best only if borrowed laptops are available. That dependency should appear beside the recommendation, not in a footnote.
A useful evaluator returns:
- Constraint status.
- Evidence quality.
- Main benefit.
- Main risk.
- Decision-changing unknown.
- Cheapest next check.
This shifts evaluation from literary criticism to decision support.
Worked example: planning a workshop under constraints
Establish the fictional brief
Consider a community organisation planning an introductory spreadsheet workshop. The following figures are supplied facts for this exercise, not market estimates.
The event must serve 24 adults. The spending limit is £600. The booking lasts two hours, including setup and clearing away. Participants need supervised practice and must not be required to own a laptop.
Available resources are:
| Resource | Supplied details |
|---|---|
| Community hall | £160; capacity 30; step-free; no computers |
| Library computer room | £240; 12 computers; capacity 24; step-free |
| Online meeting service | No extra cost; attendees need a suitable device |
| Tutor | £180 for the event |
| Assistant | £90 for the event |
| Laptop hire | £18 per laptop |
| Printed guide | £1 per participant |
| Refreshments | £2 per participant, optional |
For this exercise, both rooms are available on the required date. Treat the listed prices as inclusive and complete; a real booking would need to check taxes, delivery charges and cancellation terms.
The organisation prefers extensive practical work, but its hard requirement is supervised practice rather than one computer per person.
Stage one: generate distinct formats
Ask for materially different options, not variations in wording:
Using only the supplied resource table, propose three distinct
delivery formats.
For each format return:
- resource choices;
- device arrangement;
- itemised cost;
- hard-constraint status;
- one important limitation.
Do not invent discounts, donated equipment or participant-owned devices.
Do not build a detailed schedule yet.
An illustrative candidate set is:
A: Hall with one hired laptop per participant.
- Hall: £160.
- Tutor: £180.
- Assistant: £90.
- Laptop hire: 24 × £18 = £432.
- Printed guides: £24.
- Total: £886.
This exceeds the budget by £286. Reject it before developing a schedule.
B: Library room with paired practice.
- Library room: £240.
- Tutor: £180.
- Assistant: £90.
- Printed guides: £24.
- Refreshments: £48.
- Total: £582.
There are two participants per computer. This passes the stated resource and budget constraints, with £18 remaining.
C: Online workshop.
The staffing and material costs could fit the budget, but attendance would depend on suitable participant devices. The brief does not permit that assumption.
Mark C as failing the device requirement as currently specified. Do not silently treat device ownership as an established fact.
Stage two: prune and consider a nearby repair
Only B survives unchanged. Does that mean search is over?
Not necessarily. The hall branch failed because of one-to-one equipment hire. A bounded repair could test whether shared devices make it viable.
Hall with 12 hired laptops costs:
Hall £160
Tutor £180
Assistant £90
12 laptops × £18 £216
24 printed guides £24
Total £670
It still fails. Removing the assistant reduces the total to £580, but changes the support arrangement.
Call this repaired candidate A2. It is not automatically invalid: the brief did not require an assistant. However, the organisation must now compare a single tutor supporting 12 pairs in the hall with a tutor and assistant supporting 12 pairs in the library.
This is genuine backtracking. An earlier resource decision is changed, its consequences are recalculated, and the revised candidate is assessed.
Stage three: expand the strongest candidate
The library option has a clear support advantage under the supplied facts. Develop two schedules within that branch.
Expand candidate B into two schedules.
Schedule B1:
Emphasise guided paired practice.
Schedule B2:
Emphasise a longer demonstration followed by individual turns.
Requirements:
- Include setup and clearing away within 120 minutes.
- Allocate a keyboard turn to every participant.
- Preserve the £582 budget.
- List activity durations and the total.
- State the main instructional trade-off.
An illustrative B1 schedule is:
| Activity | Minutes |
|---|---|
| Setup and sign-in | 10 |
| Welcome and device orientation | 10 |
| Short demonstration | 15 |
| Paired exercise, first keyboard user | 25 |
| Paired exercise, second keyboard user | 25 |
| Questions and recap | 20 |
| Save work and clear room | 15 |
| Total | 120 |
B2 might allocate 30 minutes to demonstration and 35 minutes to paired practice. It could still fit the booking, but would give each participant less keyboard time.
Because the stated preference favours practical work, B1 is the stronger candidate. The evaluation rests on a visible allocation of time, not on an unexplained score.
Stage four: verify the recommendation
The final check should be separate from candidate generation.
Audit schedule B1 and its budget against the original brief.
Recalculate all totals.
Check participant capacity and device allocation.
Confirm that setup and clearing away are included.
Distinguish supplied facts from operational assumptions.
List any checks needed before a real booking.
Return findings, not a rewritten proposal.
The budget is £582 and the schedule totals 120 minutes. Twenty-four participants use 12 computers in pairs, with an explicit change of keyboard user.
Operational checks remain: spreadsheet software must work, accounts must permit saving files, and the room layout must support paired use. These are not reasons to fabricate a failure, but they are conditions to confirm before committing.
The recommendation should therefore be conditional and specific:
Choose the library paired-practice format with B1’s schedule, subject to checking software access and room layout. It meets the supplied constraints and retains both teaching staff. Keep the hall-without-assistant option as a weaker fallback.
That is a decision record someone can act on and audit.
A reusable bounded-tree prompt
A template for ordinary chat use
The following template works best when the task and evaluation criteria are already reasonably clear.
Help me solve the task below using a bounded comparison of alternatives.
TASK
[Describe the decision or problem.]
SUPPLIED FACTS
[Insert reliable inputs and identify their sources.]
HARD CONSTRAINTS
[List pass/fail requirements.]
PREFERENCES
[List ranking criteria in priority order.]
SEARCH LIMITS
- Generate at most 3 distinct initial candidates.
- Keep at most 2 candidates after feasibility checks.
- Expand each survivor into at most 2 variants.
- Allow one repair of a rejected candidate.
- Stop after the final verification pass.
PROCESS
1. Identify any missing fact that prevents useful evaluation.
2. Generate concise candidate records.
3. Reject hard-constraint failures.
4. Mark unsupported claims and unknowns explicitly.
5. Expand the strongest surviving candidates.
6. Compare them using the stated preferences.
7. Verify the recommended candidate against the original inputs.
OUTPUT
- Candidate comparison table.
- Rejection reasons.
- Final recommendation with a concise rationale.
- Verification results.
- Remaining uncertainty and the next action.
Do not provide private internal reasoning.
Provide only the decision records, evidence and checks needed
to inspect the result.
Treat the numerical limits as practical defaults, not scientifically optimal settings.
When to pause between stages
For consequential tasks, do not run everything in one response. Ask for initial candidates, review them, then approve the next stage.
This catches misunderstood constraints before they spread through the tree.
It also allows you to insert new facts: a confirmed price, a test result or a stakeholder’s priority.
Our guide to Prompt Chaining: Breaking Complex Tasks Into Reliable Steps explains how to make these hand-offs explicit.
A good hand-off contains the selected candidate, its verified facts, unresolved questions and the exact next task. It should not carry forward pages of abandoned speculation.
Worked example: debugging without guessing a fix
Start with competing explanations
Imagine a small reporting script sometimes produces duplicate customer rows after combining two tables.
A weak prompt asks:
My merge creates duplicate rows. Fix the code.
The assistant may immediately suggest removing duplicates. That could hide the symptom while discarding legitimate records.
A better approach explores causes before selecting a repair:
A reporting script produces more rows than expected after a merge.
Known facts:
- orders contains multiple orders per customer.
- customers should contain one row per customer_id.
- The intended output is one row per order.
Propose at most three testable hypotheses.
For each, give:
- the expected observation;
- a minimal diagnostic test;
- what result would weaken the hypothesis.
Do not propose a production change yet.
Useful hypotheses might include duplicate customer keys, inconsistent key formatting, and an incorrect choice of join columns.
Notice that the branches are explanations, not competing code patches.
Use a test to prune branches
A candidate diagnostic is:
duplicate_customers = customers[
customers.duplicated("customer_id", keep=False)
]
print(duplicate_customers.sort_values("customer_id"))
If duplicate keys appear, that supports one explanation. It does not prove that every duplicate is erroneous.
Two rows might represent historical customer versions. The real issue could be that the merge requires an effective-date condition rather than simple deduplication.
A further check can protect the intended relationship:
result = orders.merge(
customers,
on="customer_id",
how="left",
validate="many_to_one"
)
With pandas, validate="many_to_one" checks that the right-hand merge keys are unique. A failure becomes evidence about the data relationship.
These snippets assume that the tables and named columns exist. They are diagnostic examples, not a complete repair for every reporting system.
Branch on repairs only after understanding the data
Once duplicate customer records are confirmed, candidate repairs might be:
- Correct accidental duplicates upstream.
- Select the current customer record using a documented rule.
- Join each order to the customer version valid on its order date.
The evaluator now asks which repair matches the data’s meaning.
A command that simply keeps the first row may pass a row-count check while associating orders with the wrong customer details. Therefore verification must include representative record-level cases, not only totals.
This example shows why Tree of Thoughts and tool use work well together. Branching organises possible explanations; observations decide which deserve further attention.
Verification that adds evidence
Model agreement is a weak check
Asking the same assistant “Are you sure?” can produce reassurance rather than new information.
Generating another answer is sometimes useful, but its errors may be correlated with the first. Shared training, shared wording and shared context can lead to the same mistake.
The paper Large Language Models Cannot Self-Correct Reasoning Yet reports limitations of intrinsic self-correction in the settings it studied. It should not be read as a universal claim about every later model, but it supports an important caution: improvement is not guaranteed when revision has no external feedback.
Prefer checks that introduce independent information or a different failure-detection mechanism.
Match the check to the claim
Different claims need different verification methods.
| Claim type | Useful verification |
|---|---|
| Arithmetic total | Calculator or executable calculation |
| Code behaviour | Tests, including edge cases |
| Document quotation | Locate the passage in the source |
| Current availability | Authoritative current listing or direct confirmation |
| Constraint satisfaction | Checklist against the original brief |
| Subjective preference | Human review using explicit criteria |
A source citation is not enough by itself. Confirm that the source actually supports the particular statement.
Similarly, a passing test suite shows that tested cases passed. It does not establish that every possible input will behave correctly.
Use an evidence-aware audit prompt
Audit this proposed answer against the supplied evidence.
For each material claim, label it:
- supported by supplied evidence;
- derived and independently checkable;
- unsupported;
- contradicted;
- judgement rather than fact.
For derived claims, state the check to perform.
For unsupported claims, identify the evidence needed.
Do not repair the answer silently.
Separating the audit from the repair preserves visibility. Otherwise the assistant may change the proposal without showing what failed.
For more checking patterns, see Self-Consistency and Verification Prompts for More Reliable Answers.
Implementing a small search controller
Store candidates as data
A repeatable workflow benefits from structured candidate records.
{
"id": "B1",
"parent_id": "B",
"summary": "Library workshop with paired keyboard turns",
"cost_gbp": 582,
"duration_minutes": 120,
"constraint_status": {
"budget": "pass",
"capacity": "pass",
"device_access": "pass"
},
"unknowns": [
"Spreadsheet software access",
"Room layout for paired use"
],
"evidence_refs": [
"brief.resources.library",
"brief.resources.staffing"
],
"status": "retain"
}
A schema makes missing fields visible and allows deterministic checks. It does not make the values true.
For example, a program should recompute the budget from line items rather than trust cost_gbp merely because it is valid JSON.
See Prompting for Structured Output: JSON, Tables and Schemas for practical ways to enforce output shape.
Separate generator, checker and selector
A minimal controller might look like this:
frontier = [initial_state]
completed = []
for depth in allowed_depths:
children = expand(frontier, branch_limit)
records = []
for child in children:
mechanical_checks = run_deterministic_checks(child)
evidence_review = review_evidence(child)
records.append(
combine(child, mechanical_checks, evidence_review)
)
valid = remove_hard_failures(records)
completed += collect_verified_solutions(valid)
frontier = select_diverse_survivors(valid, beam_width)
if acceptance_rule_met(completed):
break
return best_verified_solution_or_request_for_missing_information()
The selector should not simply reward the longest candidate. It should favour constraint satisfaction, evidence quality and the user’s priorities.
Retaining some diversity can help. Two nearly identical survivors provide less protection against a mistaken premise than two meaningfully different feasible approaches.
Keep an audit trail, not a transcript dump
Store the inputs, candidate identifiers, checks, rejection reasons, tool results and final selection. These records help diagnose failure and reproduce decisions.
There is no need to store or request private internal reasoning.
For real systems, also record model versions and relevant configuration. A workflow that behaved well with one model may change when the model or prompt changes.
If tools can send messages, spend money or edit shared files, separate planning from execution. A candidate passing an evaluator does not itself authorise an external action.
Control branching, context and cost
Why small limits matter
Unrestricted branching grows quickly. With three children at every level, a full tree through four expansion levels contains 1 + 3 + 9 + 27 + 81 = 121 nodes, including the root.
That does not necessarily mean 121 model calls: implementations may batch candidates or use different calls for generation and evaluation. It does show why “explore every possibility” is usually impractical.
A beam search keeps only a limited number of candidates after each round. Retaining two candidates and expanding each into three children gives a manageable comparison, though it may discard a branch that would later have proved best.
That is a trade-off, not a bug to conceal.
Use adaptive depth
Not every candidate deserves equal effort.
Reject obvious budget failures immediately. Spend more effort where uncertainty could change the decision. Stop developing a candidate once a dominant alternative is established under the stated criteria.
A practical rule is:
Continue only if the next step could change the recommendation or resolve an important risk.
This prevents elaborate searches that add presentation detail without improving the decision.
Manage the context deliberately
Passing every previous response into every later call can bury the facts that matter.
The study Lost in the Middle found that models’ use of information could depend on its position within long contexts. Its precise findings belong to the models and evaluations studied, but the broader operational lesson remains useful: available context is not the same as reliably used context.
Keep a compact authoritative state:
- Current task and hard constraints.
- Verified facts with references.
- Surviving candidates.
- Rejected candidates and brief reasons.
- Open questions.
- Next permitted action.
Our article on What Is a Context Window and Why It Limits What AI Can Do explains why adding more history is not always helpful.
Define success before spending more
Useful stopping conditions include:
- A candidate passes all mandatory checks.
- Further ranking differences are immaterial to the user.
- The allowed number of rounds has been reached.
- Required evidence is unavailable.
- The next step exceeds the cost or time budget.
When stopping without a verified solution, report that plainly. “No acceptable candidate found within this search” is different from “no acceptable candidate exists”.
Three step-by-step exercises
Exercise one: compare a simple prompt with a bounded tree
Choose a low-stakes planning task, such as organising a study session.
- Write five supplied facts, two hard constraints and three preferences.
- Ask for a plan using a straightforward prompt.
- Save the answer without editing it.
- Run the bounded-tree template with the same inputs.
- Check both outputs against the original constraints.
- Compare factual assumptions, usability and time spent reviewing.
Do not judge only by length or sophistication. A shorter answer that satisfies the brief may be better.
Record whether branching discovered a useful alternative or merely restated the original plan in several forms.
Exercise two: build an evaluator that catches a tempting failure
Use a fictional purchasing decision.
- Set a strict budget and one non-negotiable requirement.
- Create three options with explicit prices and features.
- Make the most attractive option violate the requirement.
- Ask the assistant to rank them.
- Repeat with “filter hard failures before ranking”.
- Inspect whether the recommendation changes for the right reason.
For example, the cheapest printer might lack a required duplex scanning feature. The lesson is to distinguish attractiveness from eligibility.
As an extension, make one feature unknown rather than absent. Check whether the assistant requests evidence instead of granting a pass.
Exercise three: test whether revision actually helps
Take an answer containing a checkable calculation or code example.
- Ask the assistant to review its own answer without additional evidence.
- Save the revision.
- Supply a calculator result, failing test or source passage.
- Ask for a targeted repair.
- Recheck the repaired output.
- Compare which revision corrected the actual problem.
This exercise separates stylistic polishing from evidence-driven improvement.
If the first answer was already correct, watch for a different failure: an unsupported revision that makes it worse. A revision should earn its place through better evidence or a resolved defect.
Common mistakes and practical repairs
Generating cosmetic branches
Three proposals with different names but identical resources and assumptions are not meaningful alternatives.
Repair: specify the dimension of variation: delivery format, staffing arrangement, data structure, diagnostic hypothesis or risk strategy.
Ask the assistant to state the decisive difference between candidates in one sentence.
Allowing scores to override constraints
An overall rating can hide an impossible schedule or an unsupported prerequisite.
Repair: apply pass/fail filters first. Rank only eligible candidates, and keep unknowns visible.
If the user relaxes a constraint, record that as a changed brief rather than pretending the original candidate passed.
Asking for exhaustive exploration
Most real-world option spaces are not fully enumerable. Exhaustiveness can consume resources without establishing completeness.
Repair: specify branch limits, coverage goals and stopping rules. Say which important categories must be considered rather than asking for “all possible solutions”.
Treating model-generated facts as evidence
A branch may contain a plausible price, policy or technical capability that was never supplied or checked.
Repair: require evidence references and separate facts from assumptions. Remove unsupported details from the scoring basis until verified.
If the task depends on external documents, Retrieval-Augmented Generation (RAG) Explained for Beginners provides useful background on grounding outputs in retrieved material.
Over-pruning too early
A rough first draft may look weak because it lacks detail, while a polished candidate receives favourable evaluation.
Repair: compare candidates at a similar level of development. Allow one bounded repair where a branch’s weakness is local and clearly fixable.
Do not rescue candidates indefinitely. A repair budget prevents attachment to an attractive but unsuitable option.
Optimising for a convincing explanation
An elegant justification can distract from missing tests.
Repair: request the verification results before the recommendation. Ask what observation would falsify the preferred option.
Where possible, evaluate outputs without showing the reviewer the generator’s persuasive commentary.
Failing to test the workflow itself
A complex prompt may work on one example and fail on another.
Repair: build a small task set that includes ordinary cases, missing information, conflicting constraints and cases with no feasible answer. Compare against a simpler baseline.
Measure what matters: constraint violations, unsupported claims, successful tests, review effort and cost. Keep the more complex workflow only if it improves the outcomes you value.
A practical decision checklist
Before using an advanced reasoning pattern, check the following.
Problem: Are there meaningful alternatives, or is this really a lookup or extraction task?
Inputs: Which facts are supplied, which are verified externally, and which remain assumptions?
Constraints: What immediately disqualifies a candidate?
Branches: What must differ between options?
Evaluation: Which checks are mechanical, evidence-based or subjective?
Search: How many candidates and rounds are allowed?
Verification: What independent check supports the final recommendation?
Control: Which actions require human permission?
Stopping: When will the workflow stop, including without a solution?
Tree of Thoughts is most useful when it makes these decisions explicit. Its value is not that the assistant writes more. Its value is that alternatives become comparable, failures become visible, and the final recommendation rests on checks someone can inspect.
FAQ
Is Tree of Thoughts just brainstorming?
No. Brainstorming generates possibilities. A Tree of Thoughts workflow also evaluates candidates, develops selected branches, rejects failures and may revisit earlier choices.
If you ask for ten ideas and choose the nicest-sounding one, you have brainstorming rather than a controlled search.
Do I need to ask the model to reveal its chain of thought?
No. Ask for concise explanations, assumptions, candidate records, evidence and verification results.
These artefacts support checking without requesting private internal reasoning. For a calculation, show the relevant equations and totals. For a decision, show the criteria and evidence that distinguish the options.
Does Tree of Thoughts always outperform a direct prompt?
No. It adds generation, evaluation and review overhead. On simple tasks, that overhead may provide no benefit and can introduce extra errors.
Use a direct prompt as a baseline. Adopt branching when it improves constraint satisfaction, reveals useful alternatives or reduces costly mistakes.
How many branches should I start with?
Three distinct initial candidates is a reasonable practical starting point for ordinary chat use, followed by one or two survivors.
There is no universal optimum. Increase the budget only when additional branches cover genuinely different possibilities and you have a way to evaluate them.
Can one model generate and evaluate its own candidates?
Yes, but the evaluation may share the generator’s blind spots.
Separating stages helps organisation, not independence. Improve reliability through deterministic checks, authoritative sources, test execution or human review. Another model can provide a different perspective, but its agreement is still not proof.
What should happen when every branch fails?
Return the failure reasons and identify which constraint or missing fact prevents progress.
You can then request new information, propose a clearly labelled relaxation, or stop. Do not quietly weaken requirements or invent resources just to produce a successful-looking answer.
Can Tree of Thoughts solve hallucination problems?
It can expose inconsistent assumptions or unsupported claims, but it cannot manufacture reliable knowledge.
Branching without grounding may simply create several plausible inventions. Important factual claims still need source checking, retrieval or direct observation.
What is the simplest useful version to try?
Ask for three materially different options, reject hard-constraint failures, compare the survivors and verify the winner.
Keep the output short enough to inspect. If this reveals decisions that need revisiting, move to a staged workflow with explicit candidate records and a small search budget.
Sources
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
- ReAct: Synergizing Reasoning and Acting in Language Models
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Lost in the Middle: How Language Models Use Long Contexts
- Large Language Models Cannot Self-Correct Reasoning Yet
About the author
Editorial team · Editorial team
Author identity has not been supplied yet. Replace this record with the real writer's name, background and verifiable experience before publishing anything on the public site.
Spotted an error? Report a correction.
Related reading
A Repeatable AI Research Workflow: From Question to Verified Brief
A disciplined research process that uses AI for planning and synthesis while keeping every important claim tied to evidence you have checked.
How to Summarise Long Documents with AI Without Missing What Matters
A practical, source-grounded workflow for turning long documents into reliable summaries while preserving caveats, contradictions and important detail.
AI Meeting Notes: A Safe Workflow from Transcript to Action Items
A careful end-to-end method for using AI to draft meeting notes without inventing decisions, owners or deadlines.