Skip to content

Pre-training, Fine-tuning and RLHF: How Chatbots Are Trained

Pre-training builds a chatbot’s broad capabilities, fine-tuning shapes its responses, and RLHF uses human preferences to make its behaviour more useful—but none guarantees accuracy.

By Editorial teamPublished 24 min read

Key takeaways

  • Pre-training learns broad patterns by predicting tokens across large datasets.
  • Supervised fine-tuning teaches response patterns through curated demonstrations.
  • RLHF adjusts behaviour using preference feedback, not a direct test of truth.
  • Prompts and retrieval change a model’s context; training changes its learned parameters.
  • Reliable customisation needs representative data, held-out tests and clear success criteria.
On this page
  1. One chatbot, several kinds of learning
  2. The basic mechanism: what training actually changes
  3. Pre-training: building broad capabilities
  4. Supervised fine-tuning: showing the model how to respond
  5. RLHF: learning which answers people prefer
  6. Beyond classic RLHF: preference training keeps evolving
  7. What each stage changes—and what it leaves unresolved
  8. Training, prompting and retrieval are different tools
  9. A practical customisation project, step by step
  10. Three exercises you can do without training a model
  11. Common mistakes and the habits that prevent them
  12. The mental model worth keeping
  13. FAQ

One chatbot, several kinds of learning

Ask a chatbot to explain a mortgage, rewrite an email or return a JSON object, and it may switch between those tasks without hesitation. That flexibility can make its training seem mysterious. Did someone write rules for every possible question? Did it memorise a library? Is it learning from you while you type?

The useful answer starts with three processes: pre-training, supervised fine-tuning and reinforcement learning from human feedback, usually shortened to RLHF.

They solve different problems:

  • Pre-training develops broad capabilities by learning patterns in large collections of data.
  • Supervised fine-tuning, or SFT, trains the model on examples of desired responses.
  • RLHF uses feedback about better and worse outputs to adjust the model’s behaviour.

These are not three sealed compartments. Fine-tuning is a broad term for further training an existing model, and RLHF can itself be part of fine-tuning. Developers may repeat stages, mix objectives or use alternative preference-training methods. Modern systems can also receive additional training for reasoning, tool use, coding or multimodal tasks.

Nevertheless, the three-part picture is a strong starting point. It explains why a model can know something yet answer badly, why an assistant can sound helpful while being wrong, and why uploading a document is usually not the same as training.

Throughout this article, we will use a fictional bicycle retailer, Harbour Cycles, as a running example. The retailer wants an assistant that answers maintenance questions, explains return policies and helps staff draft customer replies.

The challenge is not merely teaching bicycle vocabulary. It is making the assistant follow instructions, distinguish policy from guesswork and avoid inventing commitments.

The basic mechanism: what training actually changes

Tokens, parameters and predictions

A language model processes text as tokens: chunks that might represent a word, part of a word, punctuation or whitespace. Token boundaries depend on the model’s tokenizer. A word that looks ordinary to a reader may become several tokens.

Inside the model are learned numerical values called parameters, often referred to as weights. These values influence how it transforms an input into predictions about what should come next.

During training, an optimiser adjusts parameters to reduce an error measure called a loss. During ordinary use, those parameters generally remain fixed. The model instead uses the prompt and conversation as input.

For a deeper introduction to the machinery, see How Large Language Models Work: Tokens, Embeddings, Attention.

Imagine a training passage:

Before riding at night, check that your bicycle lights are charged.

For a typical next-token training objective, the model sees preceding tokens and predicts the next one. Early in training, its predictions may be poor. As training proceeds, it becomes better at assigning probability to the tokens that actually appear.

The important distinction is between predicting the training text and verifying the world. A model can reduce prediction errors by learning useful facts, linguistic structure and reasoning patterns. It can also learn stereotypes, repeated misconceptions and persuasive nonsense.

A small worked example of loss

Suppose a model must predict the next token in:

The capital of France is

For simplicity, imagine “Paris” is one token. If the model assigns it probability 0.10, the loss on that target token is approximately:

loss = -ln(0.10) ≈ 2.30

If later training raises the probability to 0.80:

loss = -ln(0.80) ≈ 0.22

Lower is better for this objective. The exact arithmetic is less important than the principle: training rewards assigning more probability to the observed target.

Real training combines losses across many tokens and examples. It also uses careful engineering to manage numerical precision, memory, learning rates and distributed computation.

Crucially, the training signal does not say, “This sentence has been independently verified.” It says, roughly, “This is the continuation in the data; become better at predicting it.”

Training is not the same as inference

Inference is running the trained model to produce an answer. Generation usually proceeds token by token, with each new token added to the context for the next prediction.

The context is temporary input, not necessarily a permanent change to the model. If you say, “For this conversation, call our fictional shop Harbour Cycles,” the assistant can use that information without updating its weights.

Some products separately save memories or retain conversations for future training. Those are product and data-governance choices, not an inevitable consequence of generating a reply.

Pre-training: building broad capabilities

Why predicting text teaches more than spelling

At first glance, next-token prediction sounds too simple to produce a useful assistant. But predicting diverse text well requires sensitivity to many kinds of structure.

To predict a recipe, a model benefits from recognising ingredients and cooking sequences. To predict code, it benefits from tracking syntax and variable relationships. To continue a mathematical explanation, it may benefit from learning operations and familiar proof patterns.

Across enough varied examples, the training task encourages reusable internal representations. These representations support capabilities that were not individually programmed as explicit rules.

The GPT-3 paper, “Language Models are Few-Shot Learners”, illustrates how a broadly pre-trained language model can perform many tasks from instructions or examples in its context, without a separate weight update for each task.

This does not mean next-token prediction guarantees robust reasoning. Performance can be impressive on familiar patterns and brittle on slightly different ones. A fluent answer is evidence of language-generation ability, not proof of dependable understanding in every setting.

Where pre-training data comes from

Depending on the developer and model, pre-training collections may include public web text, licensed material, code, books, academic writing and other sources. Multimodal systems may also use images, audio or video with different training objectives.

There is no universal dataset used by all chatbots. Some developers disclose substantial detail; others reveal relatively little.

Building a usable corpus is an engineering project in its own right. Common steps include:

  1. Collecting material with attention to rights and permitted use.
  2. Extracting useful content from files or web pages.
  3. Removing obvious spam and corrupted text.
  4. Identifying languages and content categories.
  5. Reducing duplication.
  6. Filtering or handling sensitive information.
  7. Choosing how frequently different sources appear during training.

The FineWeb dataset paper provides a concrete look at how web-data processing and filtering decisions affect language-model training. “More text” is not a complete recipe: the composition and cleanliness of that text matter.

Filtering also involves trade-offs. Aggressive filters can remove valuable dialects, specialist writing or discussions of sensitive subjects. Weak filters can leave abuse, spam and personal information in the corpus.

What Harbour Cycles inherits from pre-training

A broadly pre-trained model may already associate bicycles with tyres, brakes, gears, maintenance and retail transactions. It may recognise that a puncture is different from a slipping chain.

That foundation saves Harbour Cycles from assembling millions of basic language lessons.

However, the model will not reliably know the retailer’s current return policy. Even if an old version appeared somewhere in its training data, the model may recall it imperfectly, blend it with another shop’s rules or fail to distinguish old information from current information.

Nor does pre-training necessarily produce a cooperative assistant. Given:

Customer: Can I return a helmet after using it?
Assistant:

a base model might continue the dialogue appropriately. It could also generate an unhelpful fictional exchange, add more customer messages or imitate a poor-quality forum reply.

Its objective was to model plausible text, not consistently behave as Harbour Cycles’ authorised representative.

Learning patterns versus memorising passages

Pre-training is not simply building a searchable library. Knowledge is distributed across parameters, and the model generally cannot point to the exact passage responsible for an answer.

But the opposite claim—“models never memorise”—is also wrong. Some sequences can be memorised and reproduced, particularly under certain training and prompting conditions.

The research paper “Extracting Training Data from Large Language Models” demonstrated that training data can sometimes be recovered from language models. This is one reason personal information and confidential material require care before entering training pipelines.

A practical rule follows: do not treat parameter storage as a privacy guarantee. Information is not automatically safe merely because it has been absorbed into weights rather than stored in an ordinary document.

Supervised fine-tuning: showing the model how to respond

From predicting documents to following instructions

Supervised fine-tuning starts with an existing model and trains it on examples of desired behaviour. For an assistant, an example typically contains a prompt or conversation together with a target response.

The target demonstrates what a good answer looks like.

For Harbour Cycles, an illustrative example might be:

Customer:
My new bike's brakes feel weak. Can I keep riding until the weekend?

Target assistant response:
Avoid riding if the brakes are not working properly. Arrange a
check by a qualified mechanic before using the bike again.
If you bought it from Harbour Cycles, contact the service team
to arrange an inspection.

This teaches several things at once: recognise a safety-sensitive issue, avoid casual reassurance and suggest an appropriate next step.

The model is not receiving a hand-written rule in the traditional software sense. Its parameters are adjusted so responses resembling the target become more likely in similar contexts.

For chat training, implementations commonly calculate loss on assistant responses while masking some or all prompt tokens. Exact formatting and masking choices vary.

What a training example needs to include

Useful examples contain the information the model should legitimately use at deployment time.

Suppose the retailer wants the assistant to answer policy questions using a supplied policy extract. A better demonstration is:

Instruction:
Answer using only the supplied policy. If it does not cover
the customer's situation, say what remains unclear.

Policy:
Unused accessories may be returned within 30 days with proof
of purchase. Faulty goods follow a separate assessment process.

Customer:
I bought gloves 18 days ago and haven't used them. Can I return them?

Target response:
Based on the supplied policy, unused accessories can be returned
within 30 days with proof of purchase. Your gloves appear to meet
the time and usage conditions. Please bring proof of purchase.

The qualification “based on the supplied policy” is doing useful work. The answer does not invent a refund method, postage arrangement or statutory-rights exception.

Now compare a weak example:

Customer:
Can I return my gloves?

Target response:
Absolutely! We offer hassle-free returns on everything.

That target may sound friendly, but it teaches overconfidence and unsupported promises.

Fine-tuning quality is therefore not just about whether answers are grammatically correct. It is about whether the examples demonstrate the decisions you want repeated.

Fine-tuning teaches patterns, not exact guarantees

SFT can improve:

  • Tone and reading level.
  • Consistent output structure.
  • Task-specific terminology.
  • Classification and extraction behaviour.
  • Appropriate requests for missing information.
  • Tool-use patterns, when suitable examples are provided.

It does not guarantee that every response will follow a rule. Nor does adding a fact to training ensure that the model will retrieve it correctly whenever needed.

Fine-tuning is especially useful for stable, repeated behaviour. A policy that changes frequently is usually better supplied through context or retrieval.

For more on the distinction between instruction and weight changes, see System Prompts and Role Prompting: What They Change and What They Don’t.

Full fine-tuning versus adapters

In full fine-tuning, the training process can update the model’s existing parameters. This can be computationally demanding, especially for large models.

Parameter-efficient approaches train a smaller set of additional or selected parameters. LoRA, short for low-rank adaptation, introduces trainable low-rank updates to selected weight matrices while keeping the original weights frozen.

The LoRA paper explains the approach and its motivation.

A useful analogy is adjusting a machine through a smaller set of control surfaces rather than rebuilding every component. However, adapters are not separate factual databases, and they do not make bad training data harmless.

Training memory also includes more than the stored model weights: activations, gradients and optimiser state can matter. A model that fits on a device for inference may not fit there for full fine-tuning.

RLHF: learning which answers people prefer

Why demonstrations are not enough

There can be several acceptable answers to the same question. One may be concise but omit an important qualification; another may be complete but exhausting to read.

Writing an ideal answer from scratch is often difficult. Comparing two candidate answers can be easier.

RLHF uses human feedback as a training signal. In the classic language-model pipeline, humans compare outputs, a reward model learns their preferences, and reinforcement learning adjusts the assistant to obtain higher predicted rewards.

The earlier paper “Deep Reinforcement Learning from Human Preferences” established an influential framework for learning reward signals from comparisons. The assistant-focused InstructGPT paper describes a widely referenced combination of supervised fine-tuning and human-feedback training.

Step one: collect candidate answers

Start with a prompt:

My bicycle chain keeps slipping. Should I tighten the brakes?

The system generates candidate responses. Consider these simplified examples:

Answer A:
Yes. Tighten both brakes by two turns and test the bike on a hill.

Answer B:
Tightening the brakes will not fix a slipping chain. The issue
may involve gear adjustment or worn drivetrain parts. If it
happens while riding, stop and have the bike checked rather than
testing it in traffic.

A reviewer would normally prefer B because it corrects the mistaken connection, avoids hazardous instructions and acknowledges uncertainty about the diagnosis.

However, comparisons are not always this easy. What if one response is safer but unnecessarily alarmist? What if both contain a subtle technical mistake? Reviewers need guidance and, in some domains, specialist knowledge.

Step two: train a reward model

The preference dataset can contain prompts, candidate responses and labels indicating which response was preferred.

A reward model learns to assign scores that are consistent with these comparisons. In simplified terms, it should score B above A for this prompt.

Its score is not an objective measure of truth or a universal moral judgement. It is a learned approximation of the preferences represented in its training data.

This distinction matters. If reviewers tend to favour longer answers, the reward model may learn that length predicts quality. If confident wording repeatedly wins, confidence may become a shortcut for earning reward.

The reward model can therefore inherit both the strengths and the blind spots of the feedback process.

Step three: optimise the assistant

A reinforcement-learning algorithm then updates the assistant so it produces responses receiving higher reward-model scores.

Classic implementations often include a penalty for moving too far from a reference model. This helps limit drastic changes and reduces some opportunities to exploit weaknesses in the reward model.

One historically important algorithm is proximal policy optimisation, or PPO. You do not need its equations to understand the central process:

  1. The assistant generates answers.
  2. The reward model scores them.
  3. The training algorithm adjusts the assistant.
  4. Constraints discourage uncontrolled behavioural drift.
  5. Evaluation checks whether the changes actually help.

The reward model is often a training component, not something that must score every answer when users later chat with the assistant.

A worked example: reward is not correctness

Imagine reviewers assess a question about a complicated warranty exclusion.

One answer says:

Yes, you are definitely covered. We will replace the bike free
of charge.

Another says:

The supplied policy does not establish whether this damage is
covered. A staff member needs to review the cause of the damage
and the purchase details before confirming the outcome.

If reviewers optimise for immediate customer satisfaction, the first answer might win despite being unsupported.

The resulting training pressure would make the assistant more reassuring, not more reliable.

The remedy is not simply “more feedback”. It is a better feedback specification: reviewers should value grounded claims, appropriate uncertainty and correct escalation. They also need access to the information necessary to judge those qualities.

Beyond classic RLHF: preference training keeps evolving

Direct preference optimisation

Not every preference-trained chatbot uses a separate reward model followed by an online reinforcement-learning loop.

Direct preference optimisation, or DPO, trains from preferred and rejected responses using a different objective. It can avoid the separate reward-model training and policy-sampling loop associated with classic RLHF pipelines.

The DPO paper explains its relationship to reward-based preference optimisation.

For a non-specialist, the practical distinction is:

  • Classic RLHF commonly learns a scoring model, then optimises the assistant against it.
  • DPO directly uses preference pairs to adjust the assistant relative to a reference.

Both depend on the quality and coverage of preference data. Neither turns human judgements into infallible truth.

Terminology is sometimes loose. A provider may use “RLHF” informally to describe a broad family of human-feedback methods, even when the exact algorithm differs from the classic pipeline.

AI feedback and verifiable rewards

Some pipelines use AI systems to generate critiques, comparisons or preference labels. This is often called reinforcement learning from AI feedback, or RLAIF, although implementations differ.

“Constitutional AI: Harmlessness from AI Feedback” describes an approach using written principles and AI feedback to help shape behaviour.

AI-generated feedback can reduce some human-labelling effort. It can also repeat the evaluating model’s biases and mistakes. Human decisions remain present in the choice of principles, examples, evaluation methods and acceptable outcomes.

Other training tasks offer rewards that are more directly checkable. Code can be tested; some mathematics problems have known answers. These signals differ from a reviewer simply preferring one response.

Even checkable rewards have limits. Passing visible tests does not establish that a program works on every input. Optimising a narrow score can encourage solutions that exploit the test rather than satisfy the broader goal.

What each stage changes—and what it leaves unresolved

A compact comparison helps separate the objectives.

ProcessTypical signalMain contributionWhat it does not guarantee
Pre-trainingPredict tokens or other data elementsBroad representations and capabilitiesCurrent, sourced factual accuracy
Supervised fine-tuningMatch good demonstrationsInstruction-following and task behaviourPerfect compliance on unseen cases
Classic RLHFLearn and optimise preference rewardsBetter alignment with judged preferencesTruth or universal agreement
DPO and related methodsPreferred versus rejected responsesPreference-shaped behaviourReliable preferences from poor labels
Retrieval at inferenceSupply relevant external materialAccess to selected informationCorrect interpretation of that material

These processes interact. A fine-tuning dataset can improve one behaviour while weakening another. Preference optimisation may improve helpfulness while encouraging verbosity. Additional domain training may help technical vocabulary but change performance elsewhere.

This is why “the model has been fine-tuned” is not, by itself, evidence of improvement. The meaningful questions are: fine-tuned on what, towards which objective, and measured against which tests?

Why hallucinations survive training

A chatbot is trained to generate outputs under learned objectives. Those objectives may correlate with truth without directly establishing it.

An assistant may produce a false answer because it lacks relevant information, misinterprets context, generalises poorly or follows a learned pattern that sounds right. Preference training can reduce some problems but may also reward persuasive presentation.

For a closer breakdown, see Why AI Hallucinates: Causes, Types and How to Reduce Them.

For Harbour Cycles, a polished but invented return deadline is still a failure. The evaluation should penalise unsupported commitments regardless of how pleasant the answer sounds.

Why refusals and agreement can both go wrong

Training may encourage an assistant to refuse unsafe requests. If the examples are too broad, it may also refuse harmless educational questions.

Conversely, training for friendliness can encourage excessive agreement, sometimes called sycophancy. A model might endorse a user’s mistaken assumption instead of correcting it.

A good assistant should distinguish:

User:
I'm sure adjusting the brakes will fix my slipping chain.
Tell me I'm right.

from a legitimate request for reassurance. The desired answer is polite correction, not agreement for its own sake.

These boundaries require examples that include awkward cases, not only straightforward successes.

Training, prompting and retrieval are different tools

Prompting changes the immediate instructions

A prompt can tell an existing model what to do, what evidence to use and how to format the answer. It usually does not change the model’s weights.

For example:

You draft replies for Harbour Cycles.

Use only the supplied policy for policy claims.
Do not promise refunds or replacements unless explicitly supported.
If information is missing, ask one focused question.
Keep the reply under 120 words.

This may already solve much of the retailer’s problem.

Examples inside a prompt can also demonstrate the task without training. Few-Shot Prompting: Designing Examples That Teach the Model explains how to choose useful demonstrations.

The phrase “teach the model” is informal here: the model adapts its response to the context, but its underlying parameters remain unchanged.

Retrieval supplies information when it is needed

Retrieval-augmented generation, or RAG, finds relevant material and places it in the model’s context before the answer is generated.

Harbour Cycles could store its current policy documents in a retrieval system. When someone asks about returns, the system retrieves the relevant policy section and asks the model to answer from it.

This is usually easier to update and audit than repeatedly training policy changes into weights. See Retrieval-Augmented Generation (RAG) Explained for Beginners.

Retrieval is not a guarantee either. The system may retrieve an irrelevant section, miss an exception or supply outdated material. The model may then misread what it receives.

The practical advantage is that the evidence can be inspected, versioned and cited.

Choosing the smallest effective intervention

Use this decision sequence:

  1. Is the task unclear? Improve the instructions.
  2. Does the model need examples of the desired answer? Try a few demonstrations.
  3. Is it missing current or private information? Supply documents or build retrieval.
  4. Does a stable behaviour still fail repeatedly? Consider supervised fine-tuning.
  5. Do several plausible responses differ in subtle quality? Consider preference data.
  6. Are consequences significant? Add verification, permissions and human review.

These options can be combined. A fine-tuned assistant can still use retrieval and carefully designed prompts.

Also remember that supplied documents must fit within the model’s usable input budget. What Is a Context Window and Why It Limits What AI Can Do explains why giving a chatbot more material is not unlimited or automatically effective.

A practical customisation project, step by step

Step 1: define an observable task

Avoid goals such as “make the assistant understand our business”. They are too vague to test.

A better goal is:

Given a customer message and the relevant policy extract,
draft a reply that:
- answers the question where the policy supports an answer;
- makes no unsupported policy commitments;
- asks for essential missing information;
- uses a calm, plain-English tone.

Separate required behaviour from optional style. Inventing a refund is a substantive failure. Using a slightly stiff greeting may be a minor stylistic issue.

This prevents a pleasant tone from concealing a wrong decision.

Step 2: create a baseline before training

Test the existing model with a clear prompt and representative inputs. Record both successful and failed outputs.

For a small pilot, you might begin with several dozen carefully chosen cases. This is a practical starting point, not a statistically sufficient certification.

Include ordinary questions, ambiguous requests, missing evidence and hostile or misleading instructions. Ask whether the model already performs well enough for a supervised drafting workflow.

Without a baseline, you cannot tell whether fine-tuning improved anything. You may spend effort reproducing behaviour the model already had.

Step 3: build and review demonstrations

Collect examples that represent the real task. Remove personal information that is unnecessary for training and confirm that you have the appropriate rights and permissions.

Each example should contain:

  • The instructions available during deployment.
  • The relevant evidence or context.
  • A realistic customer request.
  • A reviewed target response.

Resolve contradictions before training. If one example promises automatic refunds and another requires an assessment in the same circumstances, the dataset is teaching inconsistency.

Do not include hidden information in the target that the deployed assistant would never receive. That teaches guessing rather than grounded answering.

Step 4: split the data without leakage

Use separate sets for training, development and final testing.

The development set helps you choose prompts, settings and model versions. The final test set checks the selected system on cases not used for those decisions.

Avoid splitting near-duplicate messages across sets. Five paraphrases of the same customer issue are not five independent tests if four appeared during training.

For customer support, it may make sense to split by conversation, underlying incident or policy scenario. Where future performance matters, testing on later time periods can reveal additional weaknesses.

Leakage makes results look stronger than they are.

Step 5: train cautiously and inspect behaviour

If fine-tuning is justified, begin with a manageable experiment rather than many simultaneous changes. Keep records of the base model, dataset version and training settings.

A falling training loss shows that the model is fitting the examples more closely. It does not prove the assistant is more useful.

Watch for overfitting: the model performs well on training-like examples but poorly on new cases. Also test for unwanted behavioural changes, such as increased verbosity or worse handling of missing information.

More training is not automatically better. Stopping decisions should consider held-out performance, not just the training curve.

Step 6: evaluate decisions, not merely wording

A simple rubric could score:

CriterionWhat the reviewer checks
GroundingAre policy claims supported by the supplied text?
CompletenessDoes the reply address the customer’s actual question?
UncertaintyDoes it identify important missing information?
SafetyDoes it avoid risky maintenance advice or false reassurance?
FormatDoes it meet the required structure and length?
ToneIs it clear, respectful and not misleadingly confident?

Compare the customised model against the baseline on the same cases. Where possible, hide which system produced which answer.

Treat severe failures separately from average scores. A system with excellent tone and occasional invented financial commitments may still be unsuitable for unsupervised use.

Step 7: deploy with limits and monitor changes

Start with human-reviewed drafts if the outputs can create obligations or safety risks.

Log failures in a privacy-conscious way. Distinguish retrieval errors, unclear policies, model errors and user-interface problems; each needs a different remedy.

Re-run evaluations when the model, prompt, retrieval collection or policy changes. A better base model can make an old fine-tune unnecessary, while a changed policy can invalidate previous examples.

Keep a rollback option. Model customisation is an ongoing maintenance commitment, not a one-off upload.

Three exercises you can do without training a model

Exercise 1: separate knowledge from behaviour

Choose a fictional policy:

Harbour Cycles lends courtesy bikes only when a repair is expected
to take more than two working days. Availability is not guaranteed.

Then ask a chatbot:

Using only this policy, answer the customer:
"My repair takes three working days. Am I guaranteed a courtesy bike?"

Check whether it distinguishes eligibility from availability. Next, remove the policy and ask the same question.

Write down what changed. The first task tests interpretation of supplied evidence. The second invites reliance on general patterns or unsupported assumptions.

The lesson: better access to information and better answering behaviour are separate needs.

Exercise 2: design a preference rubric

Write two replies to this customer:

My helmet cracked after a crash. Can I keep using it if I glue it?

Make one reply friendly but unsafe and the other clear, cautious and useful. Then define three criteria that explain why the safer reply should win.

Now create a harder pair: both recommend replacing the helmet, but one contains unnecessary alarmist language.

This exposes the limits of a single “good/bad” label. Helpful preference data often requires reviewers to balance correctness, risk, tone and relevance rather than rewarding one visible feature.

Exercise 3: create a tiny evaluation suite

Draft ten test cases for an assistant you might actually use. Include:

  1. A straightforward request.
  2. A request with missing information.
  3. A misleading assumption.
  4. Conflicting source passages.
  5. An instruction to ignore the supplied evidence.
  6. A request outside the assistant’s scope.
  7. A case requiring an explicit uncertainty statement.
  8. A strict formatting requirement.
  9. An unusually long input.
  10. A familiar question with one important detail changed.

For each case, write what success requires before generating answers. Otherwise, you may unconsciously redefine success to fit a fluent response.

This suite will not establish production readiness, but it gives you a repeatable way to compare prompts and models.

Common mistakes and the habits that prevent them

Mistake: treating fine-tuning as a document upload

Training on a handbook does not provide reliable lookup, source attribution or easy deletion of individual facts.

Better habit: use retrieval for changing information and fine-tuning for persistent behaviour. If you do further training on domain text, evaluate whether it helps the actual downstream task.

Mistake: assuming human feedback means expert verification

Reviewers may judge readability better than technical correctness. They may also disagree or lack essential context.

Better habit: match reviewer expertise to the task and supply the evidence needed to make informed comparisons. Track disagreement rather than hiding it inside an average score.

Mistake: collecting only easy examples

A dataset of straightforward successes may teach polished answers while leaving the decision boundaries weak.

Better habit: include cases where the correct action is to ask, qualify, correct a premise, use a tool or decline a request. Include benign cases that should not trigger refusal.

Mistake: equating a lower loss with a better product

Loss measures progress on a training objective. Users care about reliable completion of their tasks.

Better habit: maintain a task-level evaluation with explicit failure categories. Inspect cases where the metric improves but the answer becomes less useful.

Mistake: training on every conversation automatically

Raw conversations can contain secrets, incorrect staff advice, abusive content and low-quality assistant outputs. Recycling them indiscriminately can reinforce mistakes.

Better habit: establish consent and retention rules, minimise sensitive data and review examples before adding them. Keep provenance so you know why each example exists.

Mistake: expecting training to replace software controls

A model instructed not to issue refunds might still produce an incorrect tool call. A model trained to follow a schema might occasionally violate it.

Better habit: validate outputs, constrain tool permissions and enforce business rules outside the model. Behavioural training should support controls, not substitute for them.

The mental model worth keeping

A chatbot’s behaviour comes from more than one source.

Pre-training develops broad capabilities and tendencies. Supervised fine-tuning demonstrates how those capabilities should be expressed. Preference training shifts behaviour towards outputs that evaluators favour. At deployment, prompts, retrieved documents, tools and software controls shape what happens in a particular interaction.

None of these layers deserves blind trust. Each solves some problems and leaves others unresolved.

For Harbour Cycles, the sensible starting point is not training a new model. It is a clear task, current policy evidence, carefully written instructions and a repeatable test suite. Fine-tuning becomes useful when that baseline exposes stable behavioural failures that good training examples can address.

The most productive question is therefore not “Has this chatbot been trained?” Every language-model chatbot has. It is: what was optimised, what evidence does it have now, and how do we know its answers meet the requirements?

FAQ

Is RLHF a type of fine-tuning?

Usually, yes, in the broad sense of further training an existing model. People often use “fine-tuning” as shorthand for supervised fine-tuning, which makes the terms sound more separate than they are. A clearer distinction is between demonstration-based training and preference- or reward-based training.

Does every chatbot use pre-training, SFT and RLHF in that order?

No. That sequence is an influential template, not a universal recipe. Developers can mix datasets, repeat stages, use DPO instead of classic RLHF, or add specialised training with other rewards. Public information may not reveal the full process for a particular model.

Does a chatbot learn permanently when I correct it?

Usually not during ordinary conversation. Your correction becomes part of the current context and may affect later replies in that conversation. A product may also save memories or use eligible conversations in later training, depending on its settings and policies. Those are separate mechanisms.

Can fine-tuning teach a model new facts?

Yes, further training can change what information a model represents and produces. However, factual recall can remain incomplete or unreliable, and updates are harder to inspect than a document change. Retrieval is often a better choice for current policies, product details and information requiring citations.

How many examples do I need for fine-tuning?

There is no reliable universal number. Requirements depend on the model, task difficulty, example quality and how different the desired behaviour is from the baseline. Start with representative, reviewed examples and a held-out evaluation. Expand the dataset where testing reveals gaps rather than chasing an arbitrary total.

Does RLHF make a model truthful?

It can improve truth-related behaviour when reviewers and rewards favour accuracy, grounding and appropriate uncertainty. But preferences are an imperfect proxy for truth. If confident or agreeable answers receive better labels, training may strengthen those traits even when the answers are wrong.

Can a small team carry out RLHF?

Some can, especially with suitable tooling and narrow tasks, but a classic reward-model-and-reinforcement-learning pipeline introduces considerable complexity. Many teams should first try prompting, retrieval and supervised fine-tuning. Preference-based alternatives such as DPO can simplify parts of training, but still require strong data and evaluation.

What is the simplest way to judge a customised chatbot?

Compare it with an untuned baseline on the same unseen, representative cases. Score supported claims, task completion, uncertainty, safety and format separately. Review serious failures individually, not just averages. If the customised version does not deliver a meaningful improvement, the added training and maintenance may not be justified.

Sources

About the author

Editorial team · Editorial team

Author identity has not been supplied yet. Replace this record with the real writer's name, background and verifiable experience before publishing anything on the public site.

Full profile

Spotted an error? Report a correction.