The ladder
When a team decides to "build something with AI", the first question is usually about models: which one, and whether to train our own. That is the wrong place to start. Building a language-model capability involves a series of options, each costing more than the one before:
- Prompt engineering with a hosted model such as Anthropic's Claude or OpenAI's GPT models.
- Retrieval-augmented generation (RAG): give the model the right documents at the moment it answers.
- Fine-tuning: adjust a model's weights using your own examples, often with parameter-efficient methods such as LoRA.
- Distillation: train a small model to copy a larger model's behaviour on your task.
- Training from scratch: build a new model from raw data.
Each rung needs more data, more specialist skill and more time to run. The rungs also do different jobs. RAG changes what the model knows at answer time. Fine-tuning changes how it behaves. Treating them as interchangeable is the commonest mistake we see.
Build your evaluation set first
You cannot pick a rung if you have no way to measure whether it works. Before you choose an approach, collect an evaluation set: real inputs your system will face, each paired with what a good output looks like or a clear rule for judging one. Anthropic's prompt engineering guidance says the same. It assumes you already have clear success criteria and a way to test against them before you start refining prompts. OpenAI's model optimisation guide also starts with evals that set a baseline.
A useful evaluation set:
- reflects the real mix of requests, not just the easy ones;
- includes edge cases, ambiguous inputs and things the system should refuse;
- defines success in measurable terms, such as "extracts the correct invoice total", not "gives good answers";
- can be scored automatically where possible, with exact match, code checks or a carefully checked model-based grader, plus human review for the subjective parts;
- is kept separate from any data you later use for training, so you are not marking your own homework.
A hundred or so well-chosen cases is often enough to tell the options apart. The set then becomes the regression test you run every time you change a prompt, a model or a data source.
Key point
The evaluation set is the most valuable asset in an AI project. It outlasts any model choice, and it turns "which approach?" from a matter of opinion into a measurement.
Rung 1: Prompt engineering with a hosted model
Here you use a capable general model through an API and shape its behaviour with instructions: clear task descriptions, worked examples, a defined output format, and complex jobs split into steps. Nothing is trained. You can change everything in minutes.
When to use it: almost always, as a starting point. OpenAI's own guidance says prompting may be all you need.
What it needs: a person who understands the task and writes clearly, your evaluation set, and days rather than months. No training data is required beyond a handful of examples.
Risks: per-call cost and latency can climb when prompts grow long; the provider may update or retire models under you; and sending data to a third party needs a proper data protection review.
How to evaluate: run the full evaluation set against each prompt version and each candidate model, and record the scores. Sometimes a smaller or different model fixes a cost or latency problem more easily than any amount of prompt rewriting.
Rung 2: Retrieval-augmented generation
RAG combines a language model with a searchable store of your own content. When a question arrives, the system retrieves the most relevant passages and puts them in the prompt, so the model answers from your material instead of its general training. The original 2020 RAG paper from Lewis and colleagues set out the main benefits: answers more grounded in facts, a record of which sources informed an answer, and knowledge you can update without retraining.
When to use it: when answers depend on information the model cannot know, such as policies, product documentation, contracts or case notes, especially if that information changes or users need citations. If your whole knowledge base is small, try putting it straight into the prompt before building retrieval; Anthropic notes this is often the simpler route for small corpora.
What it needs: clean, current, access-controlled source documents; engineering skills in document processing, search and data pipelines; and typically a few weeks to build a solid first version.
Risks: most RAG failures are retrieval failures. The right passage was never found, so the model answers confidently from the wrong one. Other risks are stale or contradictory documents, and leaking content to users who should not see it. Combining embedding search with keyword search, then re-ranking, often helps.
How to evaluate: score retrieval and generation separately. For each test question, note which documents should be retrieved and check whether they were. Then check whether the answer is faithful to the retrieved text and cites it correctly.
Rung 3: Fine-tuning, including LoRA
Fine-tuning continues training an existing model on your own examples, usually pairs of inputs and ideal outputs, so the behaviour becomes part of the model. Full fine-tuning updates every weight, which is expensive for large models. Parameter-efficient fine-tuning (PEFT) updates only a small number of parameters. The best-known method, LoRA (Hu et al., 2021), freezes the original weights and trains small low-rank matrices alongside them. The LoRA authors report quality on a par with full fine-tuning on the models they tested, with far fewer trainable parameters and no extra inference latency once the adapters are merged. QLoRA (Dettmers et al., 2023) adds 4-bit quantisation, bringing large models within reach of a single GPU.
When to use it: when prompting has hit a ceiling on behaviour: a strict output format, a house style, a specialist classification scheme, or a task where you have more good examples than fit in a prompt. Fine-tuning is a poor way to teach a model facts that change; use RAG for that.
What it needs: hundreds to thousands of high-quality, consistent examples (quality matters far more than quantity), ML engineering skill, GPU access or a managed tuning service, and weeks of iteration. Managed options exist: OpenAI offers supervised, preference-based (DPO) and reinforcement fine-tuning for selected models, and AWS Bedrock and Google Vertex AI offer tuning for selected models on their platforms.
Risks: the model can overfit to its training examples, or lose general abilities you relied on. Errors in the training data get built into the model. You also take on versioning and retraining work each time the base model changes.
How to evaluate: compare the fine-tuned model against your best prompted baseline on the same held-out evaluation set. Add general-capability checks so you notice if it has got worse at things it used to do well.
Key point
RAG is for knowledge; fine-tuning is for behaviour. If the problem is "it doesn't know our policies", retrieve them. If the problem is "it knows, but never answers in the shape we need", consider tuning.
Rung 4: Distillation into small models
Distillation trains a small "student" model to imitate a larger "teacher". The idea goes back to Hinton, Vinyals and Dean (2015). Today it usually means running a strong model over many real inputs, checking its outputs, and fine-tuning a small open-weight model on them.
When to use it: when a large model already does the task well but costs too much or runs too slowly at your volume, or when you need the model to run on your own hardware, at the edge or offline. It suits narrow, repetitive, high-volume tasks.
What it needs: a working large-model solution to copy, a large and representative set of inputs, fine-tuning skills, and serving infrastructure for the student. AWS Bedrock and Vertex AI both offer managed distillation for some models. Check the teacher model's terms of use, as some providers restrict using outputs to train other models.
Risks: the student is only as good as the teacher's outputs, and it can fail badly on inputs unlike its training data.
How to evaluate: measure the gap between student and teacher on your evaluation set, look closely at the cases where they disagree, and test on recent real traffic, not only the data used for distillation.
Rung 5: Training from scratch, and why it is rarely right
Pre-training a language model from raw text takes very large datasets, large GPU clusters, rare research expertise and months of work, before any of the fine-tuning and safety work that makes a model usable. Open-weight models trained by well-funded labs are freely available as starting points. A business that trains from scratch is very unlikely to match them on general ability.
There are real exceptions: models for data that is not natural language, such as sensor readings or specialised sequences; very small models for a single narrow function; or organisations whose product is the model. Even then, adapting an existing base model is usually the faster route. Before committing, ask what a fine-tuned open-weight model would fail to do.
If no rung below has been tested against an evaluation set and failed for a reason you can name, it is too early to discuss training from scratch.
Side-by-side comparison
| Approach | Best for | Data needed | Skills | Typical time to first result | Main risk |
|---|---|---|---|---|---|
| Prompting | Most tasks; quick validation | Evaluation set and a few examples | Domain knowledge, clear writing | Days | Cost and latency at scale; provider changes |
| RAG | Answers grounded in your own, changing documents | Clean, permissioned source content | Search, data engineering | Weeks | Retrieving the wrong passage |
| Fine-tuning (incl. LoRA) | Consistent format, style or specialist behaviour | Hundreds to thousands of curated examples | ML engineering | Weeks to a few months | Overfitting; built-in errors; upkeep |
| Distillation | Cheaper, faster or private serving of a proven task | Large set of representative inputs plus teacher outputs | ML engineering, model serving | Weeks to months | Fails on unfamiliar inputs |
| From scratch | Unusual data types or the model as the product | Very large corpora | Research-grade ML | Many months | Cost with no guarantee of beating open models |
Timings are rough guides for a well-scoped project, not quotes. The rungs also combine; RAG with a lightly tuned model is common.
Models and hosting
Your hosting choice affects cost, data residency and control as much as your model choice does.
Proprietary hosted models
Claude and GPT models are available directly from Anthropic and OpenAI, and through the major clouds. AWS Bedrock offers models from Anthropic, OpenAI, Meta, Mistral AI, Qwen, Google (Gemma) and others. Microsoft Foundry on Azure offers OpenAI models and Claude, among others. Google Vertex AI's Model Garden includes Gemini and Claude alongside open models. Using your existing cloud can simplify contracts and data controls. Check which regions each model is offered in before assuming UK or EU processing.
Open-weight models
Families such as Meta's Llama, Mistral, Alibaba's Qwen and Google's Gemma publish their weights, so you can fine-tune, distil and run them yourself. Licences differ: some are permissive, others carry conditions on use or attribution. Read the licence for each model before you build on it. You can run open models as managed endpoints (Bedrock and Vertex AI both offer this, and Bedrock can import your own fine-tuned weights for supported architectures) or self-host them using serving software such as vLLM on your own or rented GPUs.
Self-hosting gives you the most control over data and versions, but you then own capacity planning, security patching, monitoring and uptime.
Key point
Keep your application loosely tied to any single model. If your prompts, retrieval and evaluation set are portable, switching provider or moving to an open-weight model becomes a measured experiment rather than a rebuild.
A decision guide
Work through these questions in order and stop at the first approach that passes your evaluation set.
- Can you define success and collect test cases? If not, do that first. Nothing else is worth starting yet.
- Does a strong hosted model, well prompted, pass? If yes, ship it, monitor it, and revisit cost later.
- Does it fail because it lacks your information? Add the content to the prompt if it is small, or build RAG if it is large or changes often.
- Does it fail on behaviour despite good prompts and context? Consider fine-tuning, starting with LoRA on a suitable base model or a managed tuning service.
- Does it work, but cost too much, run too slowly or need to run privately? Consider distilling into a small open-weight model you can host where you need it.
- Have all of these failed for a reason you can clearly state? Only then look at training a custom model from scratch, and get expert advice first.
Before committing budget to any rung, confirm the basics:
- An evaluation set of real, representative cases, held separately from any training data
- Written success criteria that a non-specialist could check
- A measured baseline from a well-prompted hosted model
- A data protection assessment covering where data is processed and stored
- Licence and terms-of-use checks for every model and dataset involved
- An owner for monitoring, retraining and model upgrades after launch
- A clear reason, backed by evaluation results, for moving to the next rung
Most projects end at rung one or two, and that counts as success. The aim is a system that reliably does the job at a cost you can sustain. Whether you trained a model to get there doesn't matter.
Sources
- Prompt engineering overview — Anthropic (Claude Docs)
- Define success criteria and build evaluations — Anthropic (Claude Docs)
- Introducing Contextual Retrieval (September 2024) — Anthropic
- Model optimization — OpenAI API documentation
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al., 2020) — arXiv
- LoRA: Low-Rank Adaptation of Large Language Models (Hu et al., 2021) — arXiv
- QLoRA: Efficient Finetuning of Quantized LLMs (Dettmers et al., 2023) — arXiv
- Distilling the Knowledge in a Neural Network (Hinton, Vinyals and Dean, 2015) — arXiv
- PEFT documentation — Hugging Face
- Models at a glance — Amazon Bedrock User Guide
- Customize your model — Amazon Bedrock User Guide
- Custom model import — Amazon Bedrock User Guide
- Claude models in Microsoft Foundry — Microsoft Learn
- Overview of Model Garden — Google Cloud documentation
- vLLM documentation — vLLM project