Train

Models trainedfor one job.Owned by you.

When a general-purpose model isn’t accurate, fast, cheap or private enough, we build one that is. Every model we deliver comes with its weights, an evaluation report and a model card.

Why custom

Five reasons to stop renting intelligence.

Hosted models are a great starting point. These are the signs you’ve outgrown them.

Accuracy

Your domain is specialised

Medical coding, legal clauses, engineering logs or your own product catalogue. Models trained on your examples learn the distinctions general models miss.

Cost

Your volume is high

At millions of requests, a small model that does one job well can cost far less to run than a large general model.

Speed

Every millisecond counts

Smaller models respond faster, which matters for live chat, voice and anything that runs inside another process.

Privacy

Data can’t leave your walls

Custom models can run in your own cloud account, on your servers or fully offline, so sensitive data never goes to a third party.

Edge

It has to work offline

Factory floors, vehicles, field devices and phones. Compact, quantised models can run with no connection at all.

Not sure?

We’ll test it first

We always benchmark a hosted model on your task before recommending custom training. If it’s good enough, we’ll tell you.

How we decide

What we build

Six kinds of model, one discipline.

Model typeWhat it doesTypical usesUsual starting point
Fine-tuned language modelWrites, summarises and reasons in your format and terminologyReport drafting, case notes, support repliesOpen-weight model adapted with LoRA
Small language modelA compact model distilled for one task, cheap and fast to runHigh-volume classification, on-device assistantsSmall open-weight model plus distillation
Classifier and extractorLabels, routes and pulls structured fields out of text and documentsTicket routing, invoice fields, risk flagsEncoder model fine-tuned on your labels
Embedding and reranking modelUnderstands similarity in your domain to make search smarterKnowledge search, product matching, deduplicationOpen embedding model tuned on your pairs
Vision modelDetects, classifies and measures objects in images and videoDefect inspection, stock counts, document photosPre-trained vision backbone, fine-tuned
Speech modelTranscribes and understands audio in your accents and jargonCall centres, clinical dictation, field notesOpen speech model adapted to your audio

A technique we use often

Distillation: a large model teaches a small one.

A frontier model, such as one of Anthropic’s Claude models, helps label and generate training examples for your task. A compact model then learns from those examples, checked by your experts. The result can come close to the large model on that one task, at a fraction of the running cost.

We only use model outputs for training where the provider’s terms allow it, and we keep a full record of the data each model was trained on.

Lifecycle

Eight stages, in this order, every time.

The order matters. We build the test before we train the model, so you can see whether training is worth it.

  1. Define the task and the metric

    We agree exactly what the model must do, what “good” looks like, and the minimum score that would justify deploying it.

    Output: task spec
  2. Build the evaluation set

    Your experts help us assemble realistic test cases, including the hard and edge cases, held back from training.

    Output: evaluation set
  3. Measure a baseline

    We test leading hosted models, such as Claude, on that evaluation set with good prompts. This is the bar a custom model has to clear.

    Output: baseline report
  4. Curate the data

    We clean, de-duplicate and label your data, remove personal data the model doesn’t need, and add synthetic examples where gaps exist.

    Output: training set and data lineage
  5. Train

    We train with the lightest method that works: parameter-efficient fine-tuning (LoRA), full fine-tuning or distillation. Each run is tracked and reproducible.

    Output: candidate models
  6. Evaluate and red-team

    We score candidates against the baseline, check for bias across relevant groups, and test for misuse and failure modes.

    Output: evaluation report
  7. Deploy

    We optimise and quantise the model for its target: your cloud, your servers or the device itself. Then we ship it behind a stable API.

    Output: deployed model and API
  8. Monitor and retrain

    We track quality in production, catch drift as your data changes, and retrain on fresh examples when the numbers say it’s time.

    Output: monitoring dashboard

What you receive

You own the model. Not a licence to it.

Unless we agree otherwise, the trained weights and everything needed to run, audit and retrain them are yours.

  • Model weightsIn standard formats (Hugging Face Safetensors, GGUF or ONNX) that run on common serving stacks.
  • Model cardWhat the model is for and not for, how it was trained, its known limitations and its intended users.
  • Evaluation reportScores against the baseline on your evaluation set, broken down by case type, including where the model fails.
  • Data lineageA record of exactly which data trained the model and how it was processed, ready for audits and data protection requests.
  • Training recipeCode and configuration to reproduce or retrain the model, so you’re never dependent on us.
  • Deployment packageContainer, API and infrastructure templates for your chosen environment.

Our toolkit

Open, well-supported tools. Nothing proprietary to trap you.

PyTorchTraining
Hugging FaceTransformers, PEFT, Datasets
Open-weight modelsLlama, Mistral, Qwen, Gemma
Anthropic ClaudeBaselines, labelling, synthetic data
vLLMHigh-throughput serving
llama.cpp and GGUFCPU and edge inference
ONNX RuntimeCross-platform deployment
MLXApple silicon
AWS, Azure, GCPGPU training and hosting
Weights and BiasesExperiment tracking

Questions

Custom model FAQs

How much data do we need?

Less than most people think for fine-tuning. A few hundred to a few thousand high-quality examples is often enough for a focused task. Quality and coverage of edge cases matter more than volume. Where data is thin, we can generate synthetic examples for your experts to review.

Who owns the model and the data?

You do. Your data is only used to train your model, and the weights are yours unless we agree otherwise in writing. Open-weight base models come with their own licences, and we explain those up front.

Can it run without an internet connection?

Yes. Compact, quantised models can run on laptops, phones, edge devices and air-gapped servers. We size the model to the hardware you have.

How do you use frontier models like Claude in this work?

In three ways: as the baseline a custom model must beat, to help label data and draft synthetic examples for expert review, and often as part of the final solution alongside a custom model. We do this within each provider’s usage terms.

What if the custom model isn’t better than the baseline?

Then we don’t ship it, and we tell you. That’s why the evaluation set and baseline come before training: you find out early and cheaply.

Start a conversation

Tell us the problem.
We’ll tell you honestly if AI can solve it.