Your domain is specialised
Medical coding, legal clauses, engineering logs or your own product catalogue. Models trained on your examples learn the distinctions general models miss.
Train
When a general-purpose model isn’t accurate, fast, cheap or private enough, we build one that is. Every model we deliver comes with its weights, an evaluation report and a model card.
Why custom
Hosted models are a great starting point. These are the signs you’ve outgrown them.
Medical coding, legal clauses, engineering logs or your own product catalogue. Models trained on your examples learn the distinctions general models miss.
At millions of requests, a small model that does one job well can cost far less to run than a large general model.
Smaller models respond faster, which matters for live chat, voice and anything that runs inside another process.
Custom models can run in your own cloud account, on your servers or fully offline, so sensitive data never goes to a third party.
Factory floors, vehicles, field devices and phones. Compact, quantised models can run with no connection at all.
We always benchmark a hosted model on your task before recommending custom training. If it’s good enough, we’ll tell you.
What we build
| Model type | What it does | Typical uses | Usual starting point |
|---|---|---|---|
| Fine-tuned language model | Writes, summarises and reasons in your format and terminology | Report drafting, case notes, support replies | Open-weight model adapted with LoRA |
| Small language model | A compact model distilled for one task, cheap and fast to run | High-volume classification, on-device assistants | Small open-weight model plus distillation |
| Classifier and extractor | Labels, routes and pulls structured fields out of text and documents | Ticket routing, invoice fields, risk flags | Encoder model fine-tuned on your labels |
| Embedding and reranking model | Understands similarity in your domain to make search smarter | Knowledge search, product matching, deduplication | Open embedding model tuned on your pairs |
| Vision model | Detects, classifies and measures objects in images and video | Defect inspection, stock counts, document photos | Pre-trained vision backbone, fine-tuned |
| Speech model | Transcribes and understands audio in your accents and jargon | Call centres, clinical dictation, field notes | Open speech model adapted to your audio |
A technique we use often
A frontier model, such as one of Anthropic’s Claude models, helps label and generate training examples for your task. A compact model then learns from those examples, checked by your experts. The result can come close to the large model on that one task, at a fraction of the running cost.
We only use model outputs for training where the provider’s terms allow it, and we keep a full record of the data each model was trained on.
Lifecycle
The order matters. We build the test before we train the model, so you can see whether training is worth it.
We agree exactly what the model must do, what “good” looks like, and the minimum score that would justify deploying it.
Output: task specYour experts help us assemble realistic test cases, including the hard and edge cases, held back from training.
Output: evaluation setWe test leading hosted models, such as Claude, on that evaluation set with good prompts. This is the bar a custom model has to clear.
Output: baseline reportWe clean, de-duplicate and label your data, remove personal data the model doesn’t need, and add synthetic examples where gaps exist.
Output: training set and data lineageWe train with the lightest method that works: parameter-efficient fine-tuning (LoRA), full fine-tuning or distillation. Each run is tracked and reproducible.
Output: candidate modelsWe score candidates against the baseline, check for bias across relevant groups, and test for misuse and failure modes.
Output: evaluation reportWe optimise and quantise the model for its target: your cloud, your servers or the device itself. Then we ship it behind a stable API.
Output: deployed model and APIWe track quality in production, catch drift as your data changes, and retrain on fresh examples when the numbers say it’s time.
Output: monitoring dashboardWhat you receive
Unless we agree otherwise, the trained weights and everything needed to run, audit and retrain them are yours.
Our toolkit
Questions
Less than most people think for fine-tuning. A few hundred to a few thousand high-quality examples is often enough for a focused task. Quality and coverage of edge cases matter more than volume. Where data is thin, we can generate synthetic examples for your experts to review.
You do. Your data is only used to train your model, and the weights are yours unless we agree otherwise in writing. Open-weight base models come with their own licences, and we explain those up front.
Yes. Compact, quantised models can run on laptops, phones, edge devices and air-gapped servers. We size the model to the hardware you have.
In three ways: as the baseline a custom model must beat, to help label data and draft synthetic examples for expert review, and often as part of the final solution alongside a custom model. We do this within each provider’s usage terms.
Then we don’t ship it, and we tell you. That’s why the evaluation set and baseline come before training: you find out early and cheaply.
Start a conversation