Custom model · Privacy

Personal data,redacted offline.

We built a small AI model that finds and masks UK personal data in free text. It runs on an ordinary laptop, so sensitive text never has to leave your organisation. Here’s how we built it, and how we tested it honestly.

What it does

Ten kinds of UK personal data, replaced with clear tags.

Names, addresses, postcodes, phone numbers, emails, National Insurance and NHS numbers, sort codes, account numbers and dates of birth. Everything else stays exactly as it was.

Before

Hi, this is Fiona Gallagher from 12 Larch Way, Durham DH1 3PQ. Call me on 07700 900512 - order UK-44120.

After

Hi, this is [NAME] from [ADDRESS] [POSTCODE]. Call me on [PHONE] - order UK-44120.

Examples

Real outputs, including the one it got wrong.

Every example below is exactly what the published model produced, after the safety layers. All names and numbers are fictional.

Messy chat message

In

hiya its dan, my nan Edith lives at flat 3, 18 hope street, liverpool l1 9bq. her email is edith.lowe@example.co.uk

Out

hiya its [NAME], my nan [NAME] lives at [ADDRESS] [POSTCODE]. her email is [EMAIL]

✓ Integrity check passed

Clinical note

In

Patient: Mr Arjun Mehta
DOB: 14/02/1979
NHS number: 485 777 3456
Seen at the Royal Infirmary on 3 March 2026.

Out

Patient: Mr [NAME]
DOB: [DATE_OF_BIRTH]
NHS number: [NHS_NUMBER]
Seen at the Royal Infirmary on 3 March 2026.

✓ Integrity check passed

Bank refund request

In

Please refund £42.50 to Kate Lowe, sort code 20-41-77, account 30918274. Ref INV-55120.

Out

Please refund £42.50 to [NAME], sort code [SORT_CODE], account [ACCOUNT_NUMBER]. Ref INV-55120.

✓ Integrity check passed

Email signature

In

Kind regards,
Owen Bradshaw | Facilities Manager | Harbour Water
M: 07700 900611

Out

Kind regards,
[NAME] | Facilities Manager | Harbour Water
M: [PHONE]

✓ Integrity check passed

Trap: a person's name that's a brand

In

The Paul Smith store on King Street opens at 9am on Saturday.

Out

The Paul Smith store on King Street opens at 9am on Saturday.

✓ Integrity check passed

Flagged by the integrity check

In

Lydia's DOB is 05/06/2003 and her email's lydia.f@post.example.net

Out

[NAME]'s DOB is [DATE_OF_BIRTH] and his email's [EMAIL]

⚠ Flagged for review: the model changed “her” to “his”. The check caught it.

Run it yourself

Free to download. Runs offline.

The model is published under the Apache-2.0 licence. It is under 1 GB and runs on an ordinary laptop with no internet connection.

  • LM Studio or OllamaSearch for QuantumAiuk/Qwen2.5-1.5B-UK-PII-Redactor and choose the Q4_K_M file. Set the system prompt from the model card and temperature to 0.
  • Apple silicon (MLX)Install mlx-lm and load QuantumAiuk/Qwen2.5-1.5B-UK-PII-Redactor. The model card has a copy-and-paste example.
  • Safety layersDownload guard.py from the model page to add the pattern backstop and integrity check to your own pipeline.
  • Training data and testsThe full synthetic dataset, both hand-written test sets and the generators are published, so you can check our numbers.

Results

Tested on messages it had never seen.

The headline test is 30 messy, hand-written messages (chat logs, email signatures, lowercase text, and traps such as people’s names inside shop names), written before the final model was trained.

SystemPersonal data leaked (lower is better)Tags correct (F1)Whole message exactly right
Pattern-matching rules only55.3%61.8%36.7%
The same AI model before training78.7%32.9%23.3%
Our model2.1%97.9%83.3%
Our model + safety layers0%98.9%86.7%

30 messages containing 47 personal-data items. A second test of 200 generated records, using sentence patterns and names never seen in training, showed a 0.2% leak rate for the model alone and 0% with the safety layers. These are small test sets, and all data is synthetic, so we publish the full method and every test record alongside the model.

How we built it

The same process we use for client models.

  1. Defined the task and the metric

    Ten categories of UK personal data, and one number that matters most for privacy: how much personal data leaks through.

  2. Built the test before the model

    A generated test set using sentence patterns and names held back from training, plus hand-written messages in styles our generator never produces.

  3. Measured the baselines

    Pattern-matching rules leaked about half the personal data: they can’t recognise names or street addresses. The untrained model leaked even more.

  4. Created safe training data

    3,000 synthetic examples. Phone numbers come from Ofcom’s ranges reserved for TV and drama, and no real person’s details are used anywhere.

  5. Trained on a laptop

    We fine-tuned Qwen2.5-1.5B-Instruct (an open, Apache-licensed model) with LoRA on an Apple silicon laptop. Training took minutes and cost nothing in cloud fees.

  6. Found the weaknesses and fixed them

    The first version hid ordinary place names and sometimes dropped words. We broadened the training data, then tested again on messages written before retraining.

  7. Added safety layers

    A pattern-matching backstop catches anything structured the model misses, and an integrity check flags any output where text other than personal data was changed.

  8. Packaged it to run anywhere

    A 4-bit version under 1 GB for Mac (MLX) and for llama.cpp, Ollama and LM Studio. It leaks no more personal data than the full-size model on our tests.

Honest limitations

What it doesn’t do.

We publish these because you should know them before relying on any model.

  • It isn’t a guarantee of anonymisationText can still identify someone through context. Under UK GDPR, redacted data can still be personal data. Use it as one step in a process with human review.
  • Ten categories, UK English onlyIt doesn’t tag card numbers, vehicle registrations, passport numbers or health details.
  • Small models sometimes reword textThe integrity check flags this, but it can’t detect a tag that swallows a few neighbouring words.
  • Synthetic data has limitsReal documents contain styles our tests don’t cover. Test it on your own data first.

For your organisation

Need a model like this for your own data?

We can adapt this approach to your documents, your categories of sensitive data and your systems, running in your own environment.

Start a conversation

Tell us the problem.
We’ll tell you honestly if AI can solve it.