Skip to main content
Skip to content
Generative AI

Fine-tuning AI Models

A model trained on your own examples, so it writes in your terms and follows your rules without a page of instructions on every call.

01 / 04

The prompt has become a manual

Tone, format, the rules of your domain and three examples, sent again with every call. The model still misses your vocabulary, and your archive of good answers has never been shown to it.

CH.01 · The problem

The prompt has become a manual.

  1. Every call carries the same instructions

    Tone, format, the twenty rules of your domain, three examples. All of it is sent again with every request, and paid for every time.

  2. The model still misses your vocabulary

    Your product codes, your clause numbering, the way your analysts write a summary. A general model gets close and no closer.

  3. Your examples stay outside the model

    Thousands of good answers sit in your archive. The model has never seen one of them.

Our answer

Teach the model with your examples

Fine-tuning trains a model on pairs of input and the answer you want. The rules move from the prompt into the weights. The prompt gets short, the answers get closer to yours.

CH.02 · What we build

How we build it

Three parts. The data comes first, and most of the work is there.

Your domain, learnedCloser answersFewer invented detailsYour knowledge stays yours
01

Training data

Good examples, cleaned, and enough of them.

  • Collect real input and answer pairs from your archive
  • Remove personal data and anything you cannot keep
  • Fix the answers the team would not sign today
  • Hold out a test set the model never sees
  • Balance the set across the cases that matter
02

Training

The base model, the method, the runs.

  • Choose a base model that can run where you need it
  • Use parameter-efficient tuning when the data is small
  • Run several configurations and keep the results
  • Compare each run against the held-out set
  • Check the model still handles what it handled before
03

Serving

The tuned model in front of your users.

  • Deploy on your hardware or a no-retention route
  • Version the model with its training set
  • Keep the base model available to roll back
  • Log inputs and outputs for the next round
  • Set the retraining trigger with your team
CH.03 · How it runs

How it runs

Four phases, with a number at the end of each.

  1. 01Phase 1

    Discovery

    The task, the examples you have, and the gate.

    • Write the task the model must do
    • Count and inspect the examples available
    • Measure a prompted baseline on a sample
  2. 02Phase 2

    Data

    Build the training set and the test set.

    • Extract pairs from the archive
    • Anonymise and clean
    • Review a sample with the domain owner
  3. 03Phase 3

    Training

    Runs, evaluation, the pick.

    • Run the first configuration
    • Evaluate against the baseline on the held-out set
    • Adjust data or method and run again
  4. 04Phase 4

    Production

    Serve, watch, hand over.

    • Deploy behind the same interface as before
    • Compare live answers with the baseline for a week
    • Enable for all users
CH.04 · What changes

What changes

Shorter prompts, closer answers, and a model that is yours.

The prompt shrinks

Instructions and examples move into the weights. Each call sends the task and little else, which costs less and runs faster.

The answers match yours

Measured on the held-out set against the answers your team wrote. The number is reported before anyone relies on it.

The model belongs to you

Weights, training set and evaluation are handed over. You can retrain, move or retire it without us.

Side by sidePromptingFine-tuned model
Your vocabularyExplained in every promptLearned from the examples
EffortDays to write, endless to maintainWeeks to build, then retrain on a cadence
Accuracy on your taskWhatever the prompt reachesMeasured against your answers, gated
Tokens per callInstructions plus examples plus taskThe task
What you keepA promptThe weights, the data, the evaluation
CH.05 · Questions

Questions

It depends on the task. A narrow task with a consistent format can show a measurable gain with a few hundred clean pairs. A broad task needs thousands. We count what you have in discovery and measure a prompted baseline first, so the decision to tune rests on a number.

Training runs on your infrastructure or on a route that does not retain data. The training set is anonymised before it is used, and the weights it produces are yours. Nothing from your examples reaches a shared model.

Your domain drifts: new products, new rules, new phrasing. We log inputs and outputs, and retrain on a cadence you set, typically when the live accuracy drops below the gate. Each version is kept with its training set so any release can be rolled back.

Yes. Open base models can be tuned and served inside your perimeter. We choose the base model in discovery with your hardware in mind, so the tuned model runs where the data is allowed to be.

More in Generative AI

The map of the practice

Back to Generative AI
Start

Bring us the problem nobody has cracked yet.

We are a small team of senior specialists. We pick the right model and the right layer, and we build the least machinery that does the job. You get a call with an engineer, not a sales deck.