Skip to main content
Skip to content
Generative AI

Local LLM Deployment

An open model installed on your own hardware, inside your perimeter. Nothing leaves the building, and no vendor decides what your assistant can do next.

01 / 04

Some data cannot go to a cloud model

Patient records, claims files, personnel data. The contract says where they may go, and a cloud API is not on the list. And the per-token bill grows with every user.

CH.01 · The problem

Some data cannot go to a cloud model.

  1. The contract says it stays here

    Patient records, claims files, personnel data, tender documents. Your legal team has already said where they may go, and a cloud API is not on the list.

  2. The bill grows with use

    Per-token pricing looks cheap in a pilot. At a thousand users and a document per call, it is a line item nobody forecast.

  3. The vendor sets the terms

    Models are retired, prices change, a region moves. Your assistant changes with them, on their schedule.

Our answer

Run the model where the data is

An open model, installed on your servers or your private cloud, tuned for your hardware and connected to your systems. You choose the model, the version and the day it changes.

CH.02 · What we build

How we build it

Three parts: the model and hardware, the serving layer, and the connections.

Nothing leaves the perimeterYour model, your calendarSized to your hardwareConnected to your systems
01

Model and hardware

The right open model for your tasks and your GPUs.

  • Test candidate models on your real tasks
  • Size the hardware to the expected load
  • Quantise where it saves memory without losing accuracy
  • Benchmark latency and throughput on the target machine
  • Document the choice and what would change it
02

Serving

The model as a service inside your network.

  • Deploy an inference server with an API your apps can call
  • Batch and cache to make the hardware go further
  • Add authentication and per-team limits
  • Monitor load, latency and errors
  • Plan the upgrade path for the next model
03

Integration

The assistant your teams use, on the local model.

  • Connect documents and systems under existing permissions
  • Add retrieval where answers need your knowledge
  • Keep the option of a no-retention cloud route for public tasks
  • Log usage for cost and quality review
  • Hand over the runbook to your operations team
CH.03 · How it runs

How it runs

Four phases from a list of tasks to a model serving your teams.

  1. 01Phase 1

    Discovery

    The tasks, the data rules, the hardware you have.

    • List the tasks the model must do
    • Confirm which data may not leave
    • Inventory the GPUs and servers available
  2. 02Phase 2

    Model selection

    Candidates tested on your tasks, on your machine.

    • Shortlist open models by task and licence
    • Run them on a sample of real work
    • Measure quality against a reference
  3. 03Phase 3

    Deployment

    Serving, security, monitoring.

    • Install the inference server
    • Add authentication and limits
    • Connect the first application
  4. 04Phase 4

    Handover

    Your team runs it.

    • Write the runbook: start, stop, upgrade, roll back
    • Train the operations team
    • Agree the model review cadence
CH.04 · What changes

What changes

Compliance by construction, a cost you can forecast, and control of the calendar.

The data stays where it must

No prompt or document crosses the perimeter. The compliance argument is the network diagram.

The cost is the hardware

After the machine is paid for, a call costs electricity. Heavy use gets cheaper per call, the opposite of an API.

No surprise from a vendor

You pick the model version and the day it changes. Retirements, price changes and region moves happen to other people.

Side by sideCloud APILocal model
Where the data goesTo the vendor, under their agreementNowhere; it stays on your network
CostPer token, grows with useThe hardware, then electricity
Model changesOn the vendor's scheduleOn yours
CustomisationPrompting, some fine-tuningAny tuning, any version, any quantisation
Compliance evidenceA contract and a region settingThe network diagram
CH.05 · Questions

Questions

It depends on the model size and the number of users at peak. A capable open model for a team can run on one server with a modern GPU; a company-wide assistant needs more. We size it in discovery from your tasks and your expected load, and test on the target machine before you buy anything.

On general tasks, the largest cloud models still score higher. On your tasks, with retrieval from your documents and sometimes a fine-tune, a well-chosen open model reaches the quality your teams need. We measure this on your work in the selection phase, and report the numbers before you commit.

Yes. A local model can be tuned on your examples inside the same perimeter, and served by the same infrastructure. See Fine-tuning AI Models for how that runs.

Open models improve every few months. The serving layer is built so a new version can be tested beside the current one on your tasks, and swapped on a day you choose, with a roll-back. We agree a review cadence at handover.

More in Generative AI

The map of the practice

Back to Generative AI
Start

Bring us the problem nobody has cracked yet.

We are a small team of senior specialists. We pick the right model and the right layer, and we build the least machinery that does the job. You get a call with an engineer, not a sales deck.