Skip to main content
Skip to content
Machine Learning

Federated Learning

Train one model across several data holders, hospitals, branches, partners, without any of them handing over their data. The model travels. The records stay.

01 / 04

The data cannot be pooled

Patient files, account histories, production logs: the law or the contract says they stay. Each site alone has too few cases to learn from. A central training set is off the table.

CH.01 · The problem

The data that would make the model good cannot be pooled.

  1. The records may not leave

    Patient files, account histories, production logs. The law, the contract or the board says they stay where they are. A central training set is off the table.

  2. Compliance stops the project

    A data sharing agreement across three organisations takes a year and a legal team. Most projects do not survive the wait.

  3. Each site alone is too small

    One hospital has a few hundred cases of the condition. Five hospitals together would have enough to learn from. Apart, none of them can.

Our answer

Move the training to the data

Each site trains the model on its own records, inside its own perimeter. Only the model updates travel, combined in a way that reveals none of them. The result is one model that learned from all the sites, and no record ever moved.

CH.02 · What we build

How we build it

Three parts: training at each site, combining the updates securely, and sending the result back.

Data stays at each siteUpdates combined, never exposedCompliance built inOne model from many sites
01

Local training

The model runs at each participant, on their machine.

  • Install the training client inside each site's perimeter
  • Align data formats without moving data
  • Train on local records for a round
  • Keep the raw records on site, always
  • Log each round for the site's own audit
02

Secure aggregation

The coordinator combines what it cannot read.

  • Collect updates from the sites each round
  • Combine them so no single update is visible
  • Add noise where a stricter guarantee is required
  • Handle sites that drop out mid-round
  • Verify the combined update before it is used
03

Global model distribution

Everyone gets the improved model back.

  • Send the combined model to every site
  • Evaluate at each site on local held-out data
  • Compare against a site-only baseline
  • Repeat rounds until the gain flattens
  • Release the model to production at each site
CH.03 · How it runs

How it runs

Four phases, from the participants around the table to a model at every site.

  1. 01Phase 1

    Assessment

    Participants, data, rules, the question.

    • Define the prediction and how it will be used
    • Confirm each site's data and its constraints
    • Agree the protocol and the privacy guarantee
  2. 02Phase 2

    Infrastructure

    Clients at each site, coordinator in the middle.

    • Deploy the training client at each site
    • Set up the coordinator and secure aggregation
    • Align schemas without moving data
  3. 03Phase 3

    Model development

    Rounds, evaluation, repeat.

    • Run training rounds across the sites
    • Evaluate at each site on held-out data
    • Compare with the site-only baseline
  4. 04Phase 4

    Production

    The model at every site, monitored.

    • Release the model at each site
    • Monitor performance per site
    • Schedule new rounds as data grows
CH.04 · What changes

What changes

A model trained on data that could never be pooled, with the compliance argument written into the design.

The project can start

No data sharing agreement, because no data is shared. Legal reviews a protocol instead of a transfer.

A better model than any site could train

Measured at each site against a model trained on its own data alone. The gain from the other sites is a number, per site.

Each participant keeps control

A site can see what it contributed, pause, or leave. Its records never left its building, so there is nothing to recall.

Side by sideCentral trainingFederated learning
Where the data goesTo a central storeNowhere
PrivacyDepends on the store's securityRecords never leave; updates combined unseen
ComplianceData sharing agreements, transfers, residencyA protocol review
Data the model seesWhat could be transferredEvery participating site
InfrastructureOne large training environmentA client per site, a coordinator in the middle
CH.05 · Questions

Questions

A way to train one model on data held in several places without moving the data. Each place trains on its own records and sends back only the model update. A coordinator combines the updates into a better model and sends it back. Repeat until the model stops improving. No record leaves the place it lives.

First, raw data never leaves a site. Second, the updates are combined with secure aggregation, so the coordinator sees only the sum, never one site's contribution. Where a stricter guarantee is needed, noise is added to the updates. The protocol and the guarantee are written down and reviewed by each participant before a round runs.

Health, where patient records stay in each hospital. Financial services, where account data stays in each institution. Manufacturing groups with plants in several jurisdictions. Any case where several holders of similar data want one model and none of them can share the data.

Usually close, and always better than what any single site can train alone, which is the comparison that matters when pooling is not an option. We measure it at each site against the site-only baseline and report the numbers per site.

More in Machine Learning

The map of the practice

Back to Machine Learning
Start

Bring us the problem nobody has cracked yet.

We are a small team of senior specialists. We pick the right model and the right layer, and we build the least machinery that does the job. You get a call with an engineer, not a sales deck.