Skip to main content
Skip to content
Machine Learning

Decentralized Data Training

Train on data spread across plants, branches, subsidiaries or countries without building the central data lake first. The training goes to each site; the model comes back.

01 / 04

The data is everywhere and cannot be brought together

Plants, branches, subsidiaries, countries: each with its own system, owner and jurisdiction. Copying it all to one store is a transfer to justify, a store to protect and a year to wait.

CH.01 · The problem

The data is everywhere, and it cannot be brought together.

  1. Silos with rules of their own

    Each plant, branch or subsidiary holds its data under its own system, its own owner and often its own jurisdiction. Getting read access to all of them takes a programme of its own.

  2. Moving it is the risk

    Copying personal or confidential data to a central store creates a new transfer to justify, a new store to protect, and a new residency question in every country involved.

  3. Moving it is also slow

    Terabytes of logs or images from twenty sites over ordinary links. By the time the lake is full, the question has changed.

Our answer

Take the training to each site

A training client at each site learns from local data. A coordinator combines the learning into one model and sends it back. No copy, no transfer, no lake. The sites keep their data and their rules, and the group gets one model that knows all of them.

CH.02 · What we build

How we build it

Three parts: the network of clients, the data at each site, and the privacy of what travels.

No central store to buildEach site keeps its rulesOnly model updates travelSites can join and leave
01

Network architecture

Clients at the sites, a coordinator in the middle.

  • Install a training client inside each site's perimeter
  • Set up the coordinator where the group allows
  • Authenticate and encrypt every connection
  • Schedule rounds around each site's load
  • Handle sites that join, pause or drop
02

Data preparation

Aligned without being moved.

  • Agree a common schema across sites
  • Map each site's data to it locally
  • Check quality per site and report it
  • Hold out local test data at each site
  • Keep raw data on site, always
03

Privacy mechanisms

What travels reveals nothing.

  • Secure aggregation so no single update is visible
  • Differential privacy where the guarantee requires it
  • Bounds on each update against poisoning
  • Logs at each site for its own audit
  • Protocol reviewed by each site's security team
CH.03 · How it runs

How it runs

Four phases, from the sites on a list to a model at every one of them.

  1. 01Phase 1

    Assessment

    Sites, data, rules, the question.

    • List the sites and their data systems
    • Confirm each site's rules and owner
    • Define the prediction and how it is used
  2. 02Phase 2

    Network set-up

    Clients, coordinator, connections.

    • Deploy the client at each site
    • Set up the coordinator and secure aggregation
    • Map local data to the common schema
  3. 03Phase 3

    Model development

    Rounds across the sites.

    • Run training rounds
    • Evaluate at each site on local held-out data
    • Compare with the site-only baseline
  4. 04Phase 4

    Deployment and scaling

    The model everywhere, new sites joining.

    • Release the model at each site
    • Monitor per-site performance
    • Onboard new sites with the client
CH.04 · What changes

What changes

One model across the group, without the migration and without the transfer.

The project starts this quarter

No data lake to build first. Clients installed, schema agreed, first round run: weeks, not the year a central platform takes.

Every jurisdiction is satisfied

Data stays in the country and the system it lives in. The compliance question per site is answered by the design.

A model that knows the whole group

Patterns from every plant, branch or subsidiary, measured at each site against a model trained on its own data alone.

Side by sideCentral trainingDecentralized training
Data movementEverything to one storeNone; updates only
PrivacyA second copy to protectRecords stay on site, updates masked
ComplianceA transfer per jurisdictionSatisfied by the design
Working across organisationsNeeds a sharing agreementNeeds a protocol
Who controls the dataThe central platformEach site, still
CH.05 · Questions

Questions

Usually close, and it is the comparison that matters only when a central store is possible. When it is not, the comparison is with what each site can train alone, and the decentralized model is better by a margin we measure per site on local held-out data and report before release.

Groups with data in several plants, branches, subsidiaries or countries: manufacturing, retail chains, banking and insurance groups, health networks, public bodies with regional offices. And consortia of separate organisations that want one model without sharing data, which is federated learning across companies.

Every connection is authenticated and encrypted. Updates are combined with secure aggregation, so the coordinator sees only the sum. Where the guarantee requires it, differential privacy noise is added. Each update is bounded against poisoning. Each site logs its rounds for its own audit, and the protocol is reviewed by each site's security team before the first round.

A machine inside its perimeter able to run the training client on its local data: for most tabular and log data an ordinary server, for images or large models a GPU. Plus an outbound connection to the coordinator. The coordinator runs wherever the group allows, on premise or in a private cloud. We size it in assessment from the data and the model.

More in Machine Learning

The map of the practice

Back to Machine Learning
Start

Bring us the problem nobody has cracked yet.

We are a small team of senior specialists. We pick the right model and the right layer, and we build the least machinery that does the job. You get a call with an engineer, not a sales deck.