Skip to main content
Skip to content
Machine Learning

Privacy-Preserving Machine Learning

Models trained and run on sensitive data with a measurable privacy guarantee: differential privacy, federated training, secure computation, chosen for your data and your regulator.

01 / 04

The useful data is the data you may not use

Health, finance, HR: a regulator with specific rules, a model that could memorise a record, partners who cannot pool. Deleting the name column satisfies none of it.

CH.01 · The problem

The useful data is the data you are not allowed to use.

  1. The regulation is specific

    GDPR, sector rules, a data protection officer with questions. Anonymising a spreadsheet by deleting the name column does not satisfy any of them.

  2. The data is sensitive by nature

    Health, finance, HR, minors. A model that memorises a record can be made to repeat it, and a breach of a model is a breach of the data.

  3. Partners cannot pool

    Two banks, three hospitals, a supplier and a customer. Each has half the picture, and no lawyer will sign the transfer that would complete it.

Our answer

A guarantee you can write down

We pick the technique that fits the data and the rule: differential privacy for a bound on what any record can leak, federated learning to train without moving data, secure computation for partners who cannot share. Each comes with a number the DPO can read.

CH.02 · What we build

How we build it

Three parts: choose the technique, implement it properly, and design the system around it.

Differential privacyFederated trainingSecure computationChosen for your constraint
01

Technique selection

The right guarantee for the data and the rule.

  • Classify the data and the regulation that applies
  • Define the threat: who could learn what
  • Measure the accuracy the use case needs
  • Compare techniques on your data, on a sample
  • Record the choice, the guarantee and the trade-off
02

Implementation

The guarantee, done correctly.

  • Use audited libraries for the privacy mechanisms
  • Set the privacy budget with the DPO
  • Track the budget across training runs
  • Test that the model cannot reproduce records
  • Document the mechanism for the auditor
03

Privacy-aware architecture

The system around the model respects the same rules.

  • Least access for every component
  • Encrypt data at rest, in transit and, where needed, in use
  • Log every access to sensitive data
  • Retention and deletion built into the pipeline
  • Review with security and legal before release
CH.03 · How it runs

How it runs

Four phases, with the DPO in the room from the first.

  1. 01Phase 1

    Assessment

    Data, rules, threats, accuracy needed.

    • Classify the data and the applicable rules
    • Define the threat model
    • Measure the accuracy the use case requires
  2. 02Phase 2

    Technique implementation

    The mechanism, built and verified.

    • Implement with audited libraries
    • Configure the privacy budget
    • Build the training pipeline around it
  3. 03Phase 3

    Model development

    Accuracy within the guarantee.

    • Train under the privacy constraint
    • Measure accuracy on held-out data
    • Tune the trade-off with the business owner
  4. 04Phase 4

    Validation and governance

    Proof, then ongoing control.

    • Independent review of the guarantee
    • Release with monitoring of the budget
    • Set the review cadence with the DPO
CH.04 · What changes

What changes

Models on the data that matters, with a guarantee instead of a hope.

The project clears legal

A privacy budget, a protocol and an architecture the DPO can review. The question changes from whether to how much.

A stolen model keeps the data safe

Tested: the model cannot be made to repeat a record. What it learned is patterns, within a bound you set.

Partners can build together

Federated or secure computation lets organisations train one model without any of them seeing the others' data.

Side by sideOrdinary MLPrivacy-preserving ML
Data handlingCopied to a training environmentTrained in place or under a bound
What the model can leakUnknown; it may memorise recordsBounded by the privacy budget, tested
ComplianceArgued after the factA protocol reviewed before training
Working with partnersNeeds a data transferFederated or secure computation
If the model is stolenThe data may be at riskThe bound still holds
CH.05 · Questions

Questions

Usually some, and how much depends on the technique and the budget. We measure it: the model under the guarantee against an unconstrained baseline on the same held-out data, and the trade-off is tuned with the business owner and the DPO together. Federated training on its own often costs little; strong differential privacy costs more.

It follows from the data, the rule and the threat. One organisation with sensitive data and a regulator: differential privacy, often with federated training across its own sites. Several organisations that cannot share: federated learning with secure aggregation, or secure computation when the joint result itself is the product. The assessment tests the candidates on your data.

The mechanisms come from audited libraries, the privacy budget is tracked across every run, and the finished model is attacked: membership inference and extraction tests try to recover records from it. The results go to the DPO with the documentation. Where required, an independent reviewer checks the guarantee.

Mostly, yes. Differential privacy is a change to the training loop. Federated training adds a client at each site and a coordinator. Secure computation needs its own runtime. In every case we build on the pipelines and platforms your team already runs, and say in assessment what has to be added.

More in Machine Learning

The map of the practice

Back to Machine Learning
Start

Bring us the problem nobody has cracked yet.

We are a small team of senior specialists. We pick the right model and the right layer, and we build the least machinery that does the job. You get a call with an engineer, not a sales deck.