Performance assessment
Where the model is wrong, and why.
- Measure accuracy per site on held-out data
- Find the segments where it fails
- Compare the data mix across sites
- Identify what other sites have that this one lacks
- Set the target per site
A model trained at one site learns one site's patients, machines or customers. It is confident on those and wrong on the rest. The missing cases exist, in other silos it cannot reach.
One hospital sees its region's patients. One plant sees its machines. A model trained there is confident on those and wrong on the rest, and does not know it.
The cases your model has never seen exist, in the sister site, the partner, the other subsidiary. They would fix the gap, and they are unreachable.
The data cannot be moved: privacy, contracts, residency. The usual answer, gather everything and retrain, is the one thing you may not do.
Federated training lets the model learn from every site's cases without any record moving. We measure accuracy at each site before and after, against a model trained on its own data alone, and report the difference per site.
Three parts: know where the model fails, get the data it needs without moving it, and improve against a measured baseline.
Where the model is wrong, and why.
The cases the model is missing, reached in place.
Rounds, measured against the baseline.
Four phases, each ending with a number per site.
Baseline, failures, targets.
Aligned in place.
Rounds across the sites.
Prove it, then keep it going.
A model that is right more often, on more kinds of case, with a number per site to prove it.
The segments the assessment named, measured before and after at each site. The improvement is a per-site number.
No record moved. The privacy guarantee agreed in assessment holds through training and is documented for each site.
New sites join, data grows, rounds repeat. The model improves on a schedule, and each round is measured.
| Side by side | One site's data | Trained across sites |
|---|---|---|
| Variety of cases | One population | Every participating site |
| Privacy | Local, by default | Local still; updates masked |
| On new data | Confident and often wrong | Measured on held-out data from every site |
| Skew | One site's habits baked in | Diluted across sites, measured |
| Improvement over time | Retrain when someone remembers | Rounds on a schedule, each measured |
It depends on how much the other sites' data differs from yours and where the current model fails. We do not promise a number before measuring. The assessment reports accuracy per site and per segment on the current model; the optimisation phase reports the same after each round. The improvement is that difference, per site.
It usually does. Quality and coverage are checked per site before training, and a site's participation is weighted by what it adds. Sites with poor quality in a segment contribute less to it. The per-site evaluation shows whether a site is helping or hurting, and the schema mapping is fixed at the source when it is the cause.
It can reduce the bias that comes from one site's population dominating the training data, because more populations contribute. It does not remove bias that all sites share. We measure accuracy per segment at each site before and after, so the change in bias is a number, and we say which kind of bias it addresses.
Held-out test data at each site that the training never sees. The current model and the federated model are both evaluated on it, per site and per segment. An independent check runs before release. The report shows before and after for every site, and the privacy guarantee that held throughout.
We are a small team of senior specialists. We pick the right model and the right layer, and we build the least machinery that does the job. You get a call with an engineer, not a sales deck.