Modernizing MDM for AI workloads
MDM hubs were designed for rules authored upfront, batch reconciliation overnight, and stewards clearing queues. Those assumptions need revisiting when AI agents are among the consumers.
The requirement this answers
“Lead Master Data Management strategy, platform selection and implementation”
Models propose truth, events settle it, and stewards adjudicate policy rather than transactions.
Three shifts
From rules to models. Humans encoding match and survivorship logic upfront gives way to semantic matching and learned, attribute-level survivorship. Traditional fuzzy matching scores around 37% F1 on entity-resolution benchmarks, while an LLM adjudicating the ambiguous band reaches above 95% precision.
From batch to events. Nightly reconciliation means the golden record is at least 24 hours old. Event-driven settlement keeps it current by construction.
From queues to judgment. Stewards move from unbounded task lists to a curated set of high-value exceptions, with the system learning from each decision.
How the economics change
Traditional MDM has an ascending cost curve, because per-master-row licensing rises as you master more data and steward headcount scales with volume. An AI-native approach front-loads engineering investment and then declines as models absorb stewardship volume. Crossover typically occurs in year two, and by year three annual cost runs 30% to 40% below the traditional trajectory.
Every Gartner Leader still prices on volume, which is worth confirming before signing a multi-year renewal.
Migrating without a rewrite
A strangler-fig approach over roughly 18 months. Diagnose and baseline against ground truth. Run shadow resolution in parallel at zero production risk. Put one bounded domain into production with conservative autonomy thresholds. Expand domain by domain behind an anti-corruption layer. Then activate AI-first consumption and decommission the legacy platform.
Use the legacy rules as test cases rather than porting them, and set the legacy exit date during phase two. Running two platforms indefinitely is the most expensive available outcome.
How I apply it
- Establish ground truth and measure the incumbent before proposing a replacement
- Use a four-stage cascade: embedding blocking, fast classifier, LLM adjudication, steward exception
- Time the migration against the renewal calendar for contractual leverage
- Commit a decommission date with an accountable executive
What good looks like
- Match precision and recall are measured against ground truth
- Stewards can explain a merge to a compliance officer in plain language
- Golden records are exposed as a real-time API and event stream
- Unit cost per mastered record is falling