01 — The Signal 02 — Challenges & Gaps 03 — From the Field 04 — Vendor Reality Check 05 — The AI Advantage 06 — Reference Architecture 07 — Core Concepts Redesigned 08 — Migration Roadmap 09 — Day in the Life 10 — References
Point of View · July 2026

Master Data Management was built for a world that no longer exists.

Batch windows. Deterministic rules. Licence meters that charge by the master row. Stewardship queues nobody can clear. Traditional MDM solved the 2010 problem beautifully — and it is now the single largest bottleneck between enterprises and their agentic AI ambitions.

This is a practitioner's argument for what replaces it, how it is architected, what it costs, and how you get there without a big-bang rewrite.

Read the thesis ↓
$12.9MAverage annual cost of poor data quality per organisation (Gartner)
15%Of organisations have the data foundation agentic AI requires (HBR-AS / Reltio)
~37%F1 score of classic fuzzy matching on a standard entity-resolution benchmark
95%Precision achievable using an LLM as adjudicator on borderline match pairs
R2 Digital orbital logomark
The transformation
From To
Rules author the truthHumans encode every match, survivorship and validation rule up front
Models propose the truthSemantic matching, learned survivorship, evidence from public data
Batch reconciles itNightly windows; the golden record is always yesterday's
Events settle itStreaming resolution; the golden record is current by construction
Stewards absorb residueUnbounded task queues; the backlog is the operating model
Stewards adjudicateHumans govern policy and edge cases, not volume

01 — The Signal

The analyst community has quietly reversed itself on MDM

Six months of published research tells a coherent story: MDM was written off as legacy plumbing, then abruptly reinstated as strategic infrastructure — because generative and agentic AI cannot function without it.

APR 2026

Gartner brings the MDM Magic Quadrant back

After retiring the MQ, Gartner reinstated it in April 2026 (Kennedy, Robison, Radhakrishnan), naming Profisee, Salesforce (Informatica), Reltio, Stibo Systems and Semarchy as Leaders. The stated rationale is the arrival of agentic and augmented data management — MDM as the trusted context layer that prevents enterprise AI hallucination.

2025 – 2026

The readiness gap is the headline finding

Gartner puts the average annual cost of poor data quality at $12.9M and warns that a majority of AI projects risk abandonment without AI-ready data practices; only about 37% of organisations express confidence in their data management for AI. A Harvard Business Review Analytic Services survey cited by Reltio finds just 15% possess the data foundation agentic AI needs.

NOV 2025

The vendor landscape consolidated underneath customers

Salesforce completed its acquisition of Informatica on 18 November 2025, absorbing catalog, governance, quality, metadata and MDM into the Salesforce platform. For an installed base of 9,000+ organisations running heterogeneous SAP / Oracle / Snowflake estates, roadmap independence is now a live procurement risk, not a hypothetical.

2026 R1 · AGENTFLOW

Incumbents are bolting AI onto the same core

Profisee 2026 R1 ships automated agents, semantic matching and an MCP server; Reltio's AgentFlow Unstructured ingests PDFs, Word and HTML to surface "dark" data. Directionally right — but these are capabilities layered onto hub architectures, commercial models and stewardship paradigms designed for a rules-and-batch world.

RESEARCH

The matching evidence has moved decisively

On standard record-linkage benchmarks, classic fuzzy matching lands near 37% F1; pure embedding retrieval improves recall but suffers precision (~35%). Hybrid designs — embedding-based candidate generation with an LLM adjudicating only borderline pairs — reach 95%+ precision at negligible per-pair cost. This is the single most consequential fact in modern MDM.

DISSENT

Not everyone is convinced the MQ captured it

Independent commentary (notably from the open-source entity-resolution community) argues the reinstated MQ under-weights ML-native resolution, cost-per-record economics and the rise of composable, lakehouse-resident mastering. We include the counter-argument deliberately: a point of view that only cites confirming evidence is marketing, not analysis.

02 — Challenges & Gaps

Ten structural failures of traditional MDM

These are not product defects. They are the logical consequences of an architecture designed when compute was expensive, data was structured, and "real time" meant overnight.

GAP 01 · DETERMINISM

Truth is hard-coded by humans, up front

Every match rule, survivorship hierarchy and validation constraint must be anticipated and encoded before the first record lands. Reality then arrives in a shape nobody predicted — a new source format, a default value stuffed into a key field, a merger that doubles your naming conventions — and the rule set degrades silently.

  • Rule sets grow to thousands of conditions no one person understands
  • Every new source is a multi-week configuration project
  • Rules cannot express "this looks like the same company because both filed under the same DUNS in 2019"
GAP 02 · LATENCY

The golden record is always yesterday's record

Batch match-and-merge windows mean operational systems consume a version of truth that is hours to days stale. Fine for a monthly report. Fatal for an AI agent making a credit decision, routing a claim, or answering a customer live.

  • Nightly windows that no longer fit inside the night
  • Reprocessing full loads to correct a single upstream error
  • No credible path to sub-second entity resolution at the point of decision
GAP 03 · ECONOMICS

Licensing punishes the thing you're trying to do

When you are metered per master record, per domain, per profile or per opaque processing unit, every instinct to master more data — more domains, more history, more external enrichment — becomes a budget conversation. The commercial model is structurally hostile to the data completeness AI requires.

  • Overage clauses that convert growth into unplanned spend
  • Teams deliberately excluding data to stay under tier thresholds
  • Non-production environments metered as if they were production
GAP 04 · TCO OPACITY

You cannot forecast next year's bill

Consumption units that are not clearly defined in compute or storage terms make annual budgeting a negotiation rather than a calculation. Public analysis of Informatica's IPU model notes it "isn't clear what precisely an IPU is," with third-party commentary citing budget overruns of 40–60% against initial projections. Treat vendor-adjacent figures with appropriate scepticism — but the pattern is consistent across buyers we have worked with.

GAP 05 · STEWARDSHIP

The backlog is the operating model

Low-confidence matches fall to a human queue. Because rules cannot generalise, the queue fills faster than any team can clear it — and it fills with near-identical tasks. Vendor patent literature itself acknowledges "a large number of very similar tasks in the queue for the attention of a data steward," causing downstream delay in master data updates.

  • Stewards make the same judgement hundreds of times; the system never learns it
  • Most organisations staff stewardship part-time (10–20% of a role), guaranteeing the queue never drains
  • No feedback loop from steward decision back into matching logic
GAP 06 · MODALITY

Blind to 80–90% of enterprise information

Contracts, onboarding PDFs, support transcripts, regulatory filings, emails and scanned KYC packs contain the strongest identity evidence in the enterprise — and traditional MDM cannot read any of it. Mastering runs exclusively on structured columns while the corroborating evidence sits unused in a document store.

GAP 07 · EXTERNAL TRUTH

Enrichment is a purchased feed, not a reasoning act

Traditional enrichment means buying a reference file and joining on a key. It cannot reconcile that a customer's registered entity changed name after acquisition, that two addresses are the same building, or that a "new" prospect is an existing subsidiary — facts freely available in public registries, filings and web sources.

GAP 08 · EXPLAINABILITY

"Why did these merge?" has no good answer

Deep rule chains are technically deterministic yet practically inscrutable. Under GDPR, DORA, BCBS 239 or model-risk review, "rule 4,412 fired" is not a lineage narrative. Ironically, a well-instrumented AI pipeline that records the evidence and confidence behind each decision produces a better audit trail than the rules it replaces.

GAP 09 · TIME TO VALUE

Eighteen months to the first business outcome

Data model workshops, source profiling, rule authoring, UAT, cutover. By the time the hub is live the sponsoring executive has moved on and the business case has been re-litigated twice. Enterprise MDM programmes routinely spend $150K–$300K on implementation services in year one alone, before licence.

GAP 10 · LOCK-IN

Your logic is trapped in proprietary metadata

Match rules, survivorship, workflows and cleansing logic live in vendor-specific configuration stores with no meaningful export. Combined with platform consolidation — Salesforce absorbing Informatica being the clearest recent example — this converts a technology choice into a decade-long strategic dependency.

03 — From the Field

An anonymised enterprise MDM architecture, and what it tells us

The critique above is not theoretical. Below is a real multi-domain Informatica MDM programme R2 Digital delivered — customer, membership, certification and product domains across a dozen source systems. It was a well-executed implementation, faithful to the vendor's reference architecture. Its constraints are exactly the constraints of the era it was built in.

As-builtLayered hub topology — Landing → Staging → Base Objects → ViewsInformatica MDM / IDQ on AWS-RDS
Landing areaSource-shaped tables · truncate & load daily · no transformation, by design
Staging areaCleansing, DQ rules, reference lookups · 1-to-1 mappings
MDM Hub (Base Objects)3NF · daily incremental · tokenize → match → merge · Best Version of Truth
Data services view layerSIF / BES / DB extract for downstream consumption
Governance toolingCollibra for metadata, lineage and stewardship workflow
DQ & reportingIDQ scorecards, Address Doctor, Tableau DQ dashboards
As-builtEight enterprise integration patternsBatch · real-time · asynchronous
1 · Source extractsBatch, daily — the workhorse
2 · Downstream writesBatch to DWH / data marts
3 · Pub/Sub outMDM publishes updates to queue
4 · App → MDM syncReal-time service
5 · Lookup / XREFReal-time cross-reference
6 · Two-way RT syncBidirectional, loopback-guarded
7 · MDM → App syncReal-time push
8 · Pub/Sub inApps publish, MDM subscribes
TELL 01

The cache workaround is the architecture confessing

The design included a documented alternate solution: "if there are performance or HA issues, deploy an MDM/RDM cache on every business application, refreshed at low latency." That is a sophisticated, correct engineering answer — and it is also an admission that a batch-anchored hub cannot reliably serve real-time reads at enterprise concurrency. In the AI-native design, real-time resolution is the primary path, not the fallback, and the cache disappears along with the six replicas of truth it created.

TELL 02

"Truncate and load daily" was a deliberate choice, and it capped the ceiling

Landing and staging were both full daily reloads to keep MDM decoupled from source systems. Entirely reasonable in 2019. It also means the golden record's freshness ceiling is 24 hours no matter how much real-time integration sits on top — and no agentic use case tolerates that. Streaming CDC into an open table format removes the constraint without giving up the decoupling.

TELL 03

The Trust Framework was the right idea, statically configured

Informatica's Trust Framework assigns a trust factor per column per source — genuinely ahead of its time, and the conceptual ancestor of learned survivorship. But those factors were set by hand in workshops and then rarely revisited. The AI-native version keeps the concept and makes the weights observed: measured against outcomes, recalibrated continuously, and explainable per attribute.

TELL 04

Six tools, five consoles, one steward

Hub Console, Data Director, DQ Analyst, DQ Developer, Admin Console, Address Doctor, plus Collibra and Tableau. Each is competent; collectively they demand a specialist per console. This is why MDM programmes carry the headcount they do — and why the single steward cockpit in the AI-native design is a genuine operating-model change, not a UI preference.

TELL 05

The governance model was excellent — and it survives intact

Guiding principles, a target operating model spanning consumers, capabilities, technology, processes, governance and roles, a Business Rules Spreadsheet under version control with impact analysis, defined authorised sources, and a change-management chain of accountability. None of this is made obsolete by AI. It is precisely the control framework that makes autonomous mastering defensible. Modernise the engine; keep the governance.

TELL 06

The agile delivery model was already right

The programme was structured as domain-by-domain shippable increments, prioritised by business value, with a PoC preceding the MVP. That sequencing maps almost one-for-one onto the Phase 1–3 migration pattern later in this page. Organisations that already ran MDM this way migrate fastest, because the hard organisational muscle — incremental domain delivery — is already built.

On this material Details are drawn from a prior R2 Digital client engagement and are presented in anonymised, generalised form. Domain names, source systems, volumes and identifying specifics have been removed. Nothing here is a criticism of the client team or of Informatica's engineering — the architecture was appropriate and correctly built for its date. The point is narrower and more useful: the constraints that shaped it are no longer constraints, and an architecture optimised against removed constraints is by definition leaving value on the table.
04 — Vendor Reality Check

Where the leading platforms actually strain

All five Gartner Leaders are competent products built by serious engineering teams. The critique below is architectural and commercial, not a quality judgement — and it reflects what surfaces in the third and fourth year of an implementation, not the demo.

Dimension Informatica MDM (Salesforce) Reltio Profisee Stibo Systems Semarchy
Licence metric Hybrid IPU consumption + per-domain record counts. Enterprise entry commonly quoted around $200K/yr. High friction Subscription by consolidated profile count, tiered, invoiced annually in advance. High friction Volume-based subscription; cited by Gartner as transparent. Lower friction Record/SKU and module-based; PIM-weighted commercials. Medium Deployment + data-volume based; relatively legible. Medium
Master-row & overage exposure Growth in records and processing both consume budget. Third-party analyses report 40–60% overruns vs. initial projections. Severe Profile-count tiers create hard cliffs; teams routinely exclude domains or history to stay under tier. Severe Still volume-metered — mastering "everything" remains a cost decision. Medium Multi-domain expansion is a re-contracting event. Medium Scales more gracefully but still couples volume to price. Medium
Pricing predictability IPU definition is not publicly precise in compute/storage terms — budgeting becomes negotiation. Opaque No public price list; multiple product lines, support tiers and add-ons. Opaque Most legible of the Leaders. Good Quote-driven. Medium Quote-driven, generally simpler. Medium
Matching engine Deterministic + probabilistic rule configuration; tuning is a specialist skill and a standing cost centre. Rules-first ML-assisted matching with a strong graph model; still fundamentally configured rather than learned end-to-end. Hybrid 2026 R1 adds semantic matching — genuinely forward-looking, layered on an existing hub. Hybrid Rules-centric, PIM-optimised. Rules-first Rules with ML assists. Hybrid
Processing model Batch-anchored heritage with streaming bolted alongside; full reprocessing is expensive. Batch-anchored API-first and the most genuinely real-time of the incumbents. Strong Near-real-time; strong Azure/Fabric alignment. Good Batch-oriented for large catalog syndication. Batch-anchored Continuous, incremental. Strong
Unstructured & external evidence Limited native use of documents as identity evidence. Gap AgentFlow Unstructured ingests PDF/DOCX/HTML — the most advanced incumbent answer. Leading MCP server exposes master data to AI tools; ingestion of unstructured evidence less mature. Partial Rich asset handling, weaker identity reasoning from text. Partial Limited. Gap
Stewardship experience Functional workbench; queue-driven, minimal learning from steward decisions. Dated Modern UI, still task-queue centric. Dated paradigm Some administration still requires the FastApp Studio desktop client (web migration signalled for 2026). Friction Deep but complex. Complex Clean, developer-friendly. Good
Deployment constraints Multi-cloud, but heavy footprint and dedicated administrators required. Heavy SaaS-only; limited control over residency specifics. Medium SaaS runs on Microsoft Azure only — a hard constraint for AWS/GCP-standardised enterprises. Constraint Flexible. Good Flexible, incl. customer-managed cloud. Good
Connectivity breadth Broadest connector library in the market. Strength API-first; good coverage. Good Smaller out-of-the-box connector library than integration-led vendors. Gap Strong retail/manufacturing ecosystems. Good Native integration capability noted by Gartner as lagging competitors. Gap
Strategic / roadmap risk Salesforce completed the acquisition Nov 2025; independent commentary expects R&D to prioritise Agentforce/Data Cloud integration over standalone multi-platform capability. Elevated Independent; capital-intensive growth model. Medium Independent; deep Microsoft dependency. Medium Privately held, stable. Low Independent. Low
Portability of your logic Proprietary metadata; exit cost is effectively a re-implementation. Locked Proprietary config. Locked Proprietary config. Locked Proprietary config. Locked More code-like, better versioned. Partial
← scroll horizontally to compare all five vendors →
Methodological note Vendor pricing in this market is not publicly listed. Figures above are drawn from published analyst commentary, vendor blogs, procurement-advisory sites and buyer review platforms — several of which are themselves commercially interested. They are indicative ranges for framing a business case, not quotations. Gartner-attributed strengths and cautions are as reported in secondary coverage of the April 2026 Magic Quadrant; the full report is behind Gartner's paywall. Where we have seen a pattern repeatedly in delivery, we say so as experience, not as data.
05 — The AI Advantage

Accuracy, efficiency and cost are not a trade-off any more

The traditional MDM business case forced a choice: tighter matching meant more steward hours; broader coverage meant a bigger licence. AI-native design breaks all three constraints at once — and the mechanism is specific, not hand-waving.

ACCURACY

Semantics beat string distance

An embedding of "Intl. Business Machines Corp, Armonk NY" sits near "IBM Corporation, 1 New Orchard Rd" in vector space. No edit-distance function achieves that. Layer an LLM adjudicator over only the ambiguous band and precision on published benchmarks moves from the ~37% F1 range of fuzzy matching to 95%+ precision — at pennies per thousand adjudications.

  • Multilingual and transliterated names handled natively
  • Corporate hierarchy inferred, not just matched
  • Every decision carries a confidence score and an evidence trail
EFFICIENCY

The queue learns instead of growing

Each steward decision becomes a labelled training example. Active learning routes the most informative ambiguous cases to humans rather than an arbitrary confidence cut-off, so a few hundred decisions retune the model where thousands of rule edits previously could not. The queue shrinks structurally rather than through headcount.

  • Near-duplicate tasks collapsed and resolved as a cohort
  • Rule authoring replaced by example-based teaching
  • New source onboarding measured in days, not quarters
COST

Decouple price from record count

The largest saving is not model efficiency — it is the removal of a per-master-row licence meter. When mastering runs on infrastructure you already own (lakehouse compute, your own vector index, open-weight or API models on measured inference), cost tracks usage, not the size of your customer base. Enrichment you would previously have bought as a feed becomes a reasoning step over public sources.

  • No overage cliffs; no tier-driven data exclusion
  • Implementation services materially reduced — less to configure
  • Non-production environments effectively free

Indicative three-year cost shape

Illustrative, normalised to 100 at year-1 traditional spend. Shape, not magnitude — actual numbers depend entirely on record volume, domain count and estate complexity.

1601301007040 Year 0 (build)Year 1Year 2Year 3 licence + overage compute + inference
Traditional MDM — licence escalation, overage, tuning services, growing steward headcount AI-native MDM — higher year-0 engineering, then declining unit cost as models absorb stewardship volume

The crossover typically lands in the second year. The strategic point is the direction of the curve: traditional MDM gets more expensive precisely as you succeed at it.

06 — Reference Architecture

AI-native MDM inside an AI-first enterprise data architecture

The design principle: master data stops being a destination system and becomes a trust and context service — resolved continuously, exposed through APIs and MCP, and consumed as readily by autonomous agents as by a BI dashboard.

Tier 1Sources — structured and unstructured, treated equallyEverything is evidence
Core systemsERP · CRM · Billing · Core banking · HRIS
DocumentsContracts · KYC packs · Onboarding PDFs · Filings
InteractionsSupport transcripts · Email · Chat · Call notes
Public & third-partyRegistries · D&B / LEI · Postal authorities · Web
Streams & eventsCDC · Kafka · Webhooks · IoT
▼ ▼ ▼
Tier 2Ingestion & Representation LayerContinuous, not batch
Streaming ingestEvent-driven CDC into an open table format
Document intelligenceLayout-aware extraction; entities and attributes from text
Embedding serviceEntity-tuned vectors for names, addresses, products
Vector + graph indexANN candidate retrieval and relationship traversal
Semantic layerBusiness ontology, metrics and domain definitions
▼ ▼ ▼
Tier 3 · CoreAgentic Mastering EngineSpecialised agents, orchestrated and supervised
Profiling agentLearns each source's shape and drift; no manual profiling phase
Standardisation agentNames, addresses, classifications normalised in context
Resolution agentVector blocking → LLM adjudication of the ambiguous band
Survivorship agentLearned attribute-level trust, not fixed source hierarchy
Enrichment agentRetrieves and reconciles public evidence with citation
Quality agentAnomaly detection against learned norms, not static thresholds
Hierarchy agentLegal, ownership and household structures inferred
Explanation agentNarrates every decision for steward and auditor
▼ ▼ ▼
Tier 4Trust, Governance & Human OversightThe control plane — non-negotiable
Confidence & thresholdsPolicy-defined autonomy bands per domain and risk class
Steward cockpitException adjudication with model reasoning surfaced
Active learning loopEvery human decision retrains the resolver
Lineage & evidence storeImmutable record of inputs, prompts, scores, outcomes
Policy & consentGDPR/CCPA purpose limitation, residency, retention
Model risk controlsDrift monitoring, bias testing, challenger models, rollback
▼ ▼ ▼
Tier 5Trusted Context ServicesHow the enterprise consumes truth
Real-time resolution APISub-second "who is this?" at the point of decision
MCP serverMaster data exposed as tools to any AI agent
Golden record storeVersioned, time-travel queryable, open format
Event publicationDownstream systems notified on change, not polled
RAG grounding setCurated entity context that prevents hallucination
▼ ▼ ▼
Tier 6Consumption — human and machineSame truth, both audiences
Operational appsCRM, servicing, order management
Analytics & BILakehouse marts, executive reporting
AI agents & copilotsGrounded on resolved entities via MCP
Regulatory reportingAuditable lineage end to end
Data productsDomain-owned, contract-governed

Why this fits the AI-first enterprise

Agentic systems fail on ambiguity, not on reasoning. If an agent cannot answer "is this the same customer?" deterministically and instantly, every downstream action inherits the doubt. Tier 5 exists precisely so that agents consume resolved entities rather than re-deriving identity from raw tables.

Where the semantic layer matters

Gartner has cautioned that a majority of agentic analytics efforts relying solely on protocol-level access (MCP) will fail by 2028 without semantic foundations. Exposing tables to an agent is not the same as exposing meaning. The ontology in Tier 2 and the resolved entities in Tier 5 are what make MCP useful rather than merely convenient.

What stays human

Policy, risk appetite, autonomy thresholds, edge-case adjudication and accountability. The architecture increases human leverage; it does not remove human authority. Any design that cannot name what remains human is not ready for a regulator.

07 — Core Concepts Redesigned

Every MDM primitive, reconceived

This is where the argument becomes concrete. Below, each classical MDM capability is set against its AI-native equivalent — what changes mechanically, and what it changes commercially.

Traditional
  • Buy a reference data feed; join on a supplied key
  • Enrichment is a scheduled refresh, structurally stale
  • No ability to reconcile conflicting external sources
  • Corporate hierarchy purchased, rarely validated against filings
  • Coverage gaps in emerging markets and long-tail entities are permanent
AI-native
  • Retrieval agent gathers evidence from registries, filings, postal authorities and the open web
  • Model reconciles conflicting sources and states which it trusted and why
  • Every enriched attribute carries a citation and confidence score
  • Legal and ownership hierarchies inferred from documents, then verified
  • Triggered on change events — a filed name change updates the record the same day
Commercial impactDisplaces a material share of third-party enrichment subscription spend and converts it into measured inference cost — while improving freshness from quarterly to near-continuous.
Traditional
  • Country-specific postal libraries, licensed per country
  • Fails on informal, rural, multilingual or newly-built addresses
  • Cannot recognise that "Bldg 4, Tech Park" and "Block D, Technology Park" are the same location
  • Unparseable records dumped into a steward queue
AI-native
  • Language model parses free-form address text in any script or format
  • Geospatial reasoning plus reverse geocoding resolves informal descriptions to a coordinate
  • Sub-building and campus-level disambiguation from context
  • Postal authority validation retained as the final verification step — AI proposes, the authoritative source confirms
  • Deliverability and address-quality scores emitted per record
Practical impactIn our experience the residual "unparseable" queue — historically 3–8% of inbound records in multi-country estates — is where the majority of address stewardship effort concentrates. This is the highest-ROI single intervention in most programmes.
Traditional
  • Analyst hand-writes hundreds of DQ rules per domain
  • Static thresholds that are wrong the moment volume shifts seasonally
  • Detects the violation, cannot propose the correction
  • Rule coverage decays as sources evolve; nobody re-profiles
AI-native
  • Agent profiles each source and proposes rules for human ratification — governance by approval, not authorship
  • Anomaly detection against learned seasonal and contextual norms
  • Remediation suggested with confidence, not just flagged
  • Continuous re-profiling detects schema and semantic drift automatically
  • Rules expressed in natural language, versioned and reviewable by business owners
Governance impactThe scarce resource stops being rule-writing capacity and becomes business judgement. That is a far better use of a data owner's time — and it is auditable in a way a 4,000-line rule set never was.
Traditional
  • Blocking keys, edit distance, phonetic algorithms, weighted rule sets
  • Tuning thresholds is a specialist role and a permanent cost
  • Precision and recall traded against each other by hand
  • Benchmark-level performance for fuzzy matching sits near 37% F1
  • New data patterns cause silent degradation nobody notices for months
AI-native
  • Stage 1 — embedding-based ANN retrieval generates candidates at scale, cheaply
  • Stage 2 — a fast classifier auto-resolves the confident majority
  • Stage 3 — an LLM adjudicates only the ambiguous band, with reasoning recorded; published results reach 95%+ precision at trivial per-pair cost
  • Stage 4 — genuinely uncertain cases escalate to a steward, and that answer retrains the classifier
  • Graph context used as evidence: shared directors, addresses, payment instruments
Why the cascade mattersRunning an LLM over every pair is economically absurd. Running it over the 2–5% of pairs where it actually changes the answer is the entire design insight — and it is what makes the accuracy gain affordable.
Traditional
  • Fixed source-precedence hierarchy: "CRM wins for phone, ERP wins for address"
  • Most-recent-wins, ignoring that recency is not reliability
  • Same rule applied regardless of attribute, geography or record age
  • Changing precedence requires reconfiguration and regression testing
AI-native
  • Attribute-level trust learned from observed historical accuracy per source
  • Context-sensitive: the system learns the EU billing system is authoritative for VAT numbers but not for mobile numbers
  • Contradictions surfaced as explicit conflicts with evidence, not silently overwritten
  • Trust weights recalibrate as source quality changes — no reconfiguration project
  • Every surviving value traceable to the evidence that selected it
Regulatory impactAttribute-level provenance with a stated confidence is materially stronger under GDPR Article 5(d) accuracy obligations and BCBS 239 lineage expectations than "precedence rule 12 selected this value".
Traditional
  • Unbounded FIFO queue of low-confidence matches
  • Hundreds of near-identical tasks; no cohort resolution
  • Steward decisions vanish into the record — the system learns nothing
  • Stewardship staffed part-time (commonly 10–20% of a role), so the queue never drains
  • Reviewer sees two records side by side with no explanation of why they surfaced
AI-native
  • Active learning selects the most informative cases, not merely the least confident
  • Similar exceptions clustered and resolved as a policy decision across a cohort
  • Model reasoning presented in plain language: what evidence, what weight, what alternative
  • Every decision becomes a training example — the queue shrinks by design
  • Steward can teach by example and by natural-language instruction rather than editing rules
Role impactStewardship shifts from transaction processing to policy setting and exception adjudication. That is a promotion in substance, and it is how you retain good stewards.
Traditional
  • Static routing: every change of type X goes to approver group Y
  • Uniform scrutiny regardless of risk — a typo correction and a legal-entity merge follow the same path
  • Approvers become rubber stamps under volume
  • Workflow changes require developer involvement
AI-native
  • Risk-tiered autonomy: low-risk, high-confidence changes auto-commit with post-hoc sampling
  • Routing to the approver with demonstrated domain expertise, not just the right group membership
  • Change summarised with impact analysis — "this merge affects 14 open contracts and 3 regulatory reports"
  • Anomalous approval patterns flagged for governance review
  • Human review concentrated where it changes outcomes
Control impactCounter-intuitively this strengthens control. Uniform approval under high volume produces uniform inattention. Concentrating scrutiny on risk produces genuine review.
Traditional
  • Hierarchies maintained manually or purchased as a static file
  • Perpetually out of date after M&A activity
  • Household and beneficial-ownership structures rarely modelled well
  • One hierarchy per purpose, maintained separately, mutually inconsistent
AI-native
  • Graph inference from transactions, documents, shared attributes and public filings
  • Ownership and control structures extracted from filings and contracts, with citation
  • Multiple simultaneous perspectives — legal, sales, servicing, risk — over one resolved entity graph
  • Change detection: an acquisition announcement triggers hierarchy re-evaluation
Use case impactThis is where KYC/AML, credit exposure aggregation and account-based selling converge. A resolved entity graph serves all three from one investment.
Traditional
  • Manual crosswalk tables between every pair of source code sets
  • Breaks on every source-system upgrade or standard revision
  • Product and industry classification done by hand at scale
  • Unmapped values silently default to "Other", corrupting analytics
AI-native
  • Semantic mapping proposes crosswalks between code sets, human ratifies
  • Zero-shot classification against NAICS/SIC/UNSPSC/GS1 from descriptive text
  • New codes detected on arrival and mapped proactively rather than defaulting
  • Taxonomy drift monitored; divergence between declared and inferred classification flagged
Scale impactClassification is embarrassingly well suited to language models and is often the fastest credible pilot — narrow scope, measurable accuracy, low blast radius.
Traditional
  • Batch extracts pushed to subscribing systems on a schedule
  • Point-to-point integrations proliferate and ossify
  • Consumers re-derive their own logic because the extract does not fit
  • No mechanism for an AI agent to query master data conversationally
AI-native
  • Real-time resolution API — resolve an identity at the moment of decision
  • MCP server exposing entities, hierarchies and lineage as tools to any agent
  • Event publication on change; consumers react rather than poll
  • Curated grounding context that measurably reduces hallucination in enterprise RAG
  • Time-travel queries: "what did we believe about this customer on the day we approved the loan?"
Strategic impactThis is the capability that converts MDM from a cost centre into the enabling layer for every agentic initiative in the enterprise. It is also the easiest ROI story to tell a CFO.
08 — Migration Roadmap

Getting there without a big-bang rewrite

Nobody replaces a production MDM hub in one cutover, and nobody should try. The pattern that works is a strangler-fig migration: run the AI-native layer in shadow, prove it domain by domain, and retire the legacy hub only when it has nothing left to do.

Phase 0

Diagnose & baseline

4 – 6 weeks
  • Extract current match/survivorship logic and translate it into a documented, portable specification
  • Establish a labelled ground-truth set from historical steward decisions
  • Measure incumbent precision/recall honestly — most organisations have never done this
  • Full TCO baseline: licence, overage, services, steward FTE, opportunity cost
Exit criteriaA defensible accuracy and cost baseline, plus a golden test set. Without these, every later claim is unfalsifiable.
Phase 1

Shadow resolution

6 – 10 weeks
  • Stand up embedding, vector index and the resolution cascade against a copy of production data
  • Run in parallel with zero write-back — no production risk whatsoever
  • Compare AI decisions against the legacy hub and against ground truth
  • Publish a disagreement report: where they differ and who was right
Exit criteriaDemonstrated accuracy parity or improvement on the golden set, with the disagreement analysis reviewed by the data owners.
Phase 2

First domain in production

8 – 12 weeks
  • Select one bounded domain — supplier or product usually beats customer as a starter
  • Deploy the steward cockpit and active-learning loop
  • Set conservative autonomy thresholds; expand them on evidence
  • Stand up lineage, evidence store and model-risk controls before go-live, not after
Exit criteriaDomain running in production with measured steward-effort reduction and a clean audit trail.
Phase 3

Coexistence & expansion

4 – 9 months
  • Anti-corruption layer isolates consumers from which engine actually resolved a record
  • Migrate domain by domain with reconciliation at every wave
  • Introduce unstructured evidence and public-data enrichment
  • Begin decommissioning legacy licence tiers as record volumes shift
Exit criteriaMajority of mastered volume served by the new platform; first genuine licence reduction realised.
Phase 4

AI-first activation

Ongoing
  • Expose MCP server and real-time resolution API to enterprise agents
  • Curate grounding context for RAG and copilot use cases
  • Retire the legacy hub and its remaining commercial commitments
  • Continuous model evaluation, drift monitoring and challenger testing as business-as-usual
Exit criteriaMaster data consumed natively by agents; MDM re-positioned as the enterprise trust layer.

Indicative sequencing — 18 months

Diagnose & baseline
6 wks
Shadow resolution
10 wks
Domain 1 production
12 wks
Coexistence waves
domain by domain
Legacy decommission
staged licence exit
AI activation (MCP/API)
ongoing
M1M3M5M6M8M9M11M12M14M15M17M18

Failure mode to avoid

Permanent hybrid. Coexistence without a committed decommission date becomes two platforms, two cost lines and two versions of truth. Set the legacy exit date in Phase 2 and defend it.

Contract timing matters

Sequence the migration against your renewal calendar. The single largest financial lever in the programme is arriving at renewal with credible, demonstrated optionality — not arriving with a slide deck.

Do not migrate the rules

The instinct is to port thousands of legacy match rules. Resist it. Use them as test cases for the new resolver, not as its specification. Porting the rules ports the problem.

09 — Day in the Life

What actually changes for a data steward

Architecture diagrams do not win transformation programmes; the people who have to live inside them do. Here is the same steward, same organisation, before and after.

×Ann Marie — Traditional MDM
Opens a queue of 412 tasks

Up from 380 yesterday. It has never gone down. She sorts by age and starts at the top, knowing she will not reach the bottom.

Reviews near-identical merge candidates

Eleven variations of the same regional distributor. Same judgement, eleven times. The system will present the same eleven again next month when new records arrive.

Escalates an address she cannot parse

A rural Indonesian address the postal library rejects. She emails the local sales team and moves on. It will sit unresolved for nine days.

Cannot explain a merge to Compliance

Asked why two entities merged in March. The answer is somewhere inside a 4,000-condition rule set. She books time with the MDM developer, who is on another project.

Raises a change request for a new source

A newly acquired subsidiary's CRM. Configuration effort is estimated at six weeks. She schedules it for next quarter.

Manually reworks a bad batch merge

Last night's run merged 240 records on a default value stuffed into a tax-ID field. Unmerging is manual and takes the rest of the afternoon.

Queue closes at 448

She cleared 61 tasks. 97 arrived. The backlog grew again.

~80% transactionalQueue growingNo compounding learningDependent on developers
Ann Marie — AI-Powered MDM
Opens 14 curated exceptions

Not the least-confident cases — the most informative ones. Each arrives with the model's reasoning, the evidence considered and the alternative it rejected.

Resolves a cohort with one decision

The eleven distributor variants are clustered into a single decision. She sets the policy once; the system applies it to this cohort and to every future instance.

Reviews an AI-proposed quality rule

The profiling agent detected drift in a new source's date format and drafted a validation rule in plain language. She amends one condition and approves it. No developer involved.

Answers Compliance in four minutes

The lineage view shows the exact evidence, confidence score and reasoning behind the March merge, with citations to the source documents. She exports it as an audit artefact.

Onboards the subsidiary's CRM herself

The profiling agent maps and proposes standardisations; she reviews and ratifies. Live in three days instead of six weeks.

Does the work she was hired for

Sits with the commercial team on customer hierarchy definitions for the new market entry, and tightens the autonomy threshold on a high-risk attribute after reviewing the week's sampling report.

Queue closes at 6

Her decisions retrained the resolver. Tomorrow's queue will be smaller than today's — structurally, not heroically.

~70% judgement & policyQueue shrinkingCompounding learningSelf-service onboarding
An honest note on the human question The natural next question is whether this reduces stewardship headcount. In the engagements we have seen, it usually does not — it changes what the team does. The organisations that treat this as a cost-cutting exercise tend to lose their best stewards, which is precisely the population whose judgement the active-learning loop depends on. The organisations that treat it as capacity release — mastering domains they could never previously justify — get the better outcome on both dimensions.
Work with us

Let's pressure-test this against your estate

A free 45-minute consultation with R2 Digital. We will look at your current MDM footprint, your renewal timeline and your AI roadmap, and give you a straight answer on whether an AI-native approach is worth pursuing — including when it is not.

Email us directly
10 — References

Sources

Published analyst commentary, vendor material and research cited in this analysis. Several sources are commercially interested parties; they are included because they are the primary record of the claims made, not because they are disinterested.

01Gartner Magic Quadrant for Master Data Management Solutions (April 2026)gartner.com — paywalled 022026 Gartner MQ for MDM — Informatica reprint & commentaryinformatica.com 032026 Gartner MQ for MDM Solutions — Profisee reprintprofisee.com 04Reltio recognised as a Leader, April 2026 Gartner MQreltio.com 05Semarchy named a Leader in the 2026 Gartner MQ for MDMbusinesswire.com 06Analysis of 2026 MQ strengths and cautions by vendorqmetrix.com.au 07"The MDM Magic Quadrant Is Back. Here's What Is Missing." — counter-perspectivelearningfromdata.zingg.ai 08Salesforce completes acquisition of Informatica (18 Nov 2025)salesforce.com 09Acquisition completion — Informatica press releaseinformatica.com 10"The End of the Informatica Era" — roadmap-independence critique (vendor-authored)syncari.com 11Informatica volume tier & Flex IPU consumption pricing announcementinformatica.com 12Informatica pricing breakdown — IPU model analysisintegrate.io (commercially interested) 13Informatica pricing guide — implementation cost & overrun estimatesmammoth.io (commercially interested) 14Reltio Connected Data Platform — pricing model overviewtrustradius.com 15Practitioner discussion — Reltio pricing and cost experiencepeerspot.com 16Gartner: Lack of AI-ready data puts AI projects at riskgartner.com 17Gartner on data quality — $12.9M average annual costgartner.com 18Gartner D&A Summit 2026 — context, semantics and agentic AI takeawaysatlan.com 19Reltio AgentFlow Unstructured — dark data & the 15% readiness finding (HBR-AS)reltio.com 20"Trusted data is the foundation for agentic AI" — Reltio at Google Cloud Nextsiliconangle.com 21Profisee 2026 R1 — agents, semantic matching and MCP serverdbta.com 22"Why trusted context is becoming the currency for enterprise AI"infoworld.com 23Informatica World 2026 — CIO takeaways on data management and AIinformationweek.com 24In-context Clustering-based Entity Resolution with LLMs — design space explorationACM SIGMOD / PACMMOD 25LLM as entity-resolution judge — 95% precision benchmark vs. fuzzy matchingTowards AI 26EnsembleLink: Accurate Record Linkage Without Training DataarXiv 27Can we trust LLM Self-Explanations for Entity Resolution?arXiv — explainability caveats 28KcMF: Knowledge-compliant framework for schema and entity matchingarXiv 29Informatica MDM match rule configuration — product documentationdocs.informatica.com 30Strangler Fig pattern — incremental legacy migrationMicrosoft Azure Architecture Center 31The true cost of poor data qualityibm.com 32Gartner Peer Insights — MDM solution reviewsgartner.com
How to read this analysis This is a point of view, not a procurement recommendation. Claims sourced to vendors are labelled as such. Benchmark figures for entity resolution come from published academic and practitioner evaluations on public datasets and will not transfer unchanged to your data. Statements drawn from delivery experience are presented as experience. One widely circulated statistic — that a large majority of MDM tasks will be handled autonomously by late 2026 — is attributed to Gartner in vendor marketing but we could not verify it in primary Gartner material, so it is deliberately excluded from the argument above.