Modern Data

Master Data Starter

Starting models and matching rules for customer, product, and supplier data.

GIST stepGroundIterate
Download the PDF
What you get

One trusted record for each customer, product, and supplier.

Get the Master Data Starter overview

The problem it solves, what is included, how it works, the technical components, and how we adapt it with you.

The problem

The same customer lives in five systems, five different ways.

Names are spelled differently, emails carry tags, phones come in four formats, and two different people share a name and a ZIP code. MDM programs spend months configuring match rules before anyone sees whether they work.

What this accelerator does

The Master Data Starter tunes match rules against a labeled sample of your own records in the first cycles, so you see real results early and prove the rules before configuring a platform.

What's included

Match rules you can prove before you commit.

Four parts, each adapted to your data, platforms, and controls.

01

Domain data models

Customer, product, supplier, and reference data: golden records, cross-references, lineage, stewardship, and change log.

02

Match, merge, and survivorship rules

Rules in plain configuration files, tuned against a labeled sample of your records.

03

Stewardship workflow

A review queue for uncertain pairs. Steward decisions override the score on every later run.

04

Scorecards and data contracts

Data quality scorecards, measured match precision and recall, and a contract template per source.

You keep the models, the tuned rules, the labeled sample, and a precision check in CI that fails any rule change causing false merges.

How it works

How matching works.

  1. Map and standardize

    Names folded and nicknames expanded, emails normalized, phones in E.164, addresses and identifiers cleaned.

  2. Block

    Only records that share a key, such as email, phone, or last name and ZIP, are compared.

  3. Score

    Each field has a comparator and a weight. The score is the weighted average over shared fields.

  4. Classify

    Match, review, or no match by threshold. Hard rules override, such as never merging different tax IDs.

  5. Steward decisions

    A data steward decides the review pairs. Those decisions hold on every later run.

  6. Cluster and survive

    Matches become golden records. Survivorship picks each value, and lineage records its source.

Technical detail

Under the hood.

Vendor-neutral Python and configuration, Azure first, with tests included from the start.

Matching engine
Standardization, blocking, scoring, clustering, and survivorship in Python
Comparators
Exact, Jaro-Winkler, token set, name tokens, and numeric
Survivorship
Most recent, source priority, most frequent, and most complete
mdm evaluate
Precision and recall against a labeled sample; fails CI below a set floor
mdm explain
Field-by-field scores for any pair of records
SQL models
Warehouse landing design for golden records, cross-references, and lineage
Proof in the package

Worked example: 359 customer records from three systems.

Records from a CRM, an ERP, and an online store describe 150 real customers, with the mess real data has.

Planted in the data
  • CRM duplicates in capitals, with no email
  • ERP names as “Last, First” with nicknames
  • Store logins with +shop tags and gmail dots
  • Phones in four formats, some mistyped
  • Two different James Smiths in one ZIP code
Result

With the rules as shipped, the engine merges with no false merges. Mistyped-phone pairs go to the stewardship queue, and steward decisions raise recall on the next run. The two James Smiths stay separate.

Where it fits

Platforms, related accelerators, and limits.

Works with

  • Informatica MDM, Reltio, Profisee, Semarchy, and others
  • Custom builds in your warehouse
  • Spark for large volumes

Pairs with

  • The Data Readiness Assessment finds the duplicates and join gaps
  • Golden records land on the Data Platform Foundation
  • Clean master data feeds the Metrics Layer and models

Assumptions and limits

  • The engine suits tuning and up to a few hundred thousand records; larger volumes run on the platform or Spark
  • Standardization is US-style; other markets need country rules
  • Clustering is transitive, so the scorecard flags clusters to check by hand
Ask Datagist AI

Have a question about the Master Data Starter?

Get an answer from our pages in seconds, with links to the sources.

Take the overview with you.

Share the overview with your team, or tell us the decision you want to improve and we will tell you whether it fits.