Entity Resolution Examples
Three real-world scenarios show how to combine fuzzy comparisons, exact comparison helpers, and hard blocking boundaries.
Each scenario shows the Source Records that would appear in the resulting Shape Anomaly. Source Records surface one example row per distinct value of the distinction field within each non-compliant cluster, alongside the cluster identifier _qualytics_entity_id so the cluster boundaries are visible.
The situation: Your customers table is the master record for downstream billing. Each row has a customer_id that should be the single identifier per customer, but historic ingestions from multiple sources have produced near-duplicate records with slightly different spellings of the same person's full_name and address. You want Entity Resolution to surface customers where two different customer_id values plausibly describe the same person.
Check configuration
| Field | Value |
|---|---|
| Rule | Entity Resolution |
| Distinction Field | customer_id |
| Block Fields | (none) |
| Compare Fields | full_name (String, fuzzy, pair_substrings: true, weight: 1.0), address (String, fuzzy, weight: 0.8) |
| Composite Match Threshold | 0.75 |
| Filter | (none) |
| Custom Anomaly Description | Off |
| Status | Active |
| Owner | (check creator) |
| Anomaly Assignee | (customer-data steward) |
| Tags | pii, master-data |
| Additional Metadata | jira: DATA-4101 |
| Description | Customers with similar names and addresses must share a customer_id |
Payload
{
"description": "Customers with similar names and addresses must share a customer_id",
"rule": "entityResolution",
"fields": [],
"container_id": 145,
"filter": null,
"properties": {
"distinct_field_name": "customer_id",
"composite_match_threshold": 0.75,
"target_fields": [
{
"type": "String",
"field_name": "full_name",
"role": "compare",
"comparison_type": "fuzzy",
"pair_substrings": true,
"pair_homophones": false,
"consider_term_frequency": false,
"weight": 1.0
},
{
"type": "String",
"field_name": "address",
"role": "compare",
"comparison_type": "fuzzy",
"pair_substrings": false,
"pair_homophones": false,
"consider_term_frequency": false,
"weight": 0.8
}
]
},
"tags": ["pii", "master-data"],
"additional_metadata": {"jira": "DATA-4101"},
"anomaly_message_field": null,
"template_id": null,
"status": "Active",
"owner_id": 7,
"default_anomaly_assignee_id": 12
}
Source Records
| _qualytics_entity_id | customer_id | full_name | address |
|---|---|---|---|
| ent-a01f | 1001 | Alice Cohen | 142 Maple St |
| ent-a01f | 1057 | Alice C. | 142 Maple Street |
| ent-b73c | 1102 | Catherine Wu | 87 Elm Avenue |
| ent-b73c | 1184 | Catherine Wu | 87 Elm Ave. |
What gets flagged
Two non-compliant clusters appear in the source records:
ent-a01fresolved two records the platform considers the same customer ("Alice Cohen, 142 Maple St"↔"Alice C., 142 Maple Street"). The fuzzy match onfull_namereaches the threshold becausepair_substringspromotes"Alice C."against"Alice Cohen", and the address pair is near-identical. The cluster holds two differentcustomer_idvalues (1001and1057), so it is non-compliant.ent-b73cresolved two records with identical names and only a punctuation difference inaddress. The cluster holds two differentcustomer_idvalues (1102and1184), so it is also non-compliant.
Each non-compliant cluster contributes one row per distinct customer_id to the Source Records panel, four rows total in this scan.
Shape Anomaly
184 records were resolved to 173 distinct entities (composite threshold 0.75: full_name (w=1.0), address (w=0.8)). 2 of those entities are assigned more than one value of customer_id
Flowchart
graph TD
A["Filter: none, evaluate all customers"] --> B["Score every candidate pair on full_name + address"]
B --> C{"Composite score ≥ 0.75?"}
C -->|No| D["Pair is not a match"]
C -->|Yes| E["Connect both records in the same cluster"]
E --> F["Assign cluster _qualytics_entity_id"]
F --> G{"Cluster has more than one customer_id?"}
G -->|No| H["Cluster is compliant"]
G -->|Yes| I["Flag cluster. Source Records gets one row per distinct customer_id."]
The situation: Your businesses table combines records from several vendor feeds. Business names can have spelling variations, but similar names in different industries or countries often represent unrelated companies. You want industry and country to influence the decision without preventing a rare cross-category or cross-country match when the other evidence is overwhelming.
Check configuration
| Field | Value |
|---|---|
| Rule | Entity Resolution |
| Distinction Field | business_id |
| Block Fields | (none) |
| Compare Fields | business_name (String, fuzzy, substring and homophone matching, weight: 4.0), industry (String, exact, weight: 2.0), country (String, exact, weight: 1.0) |
| Composite Match Threshold | 0.85 |
| Filter | (none) |
| Custom Anomaly Description | Off |
| Status | Active |
| Owner | (check creator) |
| Anomaly Assignee | (business-master steward) |
| Tags | consolidation, vendor-feeds |
| Additional Metadata | jira: DATA-4207 |
| Description | Similar businesses in matching contexts should share a business_id |
Payload
{
"description": "Similar businesses in matching contexts should share a business_id",
"rule": "entityResolution",
"fields": [],
"container_id": 212,
"filter": null,
"properties": {
"distinct_field_name": "business_id",
"composite_match_threshold": 0.85,
"target_fields": [
{
"type": "String",
"field_name": "business_name",
"role": "compare",
"comparison_type": "fuzzy",
"pair_substrings": true,
"pair_homophones": true,
"consider_term_frequency": false,
"weight": 4.0
},
{
"type": "String",
"field_name": "industry",
"role": "compare",
"comparison_type": "exact",
"weight": 2.0
},
{
"type": "String",
"field_name": "country",
"role": "compare",
"comparison_type": "exact",
"weight": 1.0
}
]
},
"tags": ["consolidation", "vendor-feeds"],
"additional_metadata": {"jira": "DATA-4207"},
"anomaly_message_field": null,
"template_id": null,
"status": "Active",
"owner_id": 7,
"default_anomaly_assignee_id": 18
}
Source Records
| _qualytics_entity_id | business_id | business_name | industry | country |
|---|---|---|---|---|
| ent-c4d1 | 5001 | Catherine's Books | Retail | US |
| ent-c4d1 | 5042 | Katherine's Books | Retail | US |
| ent-c4d1 | 5108 | Catherines Books LLC | Retail | US |
What gets flagged
One non-compliant cluster appears. The name variations resolve through homophone and substring matching, while industry and country both agree. The cluster holds three different business_id values (5001, 5042, and 5108), so it contributes three rows to Source Records.
Consider two other records named "ACME Boxing" and "ACME Boxes". If one is in Entertainment in the US and the other is in Manufacturing in Canada, both exact context comparisons score 0.0. Even a perfect name score would produce only 4 / 7, or about 0.57, so the pair stays below the 0.85 threshold. Because industry and country are comparison helpers rather than block fields, a pair can still match when stronger evidence from other comparison fields is available.
Shape Anomaly
2,341 records were resolved to 2,296 distinct entities (composite threshold 0.85: business_name (w=4.0), industry (w=2.0), country (w=1.0)). 1 of those entities is assigned more than one value of business_id
Flowchart
graph TD
A["Filter: none, evaluate all businesses"] --> B["Score business_name, industry, and country"]
B --> C{"Composite score ≥ 0.85?"}
C -->|No| D["Pair is not a match"]
C -->|Yes| E["Connect both records in the same cluster"]
E --> F["Connected components collapse transitive chains<br/>(A↔B and B↔C become {A,B,C})"]
F --> G["Assign cluster _qualytics_entity_id"]
G --> H{"Cluster has more than one business_id?"}
H -->|No| I["Cluster is compliant"]
H -->|Yes| J["Flag cluster. Source Records gets one row per distinct business_id."]
The situation: Your contacts table is multi-tenant. The same email is allowed to repeat across tenants (different people, different organizations) but never within a single tenant. You want to resolve contacts within each tenant by full_name and email, and tenant_id should act as a hard boundary so cross-tenant collisions never trigger an anomaly.
Check configuration
| Field | Value |
|---|---|
| Rule | Entity Resolution |
| Distinction Field | contact_id |
| Block Fields | tenant_id (Numeric, exact) |
| Compare Fields | full_name (String, fuzzy, pair_substrings: true, weight: 1.0), email (String, fuzzy, weight: 1.0) |
| Composite Match Threshold | 0.8 |
| Filter | status = 'active' |
| Custom Anomaly Description | Off |
| Status | Active |
| Owner | (check creator) |
| Anomaly Assignee | (ingestion on-call) |
| Tags | multi-tenant, contacts |
| Additional Metadata | jira: DATA-4311 |
| Description | Within a tenant, contacts with similar name and email must share a contact_id |
Payload
{
"description": "Within a tenant, contacts with similar name and email must share a contact_id",
"rule": "entityResolution",
"fields": [],
"container_id": 318,
"filter": "status = 'active'",
"properties": {
"distinct_field_name": "contact_id",
"composite_match_threshold": 0.8,
"target_fields": [
{
"type": "Numeric",
"field_name": "tenant_id",
"role": "block",
"comparison_type": "exact"
},
{
"type": "String",
"field_name": "full_name",
"role": "compare",
"comparison_type": "fuzzy",
"pair_substrings": true,
"pair_homophones": false,
"consider_term_frequency": false,
"weight": 1.0
},
{
"type": "String",
"field_name": "email",
"role": "compare",
"comparison_type": "fuzzy",
"pair_substrings": false,
"pair_homophones": false,
"consider_term_frequency": false,
"weight": 1.0
}
]
},
"tags": ["multi-tenant", "contacts"],
"additional_metadata": {"jira": "DATA-4311"},
"anomaly_message_field": null,
"template_id": null,
"status": "Active",
"owner_id": 7,
"default_anomaly_assignee_id": 24
}
Why the blocking field matters
Because tenant_id has the Block role with exact comparison, the platform never compares a contact in tenant 7 against a contact in tenant 12. Two contacts named "Jane Doe" with the same email on different tenants are treated as separate entities and never cluster together. Blocking on tenant_id is both a correctness guarantee and a performance optimization: candidate pairs are constrained to rows that share the same tenant.
Source Records (filtered to status = 'active')
| _qualytics_entity_id | tenant_id | contact_id | full_name | |
|---|---|---|---|---|
| ent-7a2b | 7 | c-991 | Jane Doe | jane.doe@acme.com |
| ent-7a2b | 7 | c-1042 | J. Doe | jane.doe@acme.com |
The contact c-2071 (tenant_id = 12, full_name = "Jane Doe", email = "jane.doe@acme.com") does not appear in the Source Records: it is in a different tenant, so blocking prevents it from being paired with the rows in tenant 7. It is its own cluster, with its own _qualytics_entity_id, and is compliant.
Shape Anomaly
4,820 records were resolved to 4,791 distinct entities (blocked on [tenant_id], composite threshold 0.8: full_name (w=1.0), email (w=1.0)). 1 of those entities is assigned more than one value of contact_id
Flowchart
graph TD
A["Apply filter: status = 'active'"] --> B["Block pairs by tenant_id<br/>(records in different tenants never compared)"]
B --> C["Score remaining pairs on full_name + email"]
C --> D{"Composite score ≥ 0.8?"}
D -->|No| E["Pair is not a match"]
D -->|Yes| F["Connect both records in the same cluster (per tenant)"]
F --> G["Assign cluster _qualytics_entity_id"]
G --> H{"Cluster has more than one contact_id?"}
H -->|No| I["Cluster is compliant"]
H -->|Yes| J["Flag cluster. Source Records gets one row per distinct contact_id."]
Related
- Introduction: formal definition, field roles, comparison types, and field scope.
- How It Works: blocking, weighted comparison, clustering, Recipe guidance, and golden-set behavior.
- API: current public payload and field reference.
- FAQ: short answers to common configuration and remediation questions.