Skip to content

Data Diff Check FAQ

Answers to common questions about how the Data Diff check matches rows between target and reference datasets, how it handles missing or differing values, and how anomalies are reported, grouped by topic.

Behavior

What is the difference between Data Diff and Is Replica Of?

Data Diff (dataDiff) is the current rule type for two-table comparison, including checks previously called Is Replica Of. The API still accepts isReplicaOf for compatibility with older clients and treats it as dataDiff. Use dataDiff in new API requests and reporting queries. Data Diff supports Row Identifiers, Passthrough Fields, Comparators, and diff_change_types, which restricts which diff statuses fire anomalies.

What are Row Identifiers, and do I need them?

Row Identifiers are the compound key the platform uses to pair each target row with a reference row. Without them, the platform falls back to a symmetrical set difference and can only produce added and removed diffs (a row that differs in one field becomes one added plus one removed row, not one changed row). Setting Row Identifiers is the only way to get per-field, side-by-side diffs (Left vs Right on the same row) in the Comparison Source Records view. See How It Works → Row Identifiers and Passthrough Fields.

What are Passthrough Fields?

Passthrough Fields are extra columns carried into the source-records output for context. They appear alongside the diffed fields in the Comparison Source Records view but are not themselves compared, so they never cause the anomaly to fire. Typical use: showing a customer_name or created_at next to the differing column so anomaly triagers can identify the row without leaving the page.

Does the filter clause scope both target and reference?

Each side has its own clause. The check's Filter Clause narrows the target container; the Filter Clause in the Right Reference panel (properties.ref_filter) narrows the reference. A side with no clause is read in full. When both sides must cover the same slice, set both and keep them aligned, otherwise the scope mismatch shows up as added and removed rows. See How It Works → The Filter Clauses.

How are Comparators applied?

Comparators apply a per-field tolerance to the equality check. The platform supports Numeric, Duration, and String Comparators. Without a Comparator the values are compared strictly: 1.00 and 1.000001 differ, "Australia" and "australia" differ. Use Comparators when small, expected divergences (rounding, casing, whitespace) would otherwise produce noise. See How It Works → Comparators.


Anomaly Reporting

What do the row statuses mean?

Status Meaning
added The identifier exists only on the target (left). The target has a row the reference does not.
removed The identifier exists only on the reference (right). The target is missing a row the reference has.
changed The identifier exists on both sides, but at least one compared field has a different value.

Can I exclude added (or removed, or changed) rows from anomalies?

Yes. The diff_change_types property (inside properties) takes a list of statuses to allow: any subset of ["added", "removed", "changed"]. Rows with a status outside the list still get computed but do not contribute to the anomaly. Set it to ["added", "changed"] when the reference is intentionally a superset of the target (a staging table with extra QA rows, which arrive as removed), or to ["changed"] when row presence is guaranteed by an upstream contract and only value drift matters. The property defaults to all three statuses; an empty list is rejected at the API. See How It Works → Restricting Anomalies by Status.

Why do some rows show "missing" instead of a value?

When a row is added, the reference side has no row to read, so every right-side cell renders as the literal text missing. When a row is removed, the target side has no row to read, so every left-side cell renders as missing. This is the same rendering the Comparison Source Records view uses in the Qualytics app.

Why doesn't Data Diff produce a Record Anomaly?

Data Diff is a shape-only rule type: the violation is a property of the target as a whole (the set of rows that diverge from the reference), not of an individual record's value. The per-row detail you see is attached to the Shape Anomaly as Comparison Source Records, not as separate Record Anomalies.

What does the Shape Anomaly message look like?

N records differ between <container> (<datastore>) and <reference> (<reference datastore>)

Each container is named with its datastore in parentheses, the target first and the reference second. N is the number of differing rows, counting only the statuses Diff Change Types allows. When the comparison cannot run row by row, the message says why instead: that the row counts differ, that one side is empty, that a compared field is missing, or that a compared field is stored as a different type. See Anomaly Reporting for the full set.


Configuration

Can I compare two containers in the same datastore?

Yes. Set ref_datastore_id to the same datastore ID as the target's datastore and ref_container_id to the second container. The reference does not have to be in a separate datastore.

Does Custom Anomaly Description (the anomaly_message_field payload field) work for Data Diff?

No. The Custom Anomaly Description toggle (and the corresponding anomaly_message_field payload field) only affects Record Anomaly messages. Because Data Diff emits only Shape Anomalies, the field is silently ignored at evaluation time, and the resulting message uses the fixed Shape Anomaly template (see What does the Shape Anomaly message look like?).

Which fields can I edit on an existing Data Diff check?

In the UI, open the check and change what you need, then click Update; see Edit a Check for the full steps. Through the API, a PUT to /api/quality-checks/{id} updates the description, filter, tags, status, ownership, the compared fields list, every Data Diff property (Row Identifiers, Passthrough Fields, diff_change_types, Comparators), and the reference container. The rule type and the target container itself are fixed at creation. See the API page for the full editable/immutable matrix.

What permission do I need to create or edit a Data Diff check?

The Drafter team permission on the datastore covers Draft work (creating a check as Draft or editing it while it stays Draft). Anything that puts the check into evaluation, such as creating or editing an Active check, archiving, or deleting, requires the Author team permission (or above). Viewing only requires Reporter. See Permissions for the full matrix.