Skip to content

Lineage Sources

Every lineage connection in Qualytics records how it was created. There are four categories, each with different rules for how connections are created, updated, and removed:

Category source_type value Created by Updated by Removed by
Manual manual A user in the Lineage tab The same user (or anyone with edit access) A user explicitly deleting the edge
Data catalog import The data catalog name (atlan, alation, collibra, datahub, purview) A data catalog sync The next sync of the same data catalog The next sync of the same data catalog, if the relationship is gone from the data catalog
Source collection The datastore type (snowflake, bigquery, databricks) A Sync operation run with Collect lineage turned on A later Sync operation run with the same option A user explicitly deleting the edge. Collection itself never removes one
Qualytics managed inferred A computed container being created or updated, or an operation writing a table into an enrichment datastore A subsequent update to the computed container The computed container being deleted, or its definition no longer referencing the source. Lineage for a written table goes when that table is deleted

Connections from different categories coexist without conflict. A container can carry connections from all four at once, and one category never overwrites another. Each data catalog also reconciles independently, so connections imported from Atlan are untouched by an Alation sync and vice versa.

Manual

Manual connections are added directly by a user in the Lineage tab through the Add menu. Both endpoints can be containers or fields.

Manual connections are never modified by background processes. Once you create one, it stays until a user with edit access removes it explicitly.

This is the right choice when:

  • You have lineage knowledge that lives outside any data catalog (for example, a pipeline orchestrated by a custom tool).
  • You want to override or supplement data catalog-imported lineage with connections that the data catalog does not capture.
  • You are documenting field-level relationships that the source systems do not expose.

Data Catalog Import

When a data catalog integration syncs, Qualytics pulls lineage relationships from the data catalog and creates connections tagged with the data catalog's name. The data catalog is the source of truth for these connections, so each sync rebuilds them to match the current state.

Reconciliation is scoped to the data catalog that owns the connection. Each sync:

  • Adds connections that exist in the data catalog but not yet in Qualytics.
  • Removes connections that previously came from the same data catalog but are no longer present.
  • Leaves untouched every connection from a different source. Manual and Qualytics managed connections, and connections imported from other data catalogs, are never affected.

Cross-datastore Lineage

Data catalog import preserves relationships that span multiple datastores. A common pattern is a medallion architecture where the same logical entity is materialized at multiple stages:

  • customer_bronze lives in Datastore A.
  • customer_silver lives in Datastore B.
  • customer_gold lives in Datastore C.

The data catalog knows that the bronze table feeds the silver table, which feeds the gold table, even though each materialization lives in a separate Qualytics datastore. The import keeps those cross-datastore connections intact, so the lineage graph for customer_gold can walk all the way back to customer_bronze regardless of where each container is registered.

Note

Cross-datastore connections only exist when both sides have been registered as containers in Qualytics. If customer_bronze exists in the data catalog but has not been added as a container in any Qualytics datastore, that relationship is skipped.

Supported Data Catalogs

Lineage import is available for every data catalog integration that exposes lineage relationships in its API. For details on configuring each one, see the integration page below.

No Data Catalog
1. Alation Alation
2. Atlan Atlan
3. Collibra Collibra
4. DataHub DataHub
5. Microsoft Purview Microsoft Purview

Source Collection

A Sync operation can read lineage out of the source datastore itself, so relationships the source already knows about arrive without a data catalog in between. Collection is off by default and chosen per run, with Collect lineage in the sync settings, and it starts once the catalog part of the sync has succeeded. A datastore whose type has no collector ignores the option. See Sync Operation for where the option sits.

Two kinds of relationship are read:

  • Query history: which tables a statement read in order to write another one. Some sources keep their own record of that, which Qualytics reads directly. Where they do not, Qualytics parses the statement text the history holds.
  • View definitions: the tables a view is defined over.

What Each Datastore Exposes

Datastore Recorded history Parsed statement text View definitions
Snowflake SNOWFLAKE.ACCOUNT_USAGE.ACCESS_HISTORY SNOWFLAKE.ACCOUNT_USAGE.QUERY_HISTORY INFORMATION_SCHEMA.VIEWS
BigQuery None INFORMATION_SCHEMA.JOBS, in the dataset's region INFORMATION_SCHEMA.VIEWS
Databricks system.access.table_lineage system.query.history information_schema.views

The connection Qualytics uses has to be able to read whichever of these a collection relies on, so a run that returns nothing is worth checking against this table before anything else.

Snowflake also publishes its own per-object lineage graph, which Qualytics reads alongside the relations above. It answers what feeds a table right now, rather than what was seen during a window, which is what a source that has been running for years needs: history only reaches back as far as the source retains it.

A Databricks datastore on hive_metastore has no recorded lineage of its own. It reads the parsed history instead, and takes view definitions from SHOW CREATE TABLE.

Collection Never Removes a Connection

A collected connection is added or refreshed, never deleted. History proves that a relationship was exercised inside the window a run read, and proves nothing about one it did not see: a table that simply was not written during that window leaves no trace, and reconciling on that silence would drop a connection every time a pipeline paused. A relationship that is genuinely gone has to be deleted from the Lineage tab.

Assets That Are Not in Qualytics

A statement usually names tables that are not registered in Qualytics. Rather than discard the relationship, Qualytics keeps the connection whenever one side resolves to a container, and shows the other side under the name the source spelled it with. A relationship with neither side in Qualytics is skipped. See Reading the Graph for how one of these appears on the canvas.

Once that asset is registered, by a Sync that brings the table in, the next collection replaces the external side with the container rather than adding a second connection beside it.

Qualytics Managed

Qualytics managed connections are generated automatically when a computed container is created or updated, and when an operation writes a table into an enrichment datastore. They give you lineage for the assets that Qualytics itself produces, without requiring a manual entry or a data catalog sync.

Three computed container types produce Qualytics managed connections:

  • Computed table: Sources are detected from the table's SQL definition.
  • Computed join: Every container configured on the join becomes a source: the left and right containers of a pairwise join, or each declared source of a SQL-mode join.
  • Computed file: The source container configured on the computed file becomes the source.

Qualytics managed connections stay in sync with the computed container definition. When you update a computed table and remove a source, the corresponding connection is removed. When you add a new source, a new connection appears.

Note

If a source is not detected automatically, you can still add the connection manually.

Tables Qualytics Writes

Qualytics also records lineage for the tables it writes into an enrichment datastore, with nothing to configure. A materialize operation connects the table it writes back to the container it materialized, and a scan connects a remediation table back to the container it scanned. Both are recorded at container level and at field level, so a field in the written table points at the field it came from.

Only a table holding the rows of a single container carries this lineage, since it is the only case where naming one source is meaningful.

These connections are kept rather than reconciled against the current configuration: the rows did flow, and that stays true after a later materialize stops targeting the same table. They go when the written table is deleted.

Choosing a Source

You do not choose the source directly. It is determined by how the connection is created:

  • Adding a connection from the Lineage tab produces a manual connection.
  • Running a data catalog sync produces connections tagged with that data catalog's name.
  • Running a Sync with Collect lineage turned on produces connections tagged with the datastore's type.
  • Saving a computed container, or running an operation that writes a table into an enrichment datastore, produces Qualytics managed connections.

If you need a connection that the data catalog does not export and Qualytics does not detect automatically, add it manually.

See Also