Lineage Sources
Every lineage connection in Qualytics records how it was created. There are four categories, each with different rules for how connections are created, updated, and removed:
| Category | source_type value |
Created by | Updated by | Removed by |
|---|---|---|---|---|
| Manual | manual |
A user, on a container's lineage | The same user (or anyone with edit access) | A user explicitly deleting the edge |
| Data catalog import | The data catalog name (atlan, alation, collibra, datahub, purview) |
A data catalog sync | The next sync of the same data catalog | The next sync of the same data catalog, if the relationship is gone from the data catalog |
| Source collection | The datastore type (snowflake, bigquery, databricks, oracle) |
A Sync operation run with Collect lineage turned on | A later Sync operation run with the same option | A user explicitly deleting the edge. Collection itself never removes one |
| Qualytics managed | inferred |
A computed container being created or updated, or a Materialize or remediation output being written into an enrichment destination | A subsequent update to the computed container | The computed container being deleted, or its definition no longer referencing the source. Lineage for a written output goes when that output is deleted |
Connections from different categories coexist without conflict. A container can carry connections from all four at once, and one category never overwrites another. Each data catalog also reconciles independently, so connections imported from Atlan are untouched by an Alation sync and vice versa.
Manual
Manual connections are added directly by a user through the Add menu on a container's lineage. Both endpoints can be containers or fields.
Manual connections are never modified by background processes. Once you create one, it stays until a user with edit access removes it explicitly.
This is the right choice when:
- You have lineage knowledge that lives outside any data catalog (for example, a pipeline orchestrated by a custom tool).
- You want to override or supplement data catalog-imported lineage with connections that the data catalog does not capture.
- You are documenting field-level relationships that the source systems do not expose.
Data Catalog Import
When a data catalog integration syncs, Qualytics pulls lineage relationships from the data catalog and creates connections tagged with the data catalog's name. The data catalog is the source of truth for these connections, so each sync rebuilds them to match the current state.
Reconciliation is scoped to the data catalog that owns the connection. Each sync:
- Adds connections that exist in the data catalog but not yet in Qualytics.
- Removes connections that previously came from the same data catalog but are no longer present.
- Leaves untouched every connection from a different source. Manual and Qualytics managed connections, and connections imported from other data catalogs, are never affected.
Cross-datastore Lineage
Data catalog import preserves relationships that span multiple datastores. A common pattern is a medallion architecture where the same logical entity is materialized at multiple stages:
customer_bronzelives in Datastore A.customer_silverlives in Datastore B.customer_goldlives in Datastore C.
The data catalog knows that the bronze table feeds the silver table, which feeds the gold table, even though each materialization lives in a separate Qualytics datastore. The import keeps those cross-datastore connections intact, so the lineage graph for customer_gold can walk all the way back to customer_bronze regardless of where each container is registered.
Note
Cross-datastore connections only exist when both sides have been registered as containers in Qualytics. If customer_bronze exists in the data catalog but has not been added as a container in any Qualytics datastore, that relationship is skipped.
Supported Data Catalogs
Lineage import is available for every data catalog integration that exposes lineage relationships in its API. For details on configuring each one, see the integration page below.
| No | Data Catalog |
|---|---|
| 1. | |
| 2. | |
| 3. | |
| 4. | |
| 5. |
Source Collection
A Sync operation can read lineage out of the source datastore itself, so relationships the source already knows about arrive without a data catalog in between. In plain terms, Qualytics reads the record the warehouse itself keeps of which tables were used to build which. Collection is off by default and chosen per run, with Collect lineage in the sync settings, and it starts once the catalog part of the sync has succeeded. For a datastore whose type has no collector, the option cannot be selected in the sync settings or in a Flow's Sync action, and an API request that sets it is ignored. See Sync Operation for where the option sits.
Two kinds of relationship are read:
- Query history: which tables a statement read in order to write another one. Some sources keep their own record of that, which Qualytics reads directly. Where they do not, Qualytics parses the statement text the history holds.
- View definitions: the tables a view is defined over.
The rest of this section is reference for whoever manages the datastore connection.
What Each Datastore Exposes
| Datastore | Recorded history | Parsed statement text | View definitions |
|---|---|---|---|
SNOWFLAKE.ACCOUNT_USAGE.ACCESS_HISTORY |
SNOWFLAKE.ACCOUNT_USAGE.QUERY_HISTORY |
INFORMATION_SCHEMA.VIEWS of each database, with SNOWFLAKE.ACCOUNT_USAGE.VIEWS as a fallback for views whose definition the connection cannot read there |
|
| None | INFORMATION_SCHEMA.JOBS, in the dataset's region |
INFORMATION_SCHEMA.VIEWS |
|
system.access.table_lineage |
system.query.history |
information_schema.views |
|
| None | GV$SQL, recently run statements only |
ALL_VIEWS and ALL_MVIEWS (materialized views) |
The connection Qualytics uses has to be able to read whichever of these a collection relies on, so a run that returns nothing is worth checking against this table before anything else. A Snowflake view connection read through the fallback names SNOWFLAKE.ACCOUNT_USAGE.VIEWS as its source, and that relation can lag the live catalog by up to 90 minutes.
Snowflake also publishes its own per-object lineage graph, which Qualytics reads alongside the relations above. It answers what feeds a table right now, rather than what was seen during a window, which is what a source that has been running for years needs: history only reaches back as far as the source retains it.
A Databricks datastore on hive_metastore has no recorded lineage of its own. It reads the parsed history instead, and takes view definitions from SHOW CREATE TABLE.
Oracle adds three sources of its own:
- Dependency records.
ALL_DEPENDENCIESis Oracle's record of what each view and materialized view reads, so those connections hold even where a view definition cannot be parsed. Qualytics reads it like Snowflake's lineage graph, as a native lineage snapshot. - Stored code. Procedures, functions, package bodies, and triggers stored in the database are parsed for the statements that write tables, so a load that runs inside a package gets lineage even after its statements have aged out of
GV$SQL. Each Sync reads up to 200 of these code units and continues with the rest on the next one. Only code that references a table in the datastore, directly or through a synonym, or that runs dynamic SQL (EXECUTE IMMEDIATE) in one of the datastore's schemas is read. In code declaredAUTHID CURRENT_USER, a table named without its schema is left unresolved, because it depends on who calls the code. Stored code is read through the dictionary views that hold its dependencies, its source, and the rights it runs with:DBA_DEPENDENCIES,DBA_SOURCE, andDBA_PROCEDURESwhen the connection can read them, otherwiseALL_DEPENDENCIES,ALL_SOURCE, andALL_PROCEDURES. If neither set can be read in full, the Sync collects the rest of the lineage without stored code. - Synonyms. A statement that names a table through a synonym is resolved to the table the synonym points to.
Oracle keeps no lasting statement history unless the separately licensed Diagnostics Pack is in use, and Qualytics does not read it. The parsed history is therefore best effort: GV$SQL holds recently run statements and drops them on a database restart or under memory pressure, so loads that run regularly are seen and one-off statements may be missed.
Tables Created Since the Last Sync
When lineage evidence is available, a table that the same Sync brings into Qualytics for the first time gets its lineage in that Sync, instead of waiting for the next one.
Statements That Write to Several Tables
A statement that inserts into several tables at once, such as INSERT ALL or INSERT FIRST, produces a table-level connection into every target, on Snowflake, BigQuery, and Databricks as well as on Oracle. These connections are table-level: a statement that writes to several tables records no field-level lineage.
Collection Never Removes a Connection
A collected connection is added or refreshed, never deleted. History proves that a relationship was exercised inside the window a run read, and proves nothing about one it did not see: a table that simply was not written during that window leaves no trace, and reconciling on that silence would drop a connection every time a pipeline paused. A relationship that is genuinely gone has to be deleted by hand, from a container's lineage.
Assets That Are Not in Qualytics
A statement usually names tables that are not registered in Qualytics. Rather than discard the relationship, Qualytics keeps the connection whenever one side resolves to a container, and shows the other side under the name the source spelled it with. A relationship with neither side in Qualytics is skipped. See Reading the Graph for how one of these appears on the canvas.
Once that asset is registered, by a Sync that brings the table in, the next collection replaces the external side with the container rather than adding a second connection beside it.
Qualytics Managed
Qualytics managed connections are generated automatically when a computed container is created or updated, and when a Materialize or remediation output is written into an enrichment destination. They give you lineage for the assets that Qualytics itself produces, without requiring a manual entry or a data catalog sync.
Three computed container types produce Qualytics managed connections:
- Computed table: Sources are detected from the table's SQL definition.
- Computed join: Every container configured on the join becomes a source: the left and right containers of a pairwise join, or each declared source of a SQL-mode join.
- Computed file: The source container configured on the computed file becomes the source.
Qualytics managed connections stay in sync with the computed container definition. When you update a computed table and remove a source, the corresponding connection is removed. When you add a new source, a new connection appears.
Note
If a source is not detected automatically, you can still add the connection manually.
Outputs Qualytics Writes
Qualytics also records lineage for two of the outputs it writes into an enrichment destination, with nothing to configure:
- A Materialize output is connected back to the container it is a copy of.
- A remediation output is connected back to the container whose anomalous records it holds.
Both are recorded at container level and at field level, since these outputs carry the source container's fields, so a field in the output points at the field it came from.
Only an output holding the rows of a single container carries this lineage, since it is the only case where naming one source is meaningful. The other written outputs, such as scan results or exports, gather rows from many containers and get no connection.
These connections are kept rather than reconciled against the current configuration: the rows did flow, and that stays true after a later Materialize stops targeting the same output. They go when the written output is deleted.
Choosing a Source
You do not choose the source directly. It is determined by how the connection is created:
- Adding a connection on a container's lineage produces a manual connection.
- Running a data catalog sync produces connections tagged with that data catalog's name.
- Running a Sync with Collect lineage turned on produces connections tagged with the datastore's type.
- Saving a computed container produces Qualytics managed connections to every source in its definition, at container level. A join with several inputs gets one connection per input.
- Writing a Materialize or remediation output into an enrichment destination produces a Qualytics managed connection back to the one container it copies, at container and field level.
If you need a connection that the data catalog does not export and Qualytics does not detect automatically, add it manually.
See Also
-
Reading the Graph
Direction, nodes, assets outside Qualytics, where a connection came from, and expanding the graph.
-
Datastore-level Lineage
The datastore map, the lineage between two datastores, and every container of a datastore that has lineage.
-
Field-level Lineage
Expanding field lists, field metadata, and the focal field workflow.
-
Impact Analysis
What a container feeds and what feeds it, ranked across a datastore and exportable as CSV.
-
Examples
Enrolling a warehouse, a medallion pipeline across datastores, a remediation output, and a hop only a person can record.
-
Best Practices
Which source to lean on, how to run collection, and how to keep a graph you can trust.
-
Permissions
The add-on gate, what each role can do, and why a deleted connection can come back.
-
How It Works
The connection model, the two levels of granularity, the rules a connection has to obey, and how to find the containers that have one.