Skip to content

Lineage FAQ

General

What is Lineage?

Lineage lets you visualize how data flows between containers and fields across your datastores. You can explore, expand, and edit connections directly in the UI.

Is Lineage available on all deployments?

Lineage is an add-on feature. It must be enabled for your deployment before the Lineage tab becomes visible. If your deployment is on a trial of the add-on, a Trial badge appears on the tab and in the Add-ons panel. Contact support if you need to extend a trial or request full access.

Does Lineage work with all datastore types?

Yes. Lineage works with both databases (JDBC) and file stores (DFS), and any container or field in a connected datastore can take part in the graph.

Reading lineage out of the source automatically is narrower: that works on Snowflake, BigQuery, Databricks, and Oracle. Every other datastore gets its lineage from a data catalog, from the assets Qualytics writes itself, or by hand.

Where do I access Lineage?

Open any container or datastore and click the Lineage tab. A container's tab is centered on that container. A datastore's tab opens on a map of the datastores it exchanges data with, and from there on all of its containers that have lineage. If the tab is not visible, the Lineage add-on is not enabled for your deployment.

How do I get lineage without drawing it myself?

Three ways, and they work together. Connect a data catalog that publishes lineage and it is imported on each sync. Run a Sync with Collect lineage turned on and Qualytics reads the relationships the source datastore records itself. Computed containers carry lineage to every source in their definition, including each input of a join, at container level. Materialize and remediation outputs each carry lineage back to the single container they copy, at container and field level. None of it needs setup.

How do I turn on lineage collection?

It is an option of the Sync operation, off by default. Open the sync settings, turn on Collect lineage, and run. It applies to Snowflake, BigQuery, Databricks, and Oracle. For other datastore types the option is turned off and cannot be selected, with the caption "Lineage collection is not supported for this connector". Collection starts once the rest of the sync has succeeded, so a sync that fails records no lineage.

What does collection actually read?

The source's query history, to see which tables a statement read in order to write another one, and its view definitions. Where the source keeps its own record of that, Qualytics reads it directly. Where it does not, Qualytics parses the statement text the history holds. Snowflake also publishes its own per-object lineage graph, and Oracle its own record of what each view depends on, which Qualytics reads as well. Oracle keeps no lasting statement history, so there Qualytics reads the write statements the database still holds from recent runs, and also parses the procedures, functions, packages, and triggers stored in the database. The exact relations per datastore are listed in Lineage Sources.

Should I run collection once or keep it on?

Keep it on, on the scheduled Sync. A single run sees only the relationships exercised inside the window it read, so a table written monthly is invisible to a run that swept a week. Each run adds what it sees to what is already there, so the graph fills in over time.


Graph and Navigation

What does the graph show when I first open Lineage?

On a container's Lineage tab, the initial view shows the focal container and its immediate connections, one level upstream and one level downstream. When a side has more neighbors than are drawn, a + N more card stands in for the rest. From there you can expand the graph incrementally. A datastore's Lineage tab opens on the datastore map instead: this datastore in the middle, the datastores that feed it on the left and the ones it feeds on the right.

How do I expand the graph to see more connections?

Containers at the edge of the current view show expand controls on their upstream and downstream sides:

  • Single chevron expands one more level in that direction.
  • Double chevron follows that direction to the end of the chain instead of one step, up to 20 steps out from the node. If the chain runs longer, click again from a node at the new edge.

Can I rearrange nodes in the graph?

Yes. Click and drag any node to reposition it however works best for your analysis. These positions are not persisted and reset when you leave the page or click Reorganize in the toolbar, which re-applies the default left-to-right layout.

How do I re-center the view?

Use the Fit view button in the toolbar to fit all visible nodes into view. On a container's own lineage, Focus current (toolbar → More options) centers on the focal container, and on the datastore map it centers on this datastore. Focus selection centers on a selected node.

What are the anomaly badges on nodes?

Each container node shows a status badge: an alert icon when active anomalies exist, or a check icon when none are found. A container that has never been scanned shows no badge at all. Hovering over the badge shows the exact count. Clicking the badge (only available when anomalies exist) opens the anomaly panel filtered to that container, so you can investigate quality issues without leaving the lineage graph.

What is the focal field?

Clicking a field row inside a container node makes it the focal field, which highlights the full connection chain for that specific field, showing which upstream fields it derives from and which downstream fields it feeds into. If it also connects to containers that are not drawn, a +N chip on its row adds them. Click the same field again, click anywhere on the canvas, or press Esc to clear it.

How are fields listed inside a container node?

Fields that have at least one connection are listed first. When the node also has fields with no connections, those follow below a divider such as 4 fields (no lineage). Scroll down to load more fields, and use the search box to filter by name.

Can I zoom the graph?

Yes. Use the + / − buttons in the toolbar, pinch-to-zoom on trackpad, or click the current zoom percentage to choose a preset.

What is a node with a dashed border?

An asset that a relationship names but that is not registered in Qualytics, shown under the name the source spells it with. It is a reference rather than an asset: nothing opens from it and it has no fields to expand. Right-click it and use Copy Name to take the full name into a search or a ticket.

Register the table in Qualytics and collect again, and it becomes an ordinary node. The connection is not duplicated, because the external side is replaced by the container.

How do I know where a connection came from?

Open the connection's provenance. For lineage collected from a source datastore it names the relation each piece of evidence was read from, what kind of evidence it was, when the relationship was last seen, and the Sync that last saw it, with the logo of whichever side computed it. For every other category it says what the category alone can say: added by hand, imported from the data catalog, parsed from a computed table's SQL, or written by a materialize or scan operation. See Where a Connection Came From.

How do I find which containers have lineage at all?

Use the Has lineage filter on the container listing. Turning it on keeps only the containers with at least one connection. Comparing that list with the unfiltered one shows the containers without lineage, without opening tabs one by one. It answers the container's own state, not how far its graph reaches. A datastore's Lineage tab also lists all of its containers that have lineage, under View all N linked tables.

Why does a number end with +?

Counting stopped before the end, so the real number is at least that. On a card's count row, the tooltip says where it stopped: Counting stopped at 1,000, or Counting stopped 20 steps back (or ahead). The same plus sign appears on datastore map footers and impact analysis tiles.


Datastore Lineage and Impact Analysis

What is a group?

On a datastore's all-containers canvas, containers are drawn in boxes called groups. A group is some of this datastore's assets and everything that feeds them or that they feed, directly or through others. A box shows how many assets the group holds. When only part of a group is drawn, its box reads like 120 of 400 assets, and Show more at the bottom draws more. See Datastore-level Lineage.

Where do I add or delete a connection at datastore level?

Not on the datastore's own views. The map, the lineage between two datastores and the all-containers canvas are read-only. Click Open this table's lineage on a card to open that container's lineage, or go to the container's Lineage tab, and add or delete the connection there. If you opened it from the canvas, Back returns you there, reloaded with your change.

Can I export impact analysis?

Yes. Export at the bottom of the impact analysis panel saves the list as a CSV file, following the side, filter and search you set. On a single asset, Both sides in one file saves upstream and downstream together. Export is not shown in deployments where downloads are turned off. See Impact Analysis.


Edges and Connections

What is the difference between container-level and field-level connections?

A container-level connection links two containers, representing that data flows between them without specifying which fields are involved. A field-level connection links two specific fields, and can carry a note describing the transformation between them, such as UPPER(email), which is set through the API.

Can I mix containers and fields on opposite sides of a connection?

No. Both sides of a connection must be at the same granularity: either container-to-container or field-to-field. A container-to-field or field-to-container connection is not supported.

If I create a field-level connection between fields in different containers, does a container-level connection get created too?

Yes. When you manually create a field-level connection and the source and target fields belong to different containers, Qualytics automatically creates a container-level connection between those containers if one does not already exist.

What is the connection source type?

Each connection has a source type badge that indicates how it was created:

  • Manual: created directly by a user in the UI or via the API.
  • Qualytics managed: generated automatically by Qualytics when a computed container is created or updated, or when a Materialize or remediation output is written into an enrichment destination.
  • Catalog integration: imported from a connected catalog (e.g. Atlan, Alation, Collibra, Purview, DataHub), shown with that integration's logo.
  • Source collection: read out of the source datastore by a Sync run with Collect lineage turned on, shown with that datastore type's logo.

Opening a connection's provenance shows what Qualytics read to arrive at it. See Where a Connection Came From.

How do I delete a connection?

On a container's own lineage, hover over the source type badge on the connection. A delete button appears when you have the required permission. Click it to remove the connection. There is no confirmation step, and deleting a connection does not affect the underlying containers or fields.

Why did a connection come back after I deleted it?

Deleting removes the connection, not the thing that produces it. A data catalog connection returns on the next sync of that data catalog, a collected one the next time the relationship is observed, and a Qualytics managed one when the computed container is saved or the table is written again. Only a manual connection stays deleted on its own, because nothing else produces it.

Does a table Qualytics writes get lineage automatically?

Yes, for the two outputs that are copies of one container. A Materialize output is connected back to the container it copies, and a remediation output back to the container whose anomalous records it holds, at container level and at field level, with nothing to configure. Scan results and exports gather rows from many containers and get no connection, since there is no single source to name.

Why was a connection that no longer exists not removed by the next collection?

Because absence is not evidence. History can show that a relationship happened and can never show that one stopped: a table simply not written during the window leaves no trace. Reconciling on that silence would drop a connection every time a pipeline paused, so collected connections are only ever added or refreshed. Delete the ones that are genuinely gone.


Permissions

Who can view Lineage?

Any user with the Member role or above, once the Lineage add-on is enabled for the deployment. Team permissions do not filter the graph: it shows every connection of the container you opened, including containers in datastores your teams do not cover. They apply when you leave the graph: on a container's Lineage tab, Open this table's lineage opens that container's Lineage tab, which needs the Reporter team permission on its datastore. A datastore's own Lineage tab reads out the whole datastore, so it needs the Viewer team permission on that datastore, and so does the impact analysis of one of its containers. See Team Permissions.

Who can create or delete connections?

Only users with the Manager role can create or delete lineage connections.

Who can enable or disable the Lineage add-on?

Only Admin users can toggle the Lineage add-on on or off from the Add-ons panel. If the add-on is not available for your deployment at all, the toggle is locked. Contact support to request access.


API

Can I manage lineage connections via the API?

Yes. You can create, update, delete connections and query the lineage graph via the API. See the Lineage API page for endpoints and examples.

Does the API return assets that are not registered in Qualytics?

Yes, in their own arrays. The graph response keeps nodes and edges to managed containers only, so a client reading just those two is unaffected, and reports unmanaged assets in external_nodes and external_edges beside them. An external node's id is the string ext:<source_type>:<fqn>, and an external edge points at it through source_node_id or target_node_id instead of a container id.

Can I see what a connection was derived from through the API?

Yes. Every edge response carries an evidence array, populated for lineage collected from a source datastore and empty for the other categories. Each entry names the provider, the kind of evidence, which side computed it, the relation it was read from, and the Sync operations that first and last observed it.