How Lineage Works
Lineage in Qualytics is built from connections, and the graph you see on a container's Lineage tab is assembled from them each time you open it. Understanding what a connection is, and the rules one has to obey, explains most of what the canvas does.
Why Lineage Matters
Lineage exists to answer four questions that are expensive to answer any other way:
- What caused this? Walk upstream from the container that failed until you reach the one that introduced the problem.
- What does this break? Walk downstream to see which tables, files, and consumers depend on what just changed.
- How does this pipeline work? Navigate a multi-stage pipeline visually instead of reconstructing it from transformation code, including stages that live in separate datastores.
- Where did this column come from? Follow one field through the transformations that carry it, rather than reasoning about whole tables.
The graph is only as useful as it is complete, which is why Qualytics records connections wherever it can read them, rather than leaving the graph to be drawn by hand. Where each connection comes from is covered in Lineage Sources.
The Connection Model
A connection is directed and has exactly two endpoints: a source and a target. Direction is the whole meaning of it, since it is what makes one side upstream and the other downstream, and reversing it says the opposite thing. There is no undirected connection, and a relationship that genuinely runs both ways is two connections.
An endpoint is usually a container or a field registered in Qualytics. It can also be an asset that is not, which happens when a collected relationship names a table Qualytics does not manage: the connection is kept with that side recorded as the name the source spells it with. At most one of the two sides can be external, since a connection between two unmanaged assets says nothing about anything Qualytics holds.
Each connection also carries:
- Its origin, recorded when it is created and never guessed afterwards. This is what lets the graph say how Qualytics knows about a relationship, and what decides whether a later process may update or remove it.
- An optional note describing the transformation, such as
UPPER(email). It is free text for the reader, not something Qualytics interprets, and it is set through the API rather than in the Add panel.
The same pair can be connected more than once, but only once per origin. That is deliberate: a manual connection, one imported from Atlan, one imported from Alation, and one collected from the source can all describe the same pair, each reconciled by its own process without touching the others. What is refused is a second connection with the same endpoints and the same origin, which is what stops one process from duplicating its own work.
Granularity
Lineage is recorded at two levels, and both live in the same graph:
- Container-level connects whole containers:
ordersflows intodaily_revenue. It answers which assets are involved, which is what impact analysis needs first. - Field-level connects one field to another:
customers.emailflows intomarketing_contacts.email_address. It answers which column fed which, which is what you need once the question narrows to a single value being wrong.
Both endpoints of a connection have to be at the same level. A container on one side and a field on the other is rejected, because the graph traversal has no meaningful way to walk such a connection: it would have to treat a whole asset and one of its columns as equivalent steps.
Field Connections Carry Their Container
When you add a field connection by hand between fields that live in different containers, Qualytics also creates the container-level connection between those two containers, if one does not already exist. Without it the field connection would be invisible at the level most people read the graph at, and the containers would look unrelated while their columns are demonstrably not.
Fields inside the same container are the exception. A connection between two of them describes a transformation that happens within the asset, and there is no second container to connect.
Where a Connection Comes From
Every connection is tagged with what produced it, and that tag governs its life: whether it is refreshed, reconciled, or left alone by a later process. Four categories exist, each with different rules, and they coexist on the same container without overwriting each other.
The categories, their rules, and what each one can and cannot prove are covered in Lineage Sources.
Finding Containers That Have Lineage
Lineage is recorded per container, so a datastore ends up with a mix of containers that carry connections and containers that do not. The container listing has a Has lineage filter for that: keep only the containers with at least one connection, or only those with none.
The filter answers the container's own state, not how far its graph reaches. A container with a single connection passes it, whatever that connection leads to, so it is a good way to find coverage gaps and a poor way to judge depth.
See Also
-
Lineage Sources
The four categories a connection can come from, and how each is created, updated, and removed.
-
Reading the Graph
Direction, nodes, assets outside Qualytics, where a connection came from, and expanding the graph.
-
Field-level Lineage
Expanding field lists, field metadata, and the focal field workflow.
-
Examples
Enrolling a warehouse, a medallion pipeline across datastores, a remediation table, and a hop only a person can record.
-
Best Practices
Which source to lean on, how to run collection, and how to keep a graph you can trust.
-
Permissions
The add-on gate, what each role can do, and why a deleted connection can come back.