Skip to content

Unity Catalog Native Best Practices

Recommendations for getting reliable results and predictable run times from a Unity Catalog Native datastore.

Plan the connection

Grant external access before the first profile, not after. A datastore whose identity can list tables but cannot obtain their storage credential will connect, sync, and list every table without complaint, then fail on the first operation that reads records. Checking EXTERNAL USE SCHEMA and the metastore's external data access setting up front turns a confusing failure into a five-minute setup step. See Permissions.

Use one connection per catalog and reuse it. The catalog belongs to the connection, so every schema you monitor in the same catalog should share one saved connection. The credential then lives in one place, and replacing it is a single edit rather than one per datastore. Add further schemas with Add with an Existing Connection.

Confirm that storage is reachable from your deployment. The credential Unity Catalog hands out grants permission, not a network route. A bucket or storage account that only accepts traffic from Databricks cannot be read even with a valid credential. Allow the deployment's addresses on the storage side, as listed in Network access.


Choose the credential

Prefer a service principal over a personal token. A personal access token belongs to a person and stops working when that person leaves or their token expires. A service principal created for Qualytics belongs to the integration, carries only the grants it needs, and shows up under its own name in audit logs.

Prefer OAuth where the server offers it. With OAuth (Service Principal), Qualytics requests short-lived tokens on its own, so there is no expiry date to track and nothing to replace on a schedule. With Access Token, put the expiry date in your calendar: every datastore on the connection stops reading the day the token expires.

Keep the secret out of the form when you can. With Secrets Management turned on, the token or client secret can be a ${key} reference to a secret in HashiCorp Vault or another secrets manager, so a changed value takes effect without editing the connection. See HashiCorp Vault.


Choose the right connector per table

Send Delta tables through Unity Catalog Native and only the rest through Databricks. The native path is the one with no compute to start or pay for, and it reads managed and external Delta tables. What it does not read is views, tables reachable only through the Iceberg endpoint, and foreign tables. Pointing both connectors at the same workspace, and sending only those through the Databricks connector, gets you the savings for most of the schema and full coverage for the rest.

Do not treat an unread table as a failure to work around. A view is a query that only Databricks can run, and an Iceberg or foreign table is served by another engine. Rewriting them as Delta tables just to monitor them is rarely worth it. Reading them through the Databricks connector is.


Keep reads inside the credential window

Scan large tables incrementally or with a filter. The storage credential Unity Catalog hands out for a table lasts about an hour, and a single read of one table that outlives it fails. Qualytics requests a fresh credential for every table it reads, so the limit applies per table, not per operation. An incremental scan that reads only new records, or a filter on a partition column, keeps each read well inside the window. See Troubleshooting.

Partition the tables you monitor. A scan whose conditions name the partition columns reads only the matching partitions instead of the whole table. That is the difference between reading one day and reading every day the table holds, and it is also what keeps a read inside the credential window.

Narrow the datastore to the schemas you actually monitor. A schema with hundreds of tables spends real time on table definitions at every Sync, and every table it lists is one a scheduled scan may read. Selecting only the schemas you intend to profile keeps sync and scan times proportional to what you care about.


Keep the data readable

Keep Delta statistics current. When a table's Delta log carries a row count, Qualytics uses it instead of counting the rows itself, which removes a full pass over a very large table at profile time. Tables written by Databricks keep these up to date on their own.

Watch for tables that stop being Delta. A table that is converted to Iceberg, or replaced by a foreign table, drops out of the native path at the next Sync and is marked Unloadable. The Sync names it in its warning, so a glance at the Sync result after a platform migration is enough to catch it.


See Also