Introduction to the Unity Catalog Native Connector
Unity Catalog is the catalog that governs data on Databricks, and it is also available as an open-source server that runs outside Databricks. It organizes tables into catalogs and schemas, records where every table's files live in cloud storage, and decides who may read them. Tools that query Databricks normally go through a cluster or a SQL warehouse, which runs the query and returns the rows.
The Qualytics Unity Catalog Native connector takes a shorter path. It reads the table definitions from Unity Catalog, asks the catalog for a short-lived permission to read each table's files, and then reads those files directly from cloud storage. No Databricks cluster or SQL warehouse sits in the read path, so there is no compute to start, wait for, or pay for while Qualytics profiles tables, runs scheduled scans, and surfaces record- and schema-level anomalies. The connector works with a Databricks workspace or with an open-source Unity Catalog server.
Source datastores only
Unity Catalog Native is read-only, so it cannot be used as an enrichment datastore. Link a separate enrichment datastore as its destination to store the anomalies and metadata Qualytics produces, either during creation or afterwards.
Notes
Choosing between Unity Catalog Native and Databricks. Qualytics keeps both, and they are good at different things. Use Unity Catalog Native for the Delta tables that make up most Databricks estates, and the Databricks connector for what the native path does not read, such as views, and for anything that must run on Databricks compute. Both can be connected against the same workspace at the same time.
| Unity Catalog Native | Databricks | |
|---|---|---|
| How it reads data | Reads the table's files directly from cloud storage | Runs SQL queries on a cluster or SQL warehouse |
| Databricks compute | None | A cluster or SQL warehouse must be running |
| Works with | A Databricks workspace or an open-source Unity Catalog server | A Databricks workspace |
| Tables | Managed and external Delta tables | Any table or view the warehouse can query |
| Views | Not supported | Supported |
| Enrichment datastore | Not supported | Supported |
| Authentication | Access token, or OAuth with a service principal | Personal access token, or OAuth with a service principal |
Supported tables. The connector reads managed and external Delta tables, partitioned or not, in the schema you select. Tables that Unity Catalog makes available only through its Iceberg endpoint are not covered yet, and neither are views, which are query definitions that Databricks itself has to run. Neither becomes a container. The Sync skips the objects the catalog refuses, such as views, and says how many it skipped in its warning. A table the catalog lists but Qualytics cannot read is named in the warning and, if it was already a container, marked Unloadable. The rest of the schema syncs normally. Read those tables through the Databricks connector.
One connection, one catalog. The Catalog field belongs to the connection, next to the URL and the credential, and the Schema field belongs to the datastore. Every schema you pick becomes its own Qualytics datastore on that connection. To monitor a second catalog, add a second connection.
Pickers instead of typing. Once the URL and the credential are filled in, the Catalog dropdown lists the catalogs that credential can see, and the Schema dropdown lists the schemas of the chosen catalog. You can still type a name the list does not show.
Short-lived storage access. Qualytics never holds your storage keys. For each table it reads, Unity Catalog hands out a storage credential that lasts about an hour, which is what makes direct reads safe to allow. The one consequence worth knowing is that a single read of one table that runs longer than that window fails, so very large tables are best scanned incrementally or with a filter. See Best Practices.
What the form does not ask for. There is no cluster, SQL warehouse, or HTTP path to provide, because the connector never runs a query on Databricks. There is no storage credential to provide either, because Unity Catalog supplies one per table.
Deep Dive
-
Authentication
Access tokens and OAuth service principals: which credential goes in which field, the token endpoint, and how secrets are stored.
-
Examples
Worked scenarios: a Databricks workspace, several schemas of one catalog, an open-source Unity Catalog server, and a schema of mixed tables.
-
Best Practices
Recommendations for planning one connection per catalog, choosing the credential, keeping reads inside the credential window, and splitting tables with the Databricks connector.
-
Permissions
The grants and the external access your Unity Catalog must allow, plus the Qualytics user role and team permission needed to manage the datastore.
How-tos
-
Add with New Connection
Add Unity Catalog Native as a source datastore and create its connection at the same time, field by field.
-
Add with Existing Connection
Add another schema from a catalog you already connected, reusing a saved Unity Catalog Native connection.
Reference
-
Troubleshooting
Known problems and how to resolve them, grouped by the step that failed: the catalog, the credential, storage, or a table.
-
API
Create, test, read, update, and delete Unity Catalog Native datastores over the API, including bulk creation from several schemas.
-
FAQ
Short answers to the questions that come up most often about the connector, the tables it reads, and the credentials.