Hive
Adding and configuring a Hive connection within Qualytics empowers the platform to build a symbolic link with your schema to perform operations like data discovery, visualization, reporting, syncing, profiling, scanning, anomaly surveillance, and more.
This documentation provides a step-by-step guide on how to add Hive as a source datastore in Qualytics. It covers the entire process, from initial connection setup to testing and finalizing the configuration.
By following these instructions, enterprises can ensure their Hive environment is properly connected with Qualytics, unlocking the platform's potential to help you proactively manage your full data quality lifecycle.
Two ways to connect to Hive
This page covers the Hive connector, which reads your data by running queries through HiveServer2. Qualytics also offers Hive Native, which reads file-based tables such as Parquet, ORC, delimited text, JSON, and Avro straight from their files and is considerably faster on large tables. Use this connector for anything Hive Native declines to read, such as views and transactional tables, and see Choosing between the two Hive connectors for a side-by-side comparison. Both can be connected against the same cluster.

Let’s get started 🚀
Hive Setup Guide
Qualytics connects to Hive through the Hive JDBC driver (HiveServer2). It accesses the Hive metastore for schema and table discovery, and reads table data via HiveQL SELECT queries for profiling and scanning operations. Qualytics also identifies partition columns from Hive metastore metadata for optimized data reading.
Minimum Hive Permissions (Source Datastore)
| Permission | Purpose |
|---|---|
SELECT ON DATABASE <database_name> |
Access the database and its metadata |
SELECT ON TABLE <database_name>.* |
Read data from all tables for profiling and scanning |
Note
Qualytics cannot write to Hive, so a Hive datastore cannot be used as an enrichment datastore. Use an enrichment datastore on another connector instead.
Example: Source Datastore User (Read-Only)
Replace <database_name> with your actual value.
-- Grant read access to the target database and all its tables
GRANT SELECT ON DATABASE <database_name> TO USER qualytics_read;
GRANT SELECT ON ALL TABLES IN DATABASE <database_name> TO USER qualytics_read;
Note
If using Kerberos authentication, ensure the Kerberos principal Qualytics logs in as has been granted the same SELECT privileges on the target database. Enter that principal as the Client Principal in the connection form and upload its keytab and the cluster's krb5.conf there, instead of a username and password.
Info
If your Hive environment uses ZooKeeper for HiveServer2 high availability (HA), select the ZooKeeper HA checkbox in the connection form and list the ZooKeeper quorum hosts in Host. Qualytics then discovers the available HiveServer2 instances through the quorum. It works with both the Basic and Kerberos authentication types.
Troubleshooting Common Errors
| Error | Likely Cause | Fix |
|---|---|---|
User is not allowed to impersonate |
The HiveServer2 proxy user configuration does not allow the Qualytics user | Add the Qualytics user to the hadoop.proxyuser.<hive_user>.users property in core-site.xml |
Permission denied: user does not have SELECT privilege |
The user lacks SELECT on the target database or table |
Run GRANT SELECT ON DATABASE <database_name> TO USER <user> |
GSS initiate failed (Kerberos) |
The uploaded keytab no longer matches the key the KDC holds, the principal is incorrect, or the KDC is unreachable | Confirm the principal, upload the current keytab, and check that the KDC named in krb5.conf is reachable |
Kerberos Client Principal must name an identity in the keytab |
The Client Principal holds a service principal such as hive/_HOST@EXAMPLE.COM |
Enter the identity in the keytab as Client Principal (for example qualytics@EXAMPLE.COM) and move the service principal to HiveServer2 Service Principal |
Could not open client transport |
HiveServer2 is not reachable or the port (default 10000) is incorrect | Verify the host, port, and that HiveServer2 is running |
Database does not exist |
The database name (schema) in the connection form is incorrect | Verify the database name with SHOW DATABASES in HiveQL |
Detailed Troubleshooting Notes
Authentication Errors
The error GSS initiate failed (Kerberos) or Could not open client transport indicates an authentication or transport problem.
Common causes:
- Stale keytab: the credential was changed on the cluster and the keytab on the connection is the previous one. Upload the current keytab; Qualytics renews tickets from it on its own, so there is nothing to run by hand.
- Wrong principal: the Client Principal does not match the identity in the keytab, or the HiveServer2 Service Principal does not match the one configured in the Hive server.
- KDC unreachable: the Key Distribution Center named in the uploaded
krb5.confis not reachable from your Qualytics deployment. - Realm already claimed: another connection uses the same realm name with a different KDC. One realm name cannot mean two clusters, so give each a distinct realm name or connect them from separate deployments.
- HiveServer2 not running: the HiveServer2 process is not started or has crashed.
Note
If using basic authentication (username/password), ensure HiveServer2 is configured to accept password-based authentication (hive.server2.authentication=CUSTOM or LDAP).
Permission Errors
The error Permission denied: user does not have SELECT privilege means the user authenticated successfully but lacks the necessary Hive grants.
Common causes:
- Missing
SELECTon database: the user does not haveSELECTon the target database. - Missing
SELECTon table: the user has database-level access but not table-level access. - Ranger/Sentry policy: if Apache Ranger or Sentry is enabled, permissions are managed through policies rather than Hive
GRANTstatements.
Connection Errors
The error Could not open client transport or User is not allowed to impersonate indicates a transport or proxy user issue.
Common causes:
- HiveServer2 not reachable: the host or port (default 10000) is incorrect.
- Proxy user not allowed: the
hadoop.proxyuser.<hive_user>.usersproperty incore-site.xmldoes not include the Qualytics user. - SSL required: HiveServer2 requires SSL but the connection is not configured for it.
Tip
Start by confirming credentials and Kerberos configuration are valid (authentication errors), then verify Hive grants or Ranger policies (permission errors), and finally check HiveServer2 connectivity (connection errors).
Add a Source Datastore
A source datastore is a storage location Qualytics connects to so it can profile, scan, and monitor data. Adding Hive as a source lets Qualytics query it through the Hive JDBC driver and run quality operations on the tables it discovers.
Before you start, review the Minimum Hive Permissions the connecting user needs.
The destination lives on another connector
Qualytics cannot write to Hive, so a Hive datastore cannot be used as an enrichment datastore. Create one on a supported connector, then link it as the Hive datastore's destination. See Supported Connectors for the list.
Field reference
The Add Datastore page shows the sections below when Hive is selected. When reusing an existing connection, the Connection Properties and Secrets Management sections come already filled in and read-only: Qualytics has already validated those credentials, so you fill in only the Location and the General fields. To change a saved connection's credentials, edit the connection through the Manage Connections page; edits there apply to every datastore that reuses the connection.
Connection Properties
These fields define the HiveServer2 endpoint Qualytics connects to. They belong to the connection: when reusing an existing connection, they come already filled in and read-only.
| Field | Required | Type | Description |
|---|---|---|---|
| Connection Name | Text | A label for the saved connection (e.g., acme_hive_reporting), so other datastores can reuse it later. |
|
| ZooKeeper HA | Checkbox | Select it when the cluster is fronted by ZooKeeper, so Qualytics finds HiveServer2 through the ZooKeeper quorum instead of connecting to one fixed endpoint. Off by default. It works with either authentication Type. | |
| Host | Text | The hostname or address of HiveServer2. With ZooKeeper HA selected, list the ZooKeeper quorum hosts separated by commas instead, each with its client port (for example zk1:2181,zk2:2181,zk3:2181). A host written without a port gets the Port below. |
|
| Port | Number | The port HiveServer2 listens on. Defaults to 10000. With ZooKeeper HA selected, it is added to every quorum host written without its own port, so set it to the ZooKeeper client port (commonly 2181) or write the port on each host. |
|
| Max Parallelization | Number | Caps how many queries Qualytics runs at the same time against a datastore reached through this connection, from 1 to 100. Leave it empty for no explicit limit. The cap applies per datastore, so datastores sharing this connection each get their own allowance. Sits in the collapsed Advanced section of the form. See Max Parallelization. |
Authentication
Choose how Qualytics authenticates to HiveServer2. Setting Type changes the credential fields shown below it, so pick the tab that matches your choice. These fields also belong to the connection: already filled in and read-only when reusing one.
| Field | Required | Type | Description |
|---|---|---|---|
| Type | Option | Set to Basic, which is the default (BASIC in the API). |
|
| User | Text | The Hive user Qualytics connects as. Comes filled in as hive. |
|
| Password | Text | The password for that user. Leave it empty on a cluster that does not authenticate. |
| Field | Required | Type | Description |
|---|---|---|---|
| Type | Option | Set to Kerberos (KERBEROS in the API). |
|
| Client Principal | Text | The identity Qualytics logs in as, exactly as it appears in the keytab, for example qualytics@EXAMPLE.COM. Do not use a service principal of the form service/_HOST or service/_HOST@REALM here: login does not substitute a hostname, so the connection test fails. |
|
| HiveServer2 Service Principal | Text | The HiveServer2 service identity, for example hive/_HOST@EXAMPLE.COM. _HOST is replaced with the name of the HiveServer2 host Qualytics connects to. When the connection has its own keytab and this field is empty, Qualytics uses hive/_HOST@ followed by the realm of the Client Principal. Fill it in when the service name or realm differs, or when the connection has no keytab. |
|
| Keytab | File | The keytab file holding the key for the Client Principal. Must be named with a .keytab extension. Stored encrypted and never displayed back once saved. Upload a new one to change it; leaving the field empty on an edit keeps the one already stored. |
|
| krb5.conf | File | The Kerberos configuration for the cluster's realm, so Qualytics can reach its Key Distribution Center. It holds realm names and KDC hostnames rather than credentials, so it is not treated as a secret. |
No password is asked for. Qualytics authenticates with the uploaded keytab and renews its tickets on its own, so there is nothing to run by hand. Because the credentials live on the connection, changing them, or adding a second cluster, is a change to the connection rather than to your deployment.
For example, a client principal of qualytics@CLIENT.EXAMPLE.COM can connect to HiveServer2 as hive/_HOST@SERVER.EXAMPLE.COM when the Kerberos realms trust each other. Enter each value in its own field.
Connections saved with the service principal in one field
Earlier, a Kerberos connection had a single Principal field that held the HiveServer2 service principal. A connection that has its own keytab and still has a value such as hive/_HOST@EXAMPLE.COM in Client Principal now fails the connection test. Edit the connection: move that value to HiveServer2 Service Principal and enter the identity in the keytab as Client Principal. A connection without a keytab of its own keeps working as saved.
ZooKeeper connections saved earlier
Zookeeper and Zookeeper Kerberos used to be options of Type. Connections saved with them keep working: Qualytics reads them as Basic or Kerberos with ZooKeeper HA selected. Use that combination for new connections, including over the API.
Secrets Management
This group is optional: use it only if you want Qualytics to pull credentials from a secrets manager instead of typing them into the form. Turn on HashiCorp Vault to show the fields below. Despite the label, any secrets manager that exposes a compatible REST API works, not only HashiCorp Vault; see Secrets Management. It also belongs to the connection: read-only when reusing an existing connection.
| Field | Required | Type | Description |
|---|---|---|---|
| Login URL | Text | The Vault endpoint Qualytics uses to authenticate (e.g., https://vault.example.com/v1/auth/approle/login). |
|
| Credentials Payload | Text | A JSON body containing the credentials Vault expects (e.g., {"role_id":"...","secret_id":"..."}). |
|
| Token JSONPath | Text | The JSONPath that extracts the client token from Vault's response. Defaults to $.auth.client_token. |
|
| Secret URL | Text | The Vault path where the secret is stored (e.g., https://vault.example.com/v1/secret/data/hive). |
|
| Token Header Name | Text | The HTTP header name used to send the token. Defaults to X-Vault-Token. |
|
| Data JSONPath | Text | The JSONPath that extracts the secret payload from Vault's response. Defaults to $.data. |
Location
Pick the schema or schemas Qualytics should read from. You fill these in on both flows.
| Field | Required | Type | Description |
|---|---|---|---|
| Schema | Option | One or more Hive schemas to read from. Each schema you pick becomes its own Qualytics datastore. Defaults to default, and you can click the refresh icon to load the ones visible to the user. |
No catalog step
Hive has no separate catalog selection, so you pick the schema directly. Selecting more than one creates one source datastore per schema, named from the Name Template. See Multi-Schema Source Datastore Creation for details.
System schema
Hive's system schema is left out of discovery, so it does not appear in the list.
General
Common fields for every source datastore, shown below the Location section. You fill these in on both flows.
| Field | Required | Type | Description |
|---|---|---|---|
| Name Template | Text | Defines the naming pattern for each source datastore being created. Use {{schema}} as a placeholder that gets replaced with the actual schema name (e.g., hive_{{schema}} becomes hive_sales). Left empty, the datastore is named from the connection name and the schema. |
|
| Group | Option | Organizes your datastores under a shared group in the navigation tree. Select an existing group or create a new one with the Add New Group toggle. | |
| Teams | Option | Select one or more teams to associate with this source datastore. | |
| Initiate Sync | Checkbox | Automatically sync the datastore to detect containers and fields after creation. |
Below the form, an information banner lists the Public addresses and Private addresses your datastore connections originate from. Allow the ones that match your network setup through your security groups or firewall rules; see How Connections Work for where the addresses come from.
Steps
There are two ways to set up the connection: reuse a connection you already saved (Existing Connection) or create a new one from scratch (New Connection). The tabs below walk through each option; pick the one you want to follow. Each field is described in the Field reference above.
Step 1: Navigate to the Datastores page.
Step 2: Click the Add button at the top-right corner and choose Source .
Step 3: The Add Datastore page opens.
Step 4: Select New Connection next to the Search field.
Step 5: Select Hive from the connector grid. Use the search field to filter connectors by name.
Step 6: Fill in the Connection Properties: the Connection Name, Host, and Port, then fill in the Authentication fields for the Type you choose. Select ZooKeeper HA when the cluster is fronted by a ZooKeeper quorum.
Step 7: Optionally, expand Secrets Management to retrieve credentials from a secrets manager.
Step 8: Fill in the Location fields (Schema) and the General fields.
Step 9: Click Test connection. A success message confirms that the connection has been verified.
Info
The Finish and Next buttons stay disabled until the connection test passes on the current values. If the test fails, see Troubleshooting Common Errors.
Step 10: Click Finish to create the datastore.
Tip
To link a destination so Qualytics can write anomalies and metadata from the first operation, click Next instead of Finish. It has to live on a connector other than Hive; see Link Enrichment on Datastore Creation.
Step 11: A success dialog confirms that your datastore has been added. Click Go to your datastore to open its page.
Step 1: Navigate to the Datastores page.
Step 2: Click the Add button at the top-right corner and choose Source .
Step 3: The Add Datastore page opens.
Step 4: Select Existing Connection next to the Search field.
Step 5: Select the saved Hive connection from the grid. Use the search field to filter connections by name. The Connection Properties and Secrets Management sections come already filled in and read-only.
Start a new connection from this one
To use the selected connection as a starting point for a brand-new connection instead, click the Duplicate as a new connection button on the selected connection. The form switches to New Connection mode with the connection's settings already filled in for you to adjust.
Step 6: Fill in the Location fields (Schema) and the General fields. These are the only fields left to fill in.
Step 7: Click Test connection. A success message confirms that the connection has been verified.
Info
The Finish and Next buttons stay disabled until the connection test passes on the current values. If the test fails, see Troubleshooting Common Errors.
Step 8: Click Finish to create the datastore.
Tip
To link a destination so Qualytics can write anomalies and metadata from the first operation, click Next instead of Finish. It has to live on a connector other than Hive; see Link Enrichment on Datastore Creation.
Step 9: A success dialog confirms that your datastore has been added. Click Go to your datastore to open its page.
API Payload Examples
This section provides detailed examples of API payloads to guide you through the process of creating and managing datastores using Qualytics API.
Each example includes endpoint details, sample payloads, and instructions on how to replace placeholder values with actual data relevant to your setup.
Creating a Datastore
This section provides a sample payload for creating a datastore. Replace the placeholder values with actual data relevant to your setup.
Endpoint (Post)
/api/datastores (post)
{
"name": "your_datastore_name",
"teams": ["Public"],
"schema": "hive_schema",
"enrichment_only": false,
"trigger_sync": true,
"connection": {
"name": "your_connection_name",
"type": "hive",
"host": "hive_host",
"port": 10000,
"username": "hive_username",
"password": "hive_password",
"parameters": {
"zookeeper": false,
"authentication_type": "BASIC"
}
}
}
{
"name": "your_datastore_name",
"teams": ["Public"],
"schema": "hive_schema",
"enrichment_only": false,
"trigger_sync": true,
"connection": {
"name": "your_connection_name",
"type": "hive",
"host": "hive_host",
"port": 10000,
"username": "qualytics@EXAMPLE.COM",
"password": "<base64-encoded keytab>",
"parameters": {
"zookeeper": false,
"authentication_type": "KERBEROS",
"hive.metastore.kerberos.principal": "hive/_HOST@EXAMPLE.COM",
"krb5_base64": "<krb5.conf contents>"
}
}
}
The Client Principal travels as username, and the keytab as password, base64 encoded, because it is a binary file and belongs in the encrypted field. The HiveServer2 Service Principal goes in parameters as hive.metastore.kerberos.principal. When the connection sends its own keytab, you can leave it out to use hive/_HOST@ followed by the realm of username; a connection without its own keytab has no such default, so send the service principal. krb5_base64 accepts the file's text as-is. For ZooKeeper HA, set zookeeper to true and list the quorum hosts in host, with their client port on each host or in port.
# Step 1: Create a Connection
qualytics connections create \
--type hive \
--name "your_connection_name" \
--host ${HIVE_HOST} \
--port 10000 \
--username ${HIVE_USER} \
--password ${HIVE_PASSWORD}
# Step 2: Create a Source Datastore
qualytics datastores create \
--name "your_datastore_name" \
--connection-name "your_connection_name" \
--schema default
Link an Enrichment Destination
Endpoint (Patch)
/api/datastores/{datastore-id}/enrichment/{enrichment-id} (patch)