AWS Glue Native API
Everything you can do with an AWS Glue Native datastore over the API: create one with a new or an existing connection, test the connection, read and list datastores, update or delete one, and create several at once from the databases of a single Glue Data Catalog.
Complete API Reference
For the full interactive API documentation with all request and response schemas, visit the API docs.
Endpoint: POST /api/datastores
Permission: Manager
Source datastores only
AWS Glue Native is read-only and cannot be used as an enrichment datastore. To link a separate enrichment datastore as the AWS Glue Native source's destination, see Datastore API › Link Enrichment Destination.
UI ↔ API field mapping
The API names several fields differently from the connection form, and the access key pair travels in fields worth knowing about before you build a payload.
| UI form field | API field | Notes |
|---|---|---|
| Connection > Region | connection.parameters.region |
For example us-east-1 |
| Connection > Catalog ID | connection.parameters.catalog_id |
Optional; the 12-digit account ID of another account's catalog |
| Connection > Type | connection.parameters.authentication_type |
SHARED_KEY for Access Key, IAM_ROLE for Assumed Role |
| Connection > Access Key | connection.username |
Access Key only |
| Connection > Secret Key | connection.password |
Access Key only |
| Connection > Role ARN | connection.parameters.role_arn |
Assumed Role only |
| Connection > External ID | connection.parameters.external_id |
Assumed Role only, optional |
| Location > Database | schema |
The Glue database to monitor |
Why the keys travel as username and password
username and password are the fields Qualytics stores encrypted. Sending the key pair there is what keeps it out of the plain-text connection parameters. The Secret Access Key is never returned on a read: responses report only whether one is stored.
Create an AWS Glue Native Datastore
Both flows use the same endpoint and differ only in whether the connection travels inline or by reference.
Pass the region and the credentials inline. Qualytics saves the connection and creates the datastore in a single call, so the connection is then available to reuse.
Connection fields
| Property | Type | Required | Default | Description |
|---|---|---|---|---|
type |
string | Yes | glue_native. |
|
name |
string | Yes | A label for the saved connection. | |
parameters.region |
string | Yes | The AWS region of the Glue Data Catalog. | |
parameters.catalog_id |
string | No | The account ID that owns the catalog, for cross-account access. Omit it to read the credentials' own account. | |
parameters.authentication_type |
string | No | SHARED_KEY |
SHARED_KEY (Access Key) or IAM_ROLE (Assumed Role). IAM_ROLE needs an AWS or local deployment. |
parameters.role_arn |
string | With IAM_ROLE |
The ARN of the role Qualytics assumes. | |
parameters.external_id |
string | No | The External ID the role's trust policy requires, if any. | |
username |
string | With SHARED_KEY |
The Access Key ID. | |
password |
string | With SHARED_KEY |
The Secret Access Key. |
Datastore fields
| Property | Type | Required | Default | Description |
|---|---|---|---|---|
name |
string | Yes | The datastore name in Qualytics. | |
schema |
string | Yes | The Glue database to monitor (shown as Database in the form). | |
teams |
array | Yes | The teams that can work with this datastore. | |
trigger_sync |
boolean | No | false |
Ask Qualytics to run the first Sync once the datastore is created, equivalent to ticking Initiate Sync in the form. |
No Catalog to send
Connectors such as Athena take a Catalog name. Here the region and, when needed, catalog_id already identify the Glue Data Catalog, so there is no catalog field to fill in.
Example request and response (Assumed Role)
Request:
curl -X POST "https://your-instance.qualytics.io/api/datastores" \
-H "Authorization: Bearer YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "Glue Sales",
"connection": {
"name": "acme_glue_us_east_1",
"type": "glue_native",
"parameters": {
"authentication_type": "IAM_ROLE",
"region": "us-east-1",
"role_arn": "arn:aws:iam::123456789012:role/QualyticsGlueRead",
"external_id": "<your-external-id>"
}
},
"schema": "sales",
"teams": ["Data Platform"]
}'
Response (200 OK):
Example request (Access Key, cross-account catalog)
curl -X POST "https://your-instance.qualytics.io/api/datastores" \
-H "Authorization: Bearer YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "Glue Central Sales",
"connection": {
"name": "acme_glue_central",
"type": "glue_native",
"username": "<access-key-id>",
"password": "<secret-access-key>",
"parameters": {
"authentication_type": "SHARED_KEY",
"region": "us-east-1",
"catalog_id": "111122223333"
}
},
"schema": "sales",
"teams": ["Data Platform"]
}'
Reuse an AWS Glue Native connection you already created, through this API or the connection form, by passing its connection_id. The region, Catalog ID, and credentials come from the saved connection, so you supply only the datastore fields.
Reuse it whenever you are adding another database with the same identity. The credentials stay in one place, and changing them later is a single edit rather than one per datastore.
Datastore fields
| Property | Type | Required | Default | Description |
|---|---|---|---|---|
name |
string | Yes | The datastore name in Qualytics. | |
connection_id |
integer | Yes | The ID of the saved AWS Glue Native connection. | |
schema |
string | Yes | The Glue database to monitor (shown as Database in the form). | |
teams |
array | Yes | The teams that can work with this datastore. | |
trigger_sync |
boolean | No | false |
Ask Qualytics to run the first Sync once the datastore is created, equivalent to ticking Initiate Sync in the form. |
Example request and response
Request:
curl -X POST "https://your-instance.qualytics.io/api/datastores" \
-H "Authorization: Bearer YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "Glue Finance",
"connection_id": 7,
"schema": "finance",
"teams": ["Finance Team"]
}'
Response (200 OK):
Run the first Sync
A datastore is only useful once Sync has discovered its tables. Send trigger_sync as true to start it with the datastore, or run a Sync yourself once the request returns.
Test a Connection
Both tests run as the Member role. The second one also needs the Editor team permission on the datastore.
Validates a create payload without persisting anything. Send the same body you would send to create the datastore.
Endpoint: POST /api/datastores/connection
Example request and response
Request:
curl -X POST "https://your-instance.qualytics.io/api/datastores/connection" \
-H "Authorization: Bearer YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "Glue Sales",
"connection": {
"name": "acme_glue_us_east_1",
"type": "glue_native",
"parameters": {
"authentication_type": "IAM_ROLE",
"region": "us-east-1",
"role_arn": "arn:aws:iam::123456789012:role/QualyticsGlueRead"
}
},
"schema": "sales",
"teams": ["Data Platform"]
}'
Response (200 OK):
When the connection cannot be verified, the response is a 400 Bad Request whose message names the failure. A message on a 200 is a warning, not an error.
Re-checks a saved datastore. Because AWS Glue Native is read-only, the check confirms that Qualytics can read the database from the Glue Data Catalog.
Endpoint: POST /api/datastores/{id}/connection
Note
Reading the catalog is not the same as reading the data. A test can pass while the S3 objects are still unreadable, which shows up as a failed profile rather than a failed test. See Troubleshooting.
Get a Datastore
Endpoint: GET /api/datastores/{id}
Permission: Member
Example request and response
Request:
curl -X GET "https://your-instance.qualytics.io/api/datastores/42" \
-H "Authorization: Bearer YOUR_TOKEN"
Response (200 OK):
{
"id": 42,
"name": "Glue Sales",
"type": "glue_native",
"store_type": "native",
"connection_id": 7,
"schema": "sales",
"parameters": {
"authentication_type": "IAM_ROLE",
"region": "us-east-1",
"role_arn": "arn:aws:iam::123456789012:role/QualyticsGlueRead"
},
"has_password": false,
"connected": true,
"enrichment_only": false,
"enrichment_prefix": "_glue_sales",
"teams": [{ "name": "Data Platform" }]
}
The secret key never comes back
has_password reports only whether a Secret Access Key is stored on the connection. Its value is never returned.
List Datastores
Endpoint: GET /api/datastores
Permission: Member
Example request and response
Request:
curl -X GET "https://your-instance.qualytics.io/api/datastores?datastore_type=glue_native&sort_name=asc" \
-H "Authorization: Bearer YOUR_TOKEN"
Response (200 OK):
Query Parameters
| Parameter | Type | Description |
|---|---|---|
id |
list[int] |
Filter by datastore ID. |
name |
string | Filter by exact name. |
search |
string | Partial name match, case-insensitive, or an exact ID. |
datastore_type |
list[string] |
Filter by connector type. glue_native returns only AWS Glue Native datastores. |
tag |
list[string] |
Filter by tag name. |
group |
list[int] |
Filter by datastore group ID. |
enrichment_only |
boolean | false for source datastores (the default). AWS Glue Native datastores are always source datastores. |
sort_name, sort_created, sort_favorite, sort_containers, sort_active_anomalies |
string | asc or desc. |
Update a Datastore
Changes the datastore's own settings. Every request must include name, connection_id, enrichment_only, and enrichment_prefix, even when they are not changing, so send their current values from Get a Datastore. Other parameters you leave out stay as they are, except tags and teams, which replace the whole list when sent.
Endpoint: PUT /api/datastores/{id}
Permission: Member role, with the Editor team permission on the datastore
What lives on the connection
The region, the Catalog ID, and the credentials belong to the connection, not to the datastore. Change them through the connection, and the change reaches every datastore that shares it.
| Property | Type | Required | Description |
|---|---|---|---|
name |
string | Yes | The datastore name. |
connection_id |
integer | Yes | The AWS Glue Native connection the datastore reads through. |
enrichment_only |
boolean | Yes | Always false. AWS Glue Native datastores are source datastores. |
enrichment_prefix |
string | Yes | The prefix for the tables Qualytics writes to the linked enrichment destination. Send the current value to keep it. |
description |
string | No | Free text shown on the datastore. |
schema |
string | No | The Glue database to monitor. |
tags |
list[string] |
No | Replaces the datastore's tags. |
teams |
list[string] |
No | Replaces the datastore's teams. |
favorite |
boolean | No | Marks the datastore as a favorite for the calling user. |
enrichment_source_record_limit, enrichment_remediation_strategy, enrichment_auto_sync, high_count_rollup_threshold |
mixed | No | Settings for the enrichment destination linked to this source. |
Example request and response
Request:
curl -X PUT "https://your-instance.qualytics.io/api/datastores/42" \
-H "Authorization: Bearer YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "Glue Sales (EMEA)",
"connection_id": 7,
"enrichment_only": false,
"enrichment_prefix": "_glue_sales",
"schema": "sales_emea",
"tags": ["glue", "sales"],
"teams": ["Data Platform", "Sales Ops"]
}'
Response (200 OK):
Delete a Datastore
Endpoint: DELETE /api/datastores/{id}
Permission: Admin
Example request
curl -X DELETE "https://your-instance.qualytics.io/api/datastores/42" \
-H "Authorization: Bearer YOUR_TOKEN"
Response: 204 No Content
Warning
Deletion is permanent and cascading: the datastore's containers, checks, anomalies, and enrichment data are removed with it. The connection stays, and so does every other datastore that uses it.
Create Several Datastores at Once
A Glue Data Catalog usually holds more databases than one. The bulk endpoints take a list of databases and create one datastore per entry, sharing the same connection and settings.
Permission: Manager
Endpoint: POST /api/connections/datastores/bulk
Example request and response
Request:
curl -X POST "https://your-instance.qualytics.io/api/connections/datastores/bulk" \
-H "Authorization: Bearer YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"connection": {
"name": "acme_glue_us_east_1",
"type": "glue_native",
"parameters": {
"authentication_type": "IAM_ROLE",
"region": "us-east-1",
"role_arn": "arn:aws:iam::123456789012:role/QualyticsGlueRead",
"external_id": "<your-external-id>"
}
},
"schemas": ["sales", "finance", "inventory"],
"name_template": "glue_{{schema}}",
"teams": ["Data Platform"],
"trigger_sync": true
}'
Response (200 OK):
Endpoint: POST /api/connections/{connection_id}/datastores/bulk
Example request and response
Request:
curl -X POST "https://your-instance.qualytics.io/api/connections/7/datastores/bulk" \
-H "Authorization: Bearer YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"schemas": ["marketing", "logistics"],
"name_template": "glue_{{schema}}",
"teams": ["Data Platform"],
"trigger_sync": true
}'
Response (200 OK):
| Property | Type | Required | Description |
|---|---|---|---|
connection |
object | New connection only | The same connection object as in Create an AWS Glue Native Datastore. |
schemas |
list[string] |
Yes | The Glue databases to create datastores for, one datastore each. |
name_template |
string | No | Naming pattern with {{schema}} standing for the database name. Defaults to the connection name followed by the database. |
teams |
list[string] |
No | Teams assigned to every created datastore. |
tags |
list[string] |
No | Tags applied to every created datastore. |
group_id |
integer | No | An existing datastore group to place them in. |
description |
string | No | Description applied to every created datastore. |
trigger_sync |
boolean | No | Run the first Sync on each datastore as it is created. |
Note
There is no database field to send: the region and Catalog ID on the connection already identify the catalog, so schemas is the whole address. Creation is not atomic. Databases that succeed are created even if others fail, and the response lists both the created IDs and, per failed database, the reason. It also returns connection_id, the connection the datastores were created on. To retry the failed databases, send them to that connection through the Existing Connection endpoint. When no datastore was created, connection_id is null. After an Existing Connection request, retry on the same connection ID. After a New Connection request, resend it; if the response is 409 with Connection '<name>' already exists with id: <id>, a connection with that name already exists, most likely the one your first attempt saved. Check it with GET /api/connections/{id} before reusing it, then retry the databases on that ID through Existing Connection.