Skip to content

AWS Glue Native API

Everything you can do with an AWS Glue Native datastore over the API: create one with a new or an existing connection, test the connection, read and list datastores, update or delete one, and create several at once from the databases of a single Glue Data Catalog.

Complete API Reference

For the full interactive API documentation with all request and response schemas, visit the API docs.

Endpoint: POST /api/datastores

Permission: Manager

Source datastores only

AWS Glue Native is read-only and cannot be used as an enrichment datastore. To link a separate enrichment datastore as the AWS Glue Native source's destination, see Datastore API › Link Enrichment Destination.


UI ↔ API field mapping

The API names several fields differently from the connection form, and the access key pair travels in fields worth knowing about before you build a payload.

UI form field API field Notes
Connection > Region connection.parameters.region For example us-east-1
Connection > Catalog ID connection.parameters.catalog_id Optional; the 12-digit account ID of another account's catalog
Connection > Type connection.parameters.authentication_type SHARED_KEY for Access Key, IAM_ROLE for Assumed Role
Connection > Access Key connection.username Access Key only
Connection > Secret Key connection.password Access Key only
Connection > Role ARN connection.parameters.role_arn Assumed Role only
Connection > External ID connection.parameters.external_id Assumed Role only, optional
Location > Database schema The Glue database to monitor

Why the keys travel as username and password

username and password are the fields Qualytics stores encrypted. Sending the key pair there is what keeps it out of the plain-text connection parameters. The Secret Access Key is never returned on a read: responses report only whether one is stored.


Create an AWS Glue Native Datastore

Both flows use the same endpoint and differ only in whether the connection travels inline or by reference.

Pass the region and the credentials inline. Qualytics saves the connection and creates the datastore in a single call, so the connection is then available to reuse.

Connection fields

Property Type Required Default Description
type string Yes glue_native.
name string Yes A label for the saved connection.
parameters.region string Yes The AWS region of the Glue Data Catalog.
parameters.catalog_id string No The account ID that owns the catalog, for cross-account access. Omit it to read the credentials' own account.
parameters.authentication_type string No SHARED_KEY SHARED_KEY (Access Key) or IAM_ROLE (Assumed Role). IAM_ROLE needs an AWS or local deployment.
parameters.role_arn string With IAM_ROLE The ARN of the role Qualytics assumes.
parameters.external_id string No The External ID the role's trust policy requires, if any.
username string With SHARED_KEY The Access Key ID.
password string With SHARED_KEY The Secret Access Key.

Datastore fields

Property Type Required Default Description
name string Yes The datastore name in Qualytics.
schema string Yes The Glue database to monitor (shown as Database in the form).
teams array Yes The teams that can work with this datastore.
trigger_sync boolean No false Ask Qualytics to run the first Sync once the datastore is created, equivalent to ticking Initiate Sync in the form.

No Catalog to send

Connectors such as Athena take a Catalog name. Here the region and, when needed, catalog_id already identify the Glue Data Catalog, so there is no catalog field to fill in.

Example request and response (Assumed Role)

Request:

curl -X POST "https://your-instance.qualytics.io/api/datastores" \
  -H "Authorization: Bearer YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "Glue Sales",
    "connection": {
      "name": "acme_glue_us_east_1",
      "type": "glue_native",
      "parameters": {
        "authentication_type": "IAM_ROLE",
        "region": "us-east-1",
        "role_arn": "arn:aws:iam::123456789012:role/QualyticsGlueRead",
        "external_id": "<your-external-id>"
      }
    },
    "schema": "sales",
    "teams": ["Data Platform"]
  }'

Response (200 OK):

{
  "id": 42,
  "name": "Glue Sales",
  "type": "glue_native",
  "store_type": "native",
  "schema": "sales",
  "connected": true
}
Example request (Access Key, cross-account catalog)
curl -X POST "https://your-instance.qualytics.io/api/datastores" \
  -H "Authorization: Bearer YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "Glue Central Sales",
    "connection": {
      "name": "acme_glue_central",
      "type": "glue_native",
      "username": "<access-key-id>",
      "password": "<secret-access-key>",
      "parameters": {
        "authentication_type": "SHARED_KEY",
        "region": "us-east-1",
        "catalog_id": "111122223333"
      }
    },
    "schema": "sales",
    "teams": ["Data Platform"]
  }'

Reuse an AWS Glue Native connection you already created, through this API or the connection form, by passing its connection_id. The region, Catalog ID, and credentials come from the saved connection, so you supply only the datastore fields.

Reuse it whenever you are adding another database with the same identity. The credentials stay in one place, and changing them later is a single edit rather than one per datastore.

Datastore fields

Property Type Required Default Description
name string Yes The datastore name in Qualytics.
connection_id integer Yes The ID of the saved AWS Glue Native connection.
schema string Yes The Glue database to monitor (shown as Database in the form).
teams array Yes The teams that can work with this datastore.
trigger_sync boolean No false Ask Qualytics to run the first Sync once the datastore is created, equivalent to ticking Initiate Sync in the form.
Example request and response

Request:

curl -X POST "https://your-instance.qualytics.io/api/datastores" \
  -H "Authorization: Bearer YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "Glue Finance",
    "connection_id": 7,
    "schema": "finance",
    "teams": ["Finance Team"]
  }'

Response (200 OK):

{
  "id": 43,
  "name": "Glue Finance",
  "type": "glue_native",
  "store_type": "native",
  "schema": "finance",
  "connected": true
}

Run the first Sync

A datastore is only useful once Sync has discovered its tables. Send trigger_sync as true to start it with the datastore, or run a Sync yourself once the request returns.


Test a Connection

Both tests run as the Member role. The second one also needs the Editor team permission on the datastore.

Validates a create payload without persisting anything. Send the same body you would send to create the datastore.

Endpoint: POST /api/datastores/connection

Example request and response

Request:

curl -X POST "https://your-instance.qualytics.io/api/datastores/connection" \
  -H "Authorization: Bearer YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "Glue Sales",
    "connection": {
      "name": "acme_glue_us_east_1",
      "type": "glue_native",
      "parameters": {
        "authentication_type": "IAM_ROLE",
        "region": "us-east-1",
        "role_arn": "arn:aws:iam::123456789012:role/QualyticsGlueRead"
      }
    },
    "schema": "sales",
    "teams": ["Data Platform"]
  }'

Response (200 OK):

{
  "connected": true,
  "message": null
}

When the connection cannot be verified, the response is a 400 Bad Request whose message names the failure. A message on a 200 is a warning, not an error.

Re-checks a saved datastore. Because AWS Glue Native is read-only, the check confirms that Qualytics can read the database from the Glue Data Catalog.

Endpoint: POST /api/datastores/{id}/connection

Example request and response

Request:

curl -X POST "https://your-instance.qualytics.io/api/datastores/42/connection" \
  -H "Authorization: Bearer YOUR_TOKEN"

Response (200 OK):

{
  "connected": true,
  "message": null
}

A datastore that cannot be reached answers "connected": false with the reason in message.

Note

Reading the catalog is not the same as reading the data. A test can pass while the S3 objects are still unreadable, which shows up as a failed profile rather than a failed test. See Troubleshooting.


Get a Datastore

Endpoint: GET /api/datastores/{id}

Permission: Member

Example request and response

Request:

curl -X GET "https://your-instance.qualytics.io/api/datastores/42" \
  -H "Authorization: Bearer YOUR_TOKEN"

Response (200 OK):

{
  "id": 42,
  "name": "Glue Sales",
  "type": "glue_native",
  "store_type": "native",
  "connection_id": 7,
  "schema": "sales",
  "parameters": {
    "authentication_type": "IAM_ROLE",
    "region": "us-east-1",
    "role_arn": "arn:aws:iam::123456789012:role/QualyticsGlueRead"
  },
  "has_password": false,
  "connected": true,
  "enrichment_only": false,
  "enrichment_prefix": "_glue_sales",
  "teams": [{ "name": "Data Platform" }]
}

The secret key never comes back

has_password reports only whether a Secret Access Key is stored on the connection. Its value is never returned.


List Datastores

Endpoint: GET /api/datastores

Permission: Member

Example request and response

Request:

curl -X GET "https://your-instance.qualytics.io/api/datastores?datastore_type=glue_native&sort_name=asc" \
  -H "Authorization: Bearer YOUR_TOKEN"

Response (200 OK):

{
  "items": [
    { "id": 42, "name": "Glue Sales", "type": "glue_native", "store_type": "native", "connected": true },
    { "id": 43, "name": "Glue Finance", "type": "glue_native", "store_type": "native", "connected": true }
  ],
  "total": 2,
  "page": 1,
  "size": 50,
  "pages": 1
}

Query Parameters

Parameter Type Description
id list[int] Filter by datastore ID.
name string Filter by exact name.
search string Partial name match, case-insensitive, or an exact ID.
datastore_type list[string] Filter by connector type. glue_native returns only AWS Glue Native datastores.
tag list[string] Filter by tag name.
group list[int] Filter by datastore group ID.
enrichment_only boolean false for source datastores (the default). AWS Glue Native datastores are always source datastores.
sort_name, sort_created, sort_favorite, sort_containers, sort_active_anomalies string asc or desc.

Update a Datastore

Changes the datastore's own settings. Every request must include name, connection_id, enrichment_only, and enrichment_prefix, even when they are not changing, so send their current values from Get a Datastore. Other parameters you leave out stay as they are, except tags and teams, which replace the whole list when sent.

Endpoint: PUT /api/datastores/{id}

Permission: Member role, with the Editor team permission on the datastore

What lives on the connection

The region, the Catalog ID, and the credentials belong to the connection, not to the datastore. Change them through the connection, and the change reaches every datastore that shares it.

Property Type Required Description
name string Yes The datastore name.
connection_id integer Yes The AWS Glue Native connection the datastore reads through.
enrichment_only boolean Yes Always false. AWS Glue Native datastores are source datastores.
enrichment_prefix string Yes The prefix for the tables Qualytics writes to the linked enrichment destination. Send the current value to keep it.
description string No Free text shown on the datastore.
schema string No The Glue database to monitor.
tags list[string] No Replaces the datastore's tags.
teams list[string] No Replaces the datastore's teams.
favorite boolean No Marks the datastore as a favorite for the calling user.
enrichment_source_record_limit, enrichment_remediation_strategy, enrichment_auto_sync, high_count_rollup_threshold mixed No Settings for the enrichment destination linked to this source.
Example request and response

Request:

curl -X PUT "https://your-instance.qualytics.io/api/datastores/42" \
  -H "Authorization: Bearer YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "Glue Sales (EMEA)",
    "connection_id": 7,
    "enrichment_only": false,
    "enrichment_prefix": "_glue_sales",
    "schema": "sales_emea",
    "tags": ["glue", "sales"],
    "teams": ["Data Platform", "Sales Ops"]
  }'

Response (200 OK):

{
  "id": 42,
  "name": "Glue Sales (EMEA)",
  "type": "glue_native",
  "store_type": "native",
  "connection_id": 7,
  "schema": "sales_emea",
  "global_tags": [{ "name": "glue" }, { "name": "sales" }],
  "teams": [{ "name": "Data Platform" }, { "name": "Sales Ops" }]
}

Delete a Datastore

Endpoint: DELETE /api/datastores/{id}

Permission: Admin

Example request
curl -X DELETE "https://your-instance.qualytics.io/api/datastores/42" \
  -H "Authorization: Bearer YOUR_TOKEN"

Response: 204 No Content

Warning

Deletion is permanent and cascading: the datastore's containers, checks, anomalies, and enrichment data are removed with it. The connection stays, and so does every other datastore that uses it.


Create Several Datastores at Once

A Glue Data Catalog usually holds more databases than one. The bulk endpoints take a list of databases and create one datastore per entry, sharing the same connection and settings.

Permission: Manager

Endpoint: POST /api/connections/datastores/bulk

Example request and response

Request:

curl -X POST "https://your-instance.qualytics.io/api/connections/datastores/bulk" \
  -H "Authorization: Bearer YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "connection": {
      "name": "acme_glue_us_east_1",
      "type": "glue_native",
      "parameters": {
        "authentication_type": "IAM_ROLE",
        "region": "us-east-1",
        "role_arn": "arn:aws:iam::123456789012:role/QualyticsGlueRead",
        "external_id": "<your-external-id>"
      }
    },
    "schemas": ["sales", "finance", "inventory"],
    "name_template": "glue_{{schema}}",
    "teams": ["Data Platform"],
    "trigger_sync": true
  }'

Response (200 OK):

{
  "created": [101, 102, 103],
  "errors": [],
  "connection_id": 12
}

Endpoint: POST /api/connections/{connection_id}/datastores/bulk

Example request and response

Request:

curl -X POST "https://your-instance.qualytics.io/api/connections/7/datastores/bulk" \
  -H "Authorization: Bearer YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "schemas": ["marketing", "logistics"],
    "name_template": "glue_{{schema}}",
    "teams": ["Data Platform"],
    "trigger_sync": true
  }'

Response (200 OK):

{
  "created": [110, 111],
  "errors": [],
  "connection_id": 7
}
Property Type Required Description
connection object New connection only The same connection object as in Create an AWS Glue Native Datastore.
schemas list[string] Yes The Glue databases to create datastores for, one datastore each.
name_template string No Naming pattern with {{schema}} standing for the database name. Defaults to the connection name followed by the database.
teams list[string] No Teams assigned to every created datastore.
tags list[string] No Tags applied to every created datastore.
group_id integer No An existing datastore group to place them in.
description string No Description applied to every created datastore.
trigger_sync boolean No Run the first Sync on each datastore as it is created.

Note

There is no database field to send: the region and Catalog ID on the connection already identify the catalog, so schemas is the whole address. Creation is not atomic. Databases that succeed are created even if others fail, and the response lists both the created IDs and, per failed database, the reason. It also returns connection_id, the connection the datastores were created on. To retry the failed databases, send them to that connection through the Existing Connection endpoint. When no datastore was created, connection_id is null. After an Existing Connection request, retry on the same connection ID. After a New Connection request, resend it; if the response is 409 with Connection '<name>' already exists with id: <id>, a connection with that name already exists, most likely the one your first attempt saved. Check it with GET /api/connections/{id} before reusing it, then retry the databases on that ID through Existing Connection.