Skip to content

AWS Glue Native FAQ

Answers to common questions about the AWS Glue Native connector. Sections are grouped by topic; pick the heading closest to what you are working on.

Choosing a connector

Does AWS Glue Native replace the Athena connector?

No. Both stay available, and they are good at different things. AWS Glue Native reads the files behind ordinary Glue tables directly; the Athena connector runs queries through Athena, so it also covers views, partition-projection tables, Hudi tables, and Lake Formation filtering.

Can I connect both against the same catalog?

Yes, at the same time. A common arrangement is AWS Glue Native for the bulk of the Parquet, ORC, CSV, and JSON tables, and Athena for the handful of views or projected tables alongside them.

Do I have to migrate my existing Athena datastores?

No. An existing Athena datastore keeps working exactly as before. Adding AWS Glue Native is a new datastore on a new connection, not a change to the old one.

Do I need an Athena workgroup or a query-results bucket?

No. AWS Glue Native sends no queries to Athena, so there is no workgroup to choose and no results bucket to provide.


What it reads

Which table formats are supported?

Parquet, ORC, delimited text such as CSV and TSV, JSON, Avro, and other Hive SerDe formats that Qualytics ships, plus Iceberg and Delta Lake tables registered in Glue. Tables can be partitioned or not. See Supported tables and formats.

Why is one of my tables missing after a Sync?

Because reading it directly would return something different from what the catalog describes, so Qualytics declines it instead. That covers views, transactional Hive tables, Hudi tables, Athena partition-projection tables, tables backed by a storage handler or a Glue JDBC connection, tables stored outside Amazon S3, and formats Qualytics does not ship. Read those through the Athena connector.

Does one declined table break the sync?

No. Each declined table is skipped on its own, and every other table in the database syncs normally.

Will it ever read tables stored outside Amazon S3?

No. Reading Amazon S3 only is a deliberate part of how the connector works, not a first-version limit. It is what lets one AWS identity cover both the catalog and the data.

Which version of an Iceberg or Delta table is read?

The current one: the snapshot the Glue entry points to for Iceberg, and the latest version in the transaction log for Delta. Earlier versions are not read.

Can Qualytics write to my Glue catalog or S3 buckets?

No. The connector is read-only, so it cannot be used as an enrichment datastore. Link a separate enrichment datastore as its destination for the anomalies and metadata Qualytics produces.


Connecting

What is Catalog ID for?

Reading another AWS account's Glue Data Catalog. Enter that account's 12-digit ID. Leave it empty to read the catalog in the account the credentials belong to. Both accounts must allow the access; see Cross-account catalogs.

Why is there no Catalog field?

Each AWS account has one Glue Data Catalog per region, so Region and, when needed, Catalog ID already identify it. There is nothing else to name.

The connection test passed, but the profile failed. Why?

The connection test reads only the catalog, not the data files. An identity that can read Glue but not the S3 objects passes the test and then fails on the first operation that reads data. Grant the S3 permissions on every bucket the tables point to. See Troubleshooting.

Does it work with Lake Formation?

Partly. Qualytics reads with the identity's own Glue and S3 permissions. It does not use Lake Formation credential vending, and it does not apply Lake Formation row, column, or cell filters. See Lake Formation.


Authentication

Which authentication should I use?

Assumed Role in production, where it is available: nothing long-lived is stored, and the temporary credentials are renewed automatically. Access Key is the simpler option for trials and development accounts, and the only one on Azure and GCP deployments.

Can I use separate credentials for Glue and for S3?

No. One identity reads both the catalog and the data, so it needs both sets of permissions.

Can two datastores use different roles on the same bucket?

Yes. Each connection keeps its own identity, and a datastore reads only with its own connection's credentials, even when another datastore reads the same bucket.

What happens to my keys if I switch a connection to Assumed Role?

The stored Access Key ID and Secret Access Key are discarded, and the role becomes the only identity the connection uses.