AWS Glue Native Best Practices
Recommendations for getting reliable results and predictable run times from an AWS Glue Native datastore.
Plan the connection
Confirm S3 access before the first profile, not after. An identity that can read Glue but not the S3 objects passes the connection test and lists every table, then fails on the first operation that reads records. Checking the S3 permissions on every bucket your tables point to turns a confusing failure into a five-minute setup step. See Permissions.
Use one connection per identity and reuse it. Every datastore that reads with the same AWS identity should share one saved connection. The credentials then live in one place, so changing them is a single edit rather than one per database. Add further databases with Add with an Existing Connection.
Set Catalog ID only for another account's catalog. Leave it empty to read the catalog of the account the credentials belong to. A Catalog ID that names another account needs that account's resource and bucket policies in place as well, so filling it in by habit only adds a way for the connection to fail.
Choose the right connector per table
Send ordinary file-backed tables through AWS Glue Native and the rest through Athena. The native path reads the files directly and needs no Athena setup; it is also the one that declines views, partition-projection tables, and Hudi tables. Pointing both connectors at the same catalog, and splitting the tables between them, gets you direct reads where they apply and full coverage where they do not.
Do not treat a declined table as a failure to work around. A table is declined when reading its files directly would disagree with what the catalog describes. Reading it through the Athena connector is almost always simpler than reshaping the table.
Keep Lake Formation-filtered tables on Athena. AWS Glue Native does not apply Lake Formation row, column, or cell filters. A table whose protection depends on those filters should be read through Athena, which applies them.
Keep operations fast
Register partitions in Glue. Qualytics reads the partitions the catalog records, so a partition that only exists in S3 is not read until it is registered, for example by a crawler, by MSCK REPAIR TABLE, or by the job that writes it.
Filter scans on partition columns. A scan whose conditions name the partition columns lets Qualytics ask Glue for just those partitions rather than fetching the whole partition list, which is the difference between reading one day and enumerating every day the table holds.
Narrow the datastore to the databases you actually monitor. A database with a very large number of partitions spends real time on partition metadata. Selecting only the databases you intend to profile keeps sync and scan times proportional to what you care about.
Prefer partitioned, columnar tables for the data you monitor. Parquet and ORC are the formats Qualytics reads fastest, and partitioning is what lets a scan read a slice instead of the whole table.
Keep the AWS identity manageable
Prefer Assumed Role in production. A role needs no stored long-lived keys, its temporary credentials are renewed automatically, and every session it opens is recorded in CloudTrail. Require an External ID in its trust policy.
Scope the identity to what Qualytics monitors. Grant the Glue actions on the databases and tables you connect, and the S3 actions on their buckets and prefixes. An identity scoped that way makes an access problem obvious in CloudTrail rather than hidden behind a shared account.
Use a separate identity per team or data domain. Each connection keeps its identity separate, even when two datastores read the same bucket. Giving each team its own role means each datastore can read exactly what that team is allowed to, and nothing more.
See Also
-
Permissions
The Glue, Amazon S3, and STS access your AWS account must grant, plus the Qualytics user role and team permission needed.
-
Authentication
Access Key or Assumed Role: the AWS identity Qualytics uses for Glue and Amazon S3, and how its credentials are stored.
-
Examples
Worked scenarios: a partitioned Parquet lake, a database of mixed tables, a cross-account catalog, and Iceberg and Delta tables.