AWS Glue Native Examples
Real scenarios showing what AWS Glue Native does with a Glue Data Catalog arranged in different ways. Pick a tab to see what to expect from the Sync, the Profile, and the Scan in each situation.
Reading a partitioned Parquet data lake
Context. A team keeps its sales tables in a Glue database called sales, written as Parquet to s3://acme-lake/sales/ and partitioned by day. The tables were registered by a Glue crawler. They add an AWS Glue Native datastore in us-east-1 with Assumed Role authentication and select sales as the database.
What happens. The Sync lists every table in sales and records each schema from the catalog. The first Profile is where S3 is read in earnest. When a Scan filters to a single day, Qualytics asks Glue for the partitions matching that day and reads only their files, instead of listing the whole table location and discarding the rest.
Why it works. The table definitions live in Glue and the data lives in S3, and one role covers both. There is no Athena workgroup or query-results bucket involved, because no query is ever sent to Athena.
A database where only some tables can be read natively
Context. A realistic Glue database rarely holds Parquet alone. This one has seven tables, and the team wants to know what they will get before they connect it.
| Table | Registered as | Result |
|---|---|---|
orders |
Parquet, partitioned | Read |
clickstream |
JSON | Read |
vendors_csv |
CSV with a header line (OpenCSVSerde) |
Read |
orders_iceberg |
Iceberg | Read |
v_open_orders |
Athena view | Declined |
events_projected |
Parquet with partition projection enabled | Declined |
trips_hudi |
Hudi | Declined |
What happens. The Sync succeeds. The first four tables become containers to profile and scan. The other three are declined, and none of them holds up the rest of the database. The team then connects the same catalog with the Athena connector and points a second datastore at the same database for the view, the projected table, and the Hudi table.
Why it works. Each of the three is declined because reading its files directly would disagree with what Athena returns. A view has no files of its own, a projected table's partitions are computed at query time rather than registered, and a Hudi table's own metadata decides which files are current. Running both connectors against one catalog is a supported arrangement, not a workaround.
A catalog owned by another AWS account
Context. A central data account, 111122223333, owns the Glue Data Catalog and the S3 buckets. The analytics team works from account 444455556666, where the role Qualytics assumes lives. They set Catalog ID to 111122223333.
What happens. Every Glue call names the central account's catalog, and the S3 reads go to the central account's buckets, all with the analytics account's role. The central account's Glue resource policy and bucket policies allow that role; the role's own IAM policy allows the same actions on the central account's resources.
Why it works. Cross-account access in AWS needs both sides to agree. Catalog ID only chooses which account's catalog to read; it does not grant anything. Leaving it empty would read the analytics account's own, probably empty, catalog. See Cross-account catalogs.
Open table formats registered in Glue
Context. Some tables in the database are Apache Iceberg tables maintained by AWS Glue jobs, and some are Delta Lake tables written by another team. All of them are registered in Glue.
What happens. Both kinds are read. For an Iceberg table, Qualytics reads the snapshot the Glue entry currently points to. For a Delta table, it reads the table's transaction log to find the current files. When a writer commits a new version, the next Profile or Scan reads that version.
Why it works. Both formats record which files belong to the current version of the table, so reading the table directory as plain files would include deleted and replaced files. Reading through each format's own metadata avoids that. The role only needs S3 read access to the table location, which already holds that metadata.
Two teams, two roles, one bucket
Context. Finance and marketing each have their own Glue database, their own IAM role, and data under different prefixes of the same S3 bucket. Each role can read only its own prefix.
What happens. Two connections, one per role, each with its own AWS Glue Native datastore. Every read a datastore makes uses its own connection's role, including reads from the shared bucket. The finance datastore cannot read marketing's prefix, and the other way round.
Why it works. Credentials live on the connection, and Qualytics keeps each connection's identity separate even when two datastores read the same bucket. An access problem shows up in AWS CloudTrail under the role that made the call.
See Also
-
Best Practices
Recommendations for scoping the AWS identity, splitting tables between AWS Glue Native and Athena, and keeping scans fast.
-
Permissions
The Glue, Amazon S3, and STS access your AWS account must grant, plus the Qualytics user role and team permission needed.
-
Authentication
Access Key or Assumed Role: the AWS identity Qualytics uses for Glue and Amazon S3, and how its credentials are stored.