Skip to content

Introduction to the AWS Glue Native Connector

The AWS Glue Data Catalog is the metadata service most AWS data lakes are built on. It records, for every database and table, the columns, the file format, and the Amazon S3 location of each table and partition. Query services such as Amazon Athena, Amazon EMR, and AWS Glue jobs all read that catalog to find the data they query.

The Qualytics AWS Glue Native connector reads the table definitions from the Glue Data Catalog and then reads the underlying files directly from Amazon S3. No query engine sits in the read path: there is no Athena workgroup to choose, no query-results bucket to provide, and no Athena query charges. Partitions come from the catalog rather than from listing S3, so Qualytics profiles tables, runs scheduled scans, and surfaces record- and schema-level anomalies on the same rows the catalog describes.


Source datastores only

AWS Glue Native is read-only, so it cannot be used as an enrichment datastore. Link a separate enrichment datastore as its destination to store the anomalies and metadata Qualytics produces, either during creation or afterwards.

Notes

Choosing between AWS Glue Native and Athena. Both connectors read tables registered in the Glue Data Catalog, and both can be connected against the same catalog at the same time. Use AWS Glue Native for the ordinary file-backed tables that make up most data lakes, and the Athena connector for the tables the native path declines, such as views and partition-projection tables.

AWS Glue Native Athena
How it reads data Reads the table's files directly from Amazon S3 Runs SQL queries through Athena
Athena setup None Workgroup and a query-results S3 bucket
Views Not supported Supported
Partition projection tables Not supported Supported
Hudi tables Not supported Supported
Lake Formation filtering Not applied Applied by Athena
Authentication Access Key or Assumed Role Access Key or IAM Role

Supported tables and formats. The connector reads non-transactional tables whose data lives in Amazon S3, partitioned or not:

  • Parquet and ORC.
  • Delimited text (CSV and TSV) and JSON, using the SerDe settings recorded on the table, such as the field delimiter, quote character, null marker, and a header line.
  • Avro and other Hive SerDe formats, such as RCFile, SequenceFile, and regular-expression tables.
  • Apache Iceberg tables registered in Glue, read at the snapshot the catalog currently points to.
  • Delta Lake tables registered in Glue, read through their transaction log.

What is declined. Anything the connector cannot read exactly as the catalog defines it is declined rather than read incorrectly. A declined table does not become a container, and it does not hold up the rest of the database, which syncs normally.

Declined Why
Views The native path reads tables, not query definitions
Transactional (ACID) Hive tables Reading the files directly would ignore delete records
Hudi tables Hudi's own metadata decides which files are current
Athena partition-projection tables Their partitions are computed at query time rather than registered, so reading them directly could silently return only some of the rows
Tables backed by a storage handler, or registered from a Glue JDBC connection Their data is only reachable through the other system
Tables or partitions stored outside Amazon S3 This connector reads Amazon S3 only
Tables whose file format or SerDe Qualytics does not ship There is no reader for that format

Read any of these through the Athena connector instead.

Amazon S3 only, by design. The Glue Data Catalog can record locations outside Amazon S3, but this connector only reads S3 locations. That is a permanent part of how it works: it is what lets a single AWS identity cover both the catalog and the data, with no Kerberos or cluster network access to set up.

Lake Formation. Qualytics reads with the permissions of the AWS identity you configure. It does not request credentials from AWS Lake Formation, and it does not apply Lake Formation row, column, or cell filters. The identity needs direct access to the Glue tables and to their S3 locations. See Permissions before connecting a catalog governed by Lake Formation.

What the form does not ask for. There is no Catalog field of the kind Athena asks for. Each AWS account has one Glue Data Catalog per region, so Region and, for another account's catalog, Catalog ID are enough to identify it.

Row counts. Qualytics counts rows itself rather than trusting a row count recorded in the table's Glue properties, because nothing guarantees that recorded value is current.


Deep Dive

  • Authentication


    Access Key or Assumed Role: the AWS identity Qualytics uses for Glue and Amazon S3, and how its credentials are stored.

    Authentication

  • Examples


    Worked scenarios: a partitioned Parquet lake, a database of mixed tables, a cross-account catalog, and Iceberg and Delta tables.

    Examples

  • Best Practices


    Recommendations for scoping the AWS identity, splitting tables between AWS Glue Native and Athena, and keeping scans fast.

    Best Practices

  • Permissions


    The Glue, Amazon S3, and STS access your AWS account must grant, plus the Qualytics user role and team permission needed.

    Permissions


How-tos

  • Add with New Connection


    Add AWS Glue Native as a source datastore and create its connection at the same time, field by field.

    Add with New Connection

  • Add with Existing Connection


    Add another Glue database to monitor, reusing a saved AWS Glue Native connection.

    Add with Existing Connection


Reference

  • Troubleshooting


    Known problems and how to resolve them, grouped by the step that failed: the catalog, the AWS identity, Amazon S3, or a table.

    Troubleshooting

  • API


    Create, test, read, update, and delete AWS Glue Native datastores over the API, including bulk creation from several databases.

    API

  • FAQ


    Short answers to the questions that come up most often about Athena, formats, Lake Formation, and cross-account access.

    FAQ