AWS Glue Native Permissions
Two separate things have to be granted before someone can put an AWS Glue Native datastore to work, and they are granted in different places:
- In your AWS account, the identity Qualytics connects with needs read access to the Glue Data Catalog and to the S3 objects behind each table. Without it the connection cannot read your data at all.
- In Qualytics, the person doing the work needs a user role and a team permission. Without them the datastore exists but they cannot create, change, or run operations on it.
The rest of this page covers both, in that order.
Access on your AWS Account
Because AWS Glue Native reads the data files themselves, it needs more than access to the catalog. Glue access alone is enough to list databases and tables, so a datastore set up that way passes the connection test and then fails as soon as Qualytics needs to read the files. Grant the S3 access below before the first operation, not after.
Qualytics never writes to Glue or to Amazon S3 through this connector, so read access is all it asks for.
Which identity needs the permissions
| Authentication mode | Identity that carries the Glue and S3 permissions |
|---|---|
| Access Key | The IAM user behind the Access Key ID and Secret Access Key |
| Assumed Role | The target role (the Role ARN entered in the connection form). The base role (the AWS identity used by Qualytics) only needs sts:AssumeRole on that target role, and the target role's trust policy must allow it; see IAM Role Authentication. Available on AWS and local deployments only. |
Glue permissions
| Permission | Purpose |
|---|---|
glue:GetDatabases |
List the databases offered in the Database dropdown |
glue:GetDatabase |
Read the selected database |
glue:GetTables |
List the tables in the database during connection tests and Sync |
glue:GetTable |
Read a table's columns, format, and location |
glue:GetPartitions |
Read the partitions of a partitioned table, and their locations |
Glue checks each call against the resource and every level above it, so a table read needs the catalog, the database, and the table to be allowed. To scope the policy, grant the catalog ARN, the database ARN, and the tables in that database, for example arn:aws:glue:us-east-1:123456789012:catalog, arn:aws:glue:us-east-1:123456789012:database/sales, and arn:aws:glue:us-east-1:123456789012:table/sales/*.
Amazon S3 permissions
| Permission | Resource | Purpose |
|---|---|---|
s3:ListBucket |
arn:aws:s3:::<bucket> |
List the files under each table and partition location |
s3:GetObject |
arn:aws:s3:::<bucket>/<prefix>/* |
Read the data files, plus the Iceberg metadata and Delta transaction log files for those formats |
Grant them on every bucket and prefix your tables point to. Table locations in one database can span several buckets, and a partition can be registered at a location outside its table's own path. Qualytics reads both, so both need to be covered.
Encrypted buckets
If a bucket uses SSE-KMS with a customer-managed key, the identity also needs kms:Decrypt on that key. Without it, reads fail with an access error even when every Glue and S3 permission above is granted.
Example IAM policy
Replace the region, account ID, database, and bucket placeholders with your own values. The policy covers one database: repeat the two database lines for each additional database you monitor. Only the databases the policy covers appear in the Database dropdown.
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "GlueCatalogReadOnly",
"Effect": "Allow",
"Action": [
"glue:GetDatabases",
"glue:GetDatabase",
"glue:GetTables",
"glue:GetTable",
"glue:GetPartitions"
],
"Resource": [
"arn:aws:glue:<region>:<account-id>:catalog",
"arn:aws:glue:<region>:<account-id>:database/<database>",
"arn:aws:glue:<region>:<account-id>:table/<database>/*"
]
},
{
"Sid": "S3ListTableLocations",
"Effect": "Allow",
"Action": "s3:ListBucket",
"Resource": "arn:aws:s3:::<YOUR_BUCKET>"
},
{
"Sid": "S3ReadTableFiles",
"Effect": "Allow",
"Action": "s3:GetObject",
"Resource": "arn:aws:s3:::<YOUR_BUCKET>/*"
}
]
}
Cross-account catalogs
When Catalog ID names another AWS account, both accounts take part:
- The account that owns the catalog grants your identity access through the Glue Data Catalog resource policy, and grants read access to its S3 buckets through their bucket policies.
- Your account grants the same Glue and S3 actions in the identity's IAM policy, using the owning account's ID in the Glue ARNs.
A call succeeds only when both sides allow it.
Lake Formation
Qualytics reads with the identity's own IAM permissions. It does not request temporary credentials from AWS Lake Formation, and it does not apply Lake Formation row, column, or cell filters.
- Catalogs that use IAM access control only (the default for Glue, where Lake Formation grants the
IAMAllowedPrincipalsgroup access) work with the permissions above and nothing more. - Tables governed by Lake Formation permissions also need Lake Formation to let the identity see them in the catalog, and the identity still needs the direct S3 permissions above, because Qualytics reads the files with its own credentials.
Filtered tables
If a table relies on Lake Formation data filters to hide rows or columns, those filters are not applied to what AWS Glue Native reads: the identity reads whatever its S3 permissions allow. Read such tables through the Athena connector, which applies Lake Formation filtering.
Permissions in Qualytics
Qualytics checks two layers, and an action is allowed only when both pass.
- User roles are workspace-wide: Member, Manager, and Admin.
- Team permissions are granted per datastore through the teams a user belongs to: Reporter, Viewer, Drafter, Author, and Editor.
For the full reference, see User Roles and the Team Permissions Overview.
User Roles
| Action | Member | Manager | Admin |
|---|---|---|---|
| View the datastore and its activity | |||
| Create an AWS Glue Native datastore | |||
| Edit the datastore's settings | |||
| Run and schedule Sync, Profile, and Scan | |||
| Link an enrichment destination | |||
| Unlink an enrichment destination | |||
| Delete the datastore |
Team Permissions
Checked on the datastore itself, in addition to the user role above.
| Action | Reporter | Viewer | Drafter | Author | Editor |
|---|---|---|---|---|---|
| View the datastore and its activity | |||||
| Preview the datastore's data | |||||
| Edit the datastore's settings | |||||
| Run and schedule Sync, Profile, and Scan | |||||
| Link an enrichment destination |
How the two layers meet
To create an AWS Glue Native datastore, a user needs the Manager role. Team permissions are not consulted, because the datastore does not exist yet.
To edit it, run an operation on it, or link an enrichment destination, a user needs the Member role or higher and the Editor team permission on that datastore, through at least one of the teams the datastore belongs to.
To delete it, or to unlink its enrichment destination, a user needs the Admin role. No team permission is consulted.
Admins bypass team permissions
A user with the Admin role passes every team permission check, on any datastore, whether or not they belong to a team that owns it.
Who should hold what
Give the Editor team permission to the people who own the Glue datastores day to day, since it is what gates running a Sync after a schema change. Reserve the Manager role for whoever onboards new catalogs, and the Admin role for the few people who should be able to remove a datastore.
See Also
-
Authentication
Access Key or Assumed Role: the AWS identity Qualytics uses for Glue and Amazon S3, and how its credentials are stored.
-
Examples
Worked scenarios: a partitioned Parquet lake, a database of mixed tables, a cross-account catalog, and Iceberg and Delta tables.
-
Best Practices
Recommendations for scoping the AWS identity, splitting tables between AWS Glue Native and Athena, and keeping scans fast.