Getting Started with Source Datastores
What is a Source Datastore?
In Qualytics, a Source Datastore is where your data lives. It is the connection Qualytics opens to look at the tables or files inside a system so it can profile them, run quality checks, and detect anomalies.
You may also encounter this concept under names like data source or data connection. In Qualytics, we always call it a Source Datastore, and every underlying technology is treated the same way: whether the system is a relational database, a data warehouse, or a cloud object storage bucket, a Source Datastore is the parent that owns a collection of containers (tables, files, or other data groupings).
Everything else in the platform starts from the Source Datastore: once it exists, you can Sync it to discover the containers inside, Profile those containers to understand their contents, and Scan them for anomalies.
Enrichment Datastores
Optionally link an enrichment datastore as the destination for scan results, source record examples, remediation snapshots, materialized copies, and exports. Create it on a supported connector using a dedicated schema, database, or storage location, separate from the source data you govern. See Getting Started with Enrichment for details.
How a datastore maps to your storage system
The exact slice of data a datastore represents depends on the connector:
- For relational databases (PostgreSQL, Snowflake, and other JDBC engines), a datastore points at one schema inside a specific database. If a database holds three schemas that you want to monitor, you create three datastores. The Add Datastore form asks for both the database and the schema name so Qualytics knows exactly where to look.
- For cloud object storage (Amazon S3, Azure Data Lake Storage, Google Cloud Storage), a datastore points at a folder path inside a bucket, so files under the same prefix are grouped together.
In every case the mental model is the same: the datastore is the boundary Qualytics reads from, and the tables or files inside that boundary are the containers you will work with once Sync completes.
Datastore Lifecycle
Every datastore in Qualytics follows a structured lifecycle, from initial creation through ongoing data quality operations.
graph TD
A["<b>Create</b><br/>Add datastore & test connection"] --> B["<b>Link Enrichment</b><br/>Choose the destination"]
B --> C["<b>Sync</b><br/>Discover tables, files & fields"]
C --> D["<b>Profile</b><br/>Analyze data & infer checks"]
D --> E["<b>Scan</b><br/>Run checks & detect anomalies"]
E -->|"Repeat"| C
| Stage | Description |
|---|---|
| Create | Add a new source datastore by selecting a connector, providing connection credentials, and testing the connection. |
| Link Enrichment | Optionally link an enrichment destination so scan results, source record examples, remediation snapshots, and exports are written to your own infrastructure. |
| Sync | Discover the schema (tables, files, views, and fields) from your source datastore. This is the first operation after creation. |
| Profile | Analyze field patterns across records, detect data types, compute statistics, and automatically infer quality checks. |
| Scan | Execute quality checks against the data, measure data quality metrics, and detect anomalies from failed checks. |
Tip
The Sync → Profile → Scan cycle is repeatable. As your data evolves, re-running these operations keeps your quality checks and anomaly detection up to date.
Run Sync from your source datastore
Start with Sync to discover tables and files. You can then run Profile and Scan from the datastore or selected containers. See the Sync Operation page for step-by-step instructions.
Guided First Runs
A newly created datastore shows a Get started in three steps strip at the top of its overview page, with Sync, Profile, and Scan as connected step cards. Each step unlocks when the previous one completes, and its card reflects the operation's live status: waiting as To do, then Queued, Running, and finally Completed, or Failed with a Retry button. Users who can run operations launch each step directly from its card. Once all three operations have completed successfully at least once, the datastore is fully initialized and the strip no longer appears.
Architecture
The Qualytics Datastore framework is organized in four layers:
| Layer | Description |
|---|---|
| Data access | Qualytics reads fields and records through the connector and writes operation results to a linked enrichment destination. |
| Namespace | How data is organized: tables in a schema for databases, or files in a folder for file storage. |
| Schema | Field definitions and types supplied by the database or discovered from supported file formats. |
| Storage Engine | The actual platform where data resides (see supported engines below). |
Supported Storage Engines
Supported Storage Engines
For the full list of supported JDBC and DFS connectors, see the Available Datastore Connectors page.
How It All Connects
graph LR
IO["<b>Data access</b><br/>Read fields and records<br/>Write results to a linked enrichment destination"]
IO --> NS1["<b>Tables in a Schema</b>"]
IO --> NS2["<b>Files in a Folder</b>"]
NS1 --> SC1["<b>RDBMS DDL</b>"]
NS2 --> SC2["<b>AVRO, Parquet, CSV, JSON</b>"]
SC1 --> SE1["Oracle DB"]
SC1 --> SE2["PostgreSQL"]
SC1 --> SE3["MySQL"]
SC1 --> SE4["Snowflake"]
SC1 --> SE5["Redshift"]
SC1 --> SE6["...and more"]
SC2 --> SE7["Amazon S3"]
SC2 --> SE8["Azure Storage"]
SC2 --> SE9["GCS"]
The key insight is that Qualytics treats all datastores uniformly at the data-access layer. Whether you connect a PostgreSQL database or an S3 bucket, Qualytics reads and writes Fields and Records the same way, enabling consistent profiling, scanning, and anomaly detection across your entire data landscape.
Available Connectors
Qualytics supports 20 JDBC relational database connectors and 3 DFS cloud storage connectors out of the box.
-
All Connectors
See the complete list of supported connectors with links to their individual setup guides.
Deep Dive
Explore the details of each datastore type: how JDBC and DFS connectors work, their configuration, and supported features.
-
JDBC
Learn how to connect relational databases using JDBC connectors like PostgreSQL, Snowflake, Oracle, and more.
-
DFS
Learn how to connect distributed file systems like Amazon S3, Azure Data Lake Storage, and Google Cloud Storage.
Connection
Learn how to set up and manage connections to your datastores: create new connections from scratch or reuse existing credentials.
-
Connections Overview
Set up new connections or reuse existing credentials to connect your datastores.
Multiple-Schema
Discover and onboard multiple schemas from a single connection at once, including schema discovery, name templates, supported connectors, and API reference.
-
Introduction
Discover and onboard multiple schemas from a single connection in one step.
-
How It Works
Understand the multi-schema creation flow, schema discovery, and name templates.
-
Supported Connectors
See which connectors support multi-schema discovery and their catalog/schema mappings.
-
Permissions
Understand the roles and permissions required for multi-schema creation.
-
API
API endpoints for bulk datastore creation, schema discovery, and validation.
-
FAQ
Answers to common questions about multi-schema source datastore creation.
How-tos
Add, edit, and delete datastores, whether creating from scratch with a new connection or reusing an existing one.
-
Add Datastore with new connection
Create a new source datastore by setting up a new connection from scratch.
-
Add Datastore with existing connection
Create a new source datastore by reusing credentials from an existing connection.
-
Edit Datastore
Modify the settings and connection details of an existing datastore.
-
Delete Datastore
Remove a datastore and its associated configuration from Qualytics.
Data Catalog Links
View and correct how the datastore links to assets in connected data catalogs.
-
Data Catalog Links
The link card, manual and automatic links, anchoring, and the scoped synchronization.