Add a Computed File
Add a Computed File to a DFS source datastore. The Computed File becomes a container inside that datastore, alongside the base file patterns, and can be profiled and monitored like any other container.
Related
- Concept: Introduction for what a Computed File is.
- Mechanics: How It Works for validation, referencing rules, and execution.
- Comparison: Computed File vs Computed Table if you are not sure which one you need.
Profile the source first
The base file pattern you want to build on must already have been profiled at least once. Qualytics uses the profiled schema to validate every field reference in the Select Expression. If the source is not yet profiled, open it in the tree view, run Scan on it once, then come back to the Add Computed File modal.
Field reference
The Add Computed File modal has four sections. Only Name, Source File, and Select Expression are required; the rest can be filled in later with Edit.
Name, Source File, and Select Expression
| Field | Required | Type | Description |
|---|---|---|---|
| Name | Text | A descriptive, meaningful name for the Computed File (for example, customer_order_statistics). Spaces are replaced with underscores automatically. Must be unique inside the datastore. |
|
| Source File | Option | The base file pattern this Computed File is built on. The dropdown lists file patterns in this datastore that have been profiled at least once; unprofiled files and other computed containers are filtered out. | |
| Select Expression | Text | The Spark SQL projection defining the shape of the Computed File. Field references must exist in the source's profiled schema. The DISTINCT keyword is not allowed in any form (including COUNT(DISTINCT ...), APPROX_COUNT_DISTINCT(...), and DISTINCT(col)); use GROUP BY on the same columns for de-duplication, or explode(collect_set(field)) AS distinct_field_values for distinct-value semantics. Press Ctrl+Space inside the editor for hints. See How It Works for the full referencing rules. |
Filter Clause, Group By Clause, and Description
| Field | Required | Type | Description |
|---|---|---|---|
| Filter Clause | Text | A Spark SQL WHERE predicate to narrow the rows. Standard Spark SQL boolean expressions are supported. |
|
| Group By Clause | Text | A Spark SQL GROUP BY clause. Required when the Select Expression uses aggregation functions (COUNT, SUM, AVG, MIN, MAX). Every non-aggregated column in the Select Expression must appear here. |
|
| Description | Text | A short business summary of what the Computed File represents. The field opens a Markdown editor (up to 8,000 characters) that can be pinned to the side of the page, so the expression and the description stay visible together. |
Let AgentQ write the description
When AgentQ is configured for your deployment, a Suggest button appears next to Description. Click it and AgentQ drafts a description from the Computed File's current form values; review the text, adjust it if needed, and save as usual. See Description Suggestions for the full workflow.
Owner and Additional Metadata
| Field | Required | Type | Description |
|---|---|---|---|
| Owner | Option | The Qualytics user who owns this Computed File. Defaults to you when left empty. Only users with the Editor team permission can assign a different owner; users with the Author team permission can create Computed Files they own but cannot reassign later. | |
| Additional Metadata | Key-value | Custom key-value pairs. Click the plus icon to open the editor and add rows. Keys also act as default values for runtime variables in this container's checks. |
Advanced Options
| Field | Required | Type | Description |
|---|---|---|---|
| Lateral Views | Text | Located under Advanced Options. Each entry is a Spark SQL expression such as EXPLODE(tags) t AS tag; Qualytics prepends LATERAL VIEW automatically. Use this when the source has array-valued columns you want to expand into rows. |
|
| Add Lateral View | Button | Adds another lateral view entry to the list so you can chain multiple explodes or posexplode operations in the same Computed File. Each entry has a remove icon. |
Each entry must begin with one of the supported generator functions:
EXPLODE, POSEXPLODE, INLINE, STACK, JSON_TUPLE, and PARSE_URL_TUPLE.
Any of these can be written with the OUTER keyword in front of it, for example OUTER EXPLODE(tags) t AS tag, which keeps rows whose array is empty or null instead of dropping them. EXPLODE_OUTER, POSEXPLODE_OUTER, and INLINE_OUTER are accepted as single words and mean the same thing.
Steps
Step 1: Open the datastore and click the Add button in the top-right of the overview. A dropdown lists the containers you can add to a DFS datastore.
Step 2: Select Computed File. This option appears only when the parent datastore is a DFS type (S3, GCS, or Azure Data Lake Storage). The Add Computed File modal opens.
Step 3: Fill in the fields. See the Field reference above for what each input controls.
Step 4: Click Validate. A validation success message appears when every clause parses and resolves against the source; otherwise the modal shows an inline error that quotes the failed clause so you can fix it in place.
What Validate checks
Qualytics parses every clause as Spark SQL and confirms that every referenced field exists in the source's profiled schema and that the source is still reachable on storage. Running Validate is optional: if you skip it, the same validation runs as part of Save and any error is shown inline.
Validation timeouts
Validate returns 408 Request Timeout when the check does not complete within 150 seconds; Save allows 500 seconds for the same validation. Simplifying the filter or narrowing the source pattern usually resolves timeouts on either path.
Step 5: Once the validation success message appears, click Save.
Step 6: A success message confirms the creation, the modal closes, and the new Computed File appears in the datastore's container list. Qualytics automatically runs a slim profile of up to 1000 records per partition so you can see field statistics right away, followed by a full profile in the background.