How Before Date Time Checks Work
Definition
Asserts that every value in a Date or Timestamp field is strictly earlier than a chosen cutoff date and time.
Overview
The Before Date Time rule defines an upper time boundary on a single date or timestamp field. Each row in the target container must hold a value that is less than the cutoff; any value equal to or later than the cutoff fires an anomaly. The comparison is strict (<), not inclusive (<=), and the cutoff is stored as a UTC instant.
Typical use cases:
- Validate historical data cutoffs, such as an archive that must not contain rows past a migration date.
- Ensure a date precedes a business boundary, such as a creation date before an effective date or a period close.
- Catch implausible future-dated values, such as a birth date entered as today.
Field Scope
Single: The rule evaluates exactly one field per check.
Accepted Types
| Type | Supported |
|---|---|
Date |
|
Timestamp |
General Properties
| Name | Supported |
|---|---|
Filter Allows the targeting of specific data based on conditions |
|
Coverage Customization Allows adjusting the percentage of records that must meet the rule's conditions |
The filter allows you to define a subset of data upon which the rule will operate.
It requires a valid Spark SQL expression that determines the criteria rows in the DataFrame should meet. This means the expression specifies which rows the DataFrame should include based on those criteria. Since it's applied directly to the Spark DataFrame, traditional SQL constructs like WHERE clauses are not supported.
Examples
Direct Conditions
Simply specify the condition you want to be met.
Combining Conditions
Combine multiple conditions using logical operators like AND and OR.
Correct usage" collapsible="true
Incorrect usage" collapsible="true
Utilizing Functions
Leverage Spark SQL functions to refine and enhance your conditions.
Correct usage" collapsible="true
Incorrect usage" collapsible="true
Using scan-time variables
To refer to the current dataframe being analyzed, use the reserved dynamic variable {{_qualytics_self}}.
Correct usage" collapsible="true
Incorrect usage" collapsible="true
While subqueries can be useful, their application within filters in our context has limitations. For example, directly referencing other containers or the broader target container in such subqueries is not supported. Attempting to do so will result in an error.
Important Note on {{_qualytics_self}}
The {{_qualytics_self}} keyword refers to the dataframe that's currently under examination. In the context of a full scan, this variable represents the entire target container. However, during incremental scans, it only reflects a subset of the target container, capturing just the incremental data. It's crucial to recognize that in such scenarios, using {{_qualytics_self}} may not encompass all entries from the target container.
Specific Properties
Before Date Time has one rule-specific property:
| Name | Description |
|---|---|
Date |
The cutoff date and time. The check flags every row whose field value is not strictly earlier than this value. |
Anomaly Types
| Type | Supported |
|---|---|
| Record Flag inconsistencies at the row level |
|
| Shape Flag inconsistencies in the overall patterns and distributions of a field |
Evaluation Flow
Every Before Date Time check follows the same four-step evaluation flow:
- Apply the filter clause. If the check has a
filterset, only rows matching the filter expression continue to the next step. Rows outside the filter are ignored and cannot contribute to the violation count. - Read the field value as a timestamp. The platform interprets each row's value as a timestamp for comparison; values that cannot be interpreted are treated as failing the comparison (see below).
- Compare against the cutoff. The cutoff (stored as a UTC instant) is compared to the field value using strict less-than (
<). A row passes only when the field value is strictly earlier than the cutoff. - Apply coverage. At 100% coverage, any failing row causes the check to fail. Below 100% coverage, the check fails only when the passing fraction drops below the threshold (see Coverage and Tolerance).
The comparison is strict, so a row whose field value equals the cutoff is flagged. To get inclusive semantics, set the cutoff one second (for timestamps) or one day (for dates) later than the boundary you want to allow.
NULL Handling
The check passes NULL values: a row with NULL in the evaluated field is not counted as a violation. This is intentional. Before Date Time only asserts that present values fall before the cutoff and does not enforce mandatory presence on the field.
If the field must also be populated, pair Before Date Time with a Not Null check on the same field. Together they enforce "the field must be present AND must be earlier than the cutoff".
Values that cannot be interpreted as a timestamp are flagged as violations. Because the field picker accepts only Date and Timestamp columns, this rarely happens in practice.
Cutoff and Time Zone Behavior
The cutoff is stored as a UTC instant. When you set the cutoff to 2025-12-01 06:15:00 in the UI, the platform records it as 2025-12-01T06:15:00Z and uses that exact instant for every evaluation.
The platform evaluates the comparison in UTC:
Timestampfields stored with explicit time zone offsets are compared directly against the UTC cutoff.Timestampfields without an explicit time zone are interpreted as UTC instants before the comparison.Datefields are interpreted as midnight UTC of that calendar day, then compared to the cutoff instant.
If your source data carries values in mixed or local time zones, normalize the field upstream with a Computed Field before running Before Date Time so the comparison is unambiguous.
The Filter Clause
The filter clause is a SQL WHERE expression applied before the evaluation. Filtered-out rows are ignored entirely (they cannot trigger a violation and are not counted in the totals).
Common uses:
- Restricting the boundary to a subset of the dataset (
tenant_id = 42,status = 'active'). - Excluding known-bad legacy rows that are tracked by a separate clean-up task.
- Scoping the check to a partition (
event_date >= '2025-12-01').
When a filter is set, both the Record Anomaly and the Shape Anomaly messages end with [filter: <expression>] so the evaluated scope is visible in the alert.
Coverage and Tolerance
Coverage is a fractional value between 0 and 1 that defines the minimum fraction of evaluated rows that must pass:
1.0(100%, default): every row in the filtered set must pass. Any failing row causes the check to fail. This is the strictest setting.< 1.0: the check tolerates a fraction of rows failing the comparison. The check fails only when the fraction of passing rows drops below the threshold.
Lower coverage values are useful when a small, known fraction of pre-cutover rows is expected (for example, during a slow migration). Use coverage carefully: a 0.5% tolerance can mask a real regression that happens to fall just under the threshold.
See Also
-
Anomaly Reporting
The anomaly messages the check produces, what the numbers mean, Source Records highlighting, and Custom Anomaly Description.
-
Examples
Three production scenarios with sample data, anomaly messages, and the SQL equivalent of what the check evaluates.
-
Best Practices
Guidelines for choosing the cutoff, normalizing time zones, pairing rules, and keeping the signal clean.
-
Permissions
The team permission each action needs: view, create, edit, archive, restore, and delete.