How Max Length Checks Work
Definition
Asserts that the number of characters in every value of a text field is at most a configured length.
Overview
The Max Length rule measures the character count of a single text field and compares it against the configured Length. Each row passes when its value's length is at most that number, using <=. The comparison counts characters as stored, including spaces and punctuation.
Typical use cases:
- Guard a fixed-width code such as a country, currency, or status code.
- Keep values inside the limit of a downstream system that would silently truncate them.
- Catch pasted content in a column meant for short notes.
Field Scope
Single: The rule evaluates exactly one field per check.
Accepted Types
| Type | Supported |
|---|---|
String |
|
Array |
General Properties
| Name | Supported |
|---|---|
Filter Allows the targeting of specific data based on conditions |
|
Coverage Customization Allows adjusting the percentage of records that must meet the rule's conditions |
The filter allows you to define a subset of data upon which the rule will operate.
It requires a valid Spark SQL expression that determines the criteria rows in the DataFrame should meet. This means the expression specifies which rows the DataFrame should include based on those criteria. Since it's applied directly to the Spark DataFrame, traditional SQL constructs like WHERE clauses are not supported.
Examples
Direct Conditions
Simply specify the condition you want to be met.
Combining Conditions
Combine multiple conditions using logical operators like AND and OR.
Correct usage" collapsible="true
Incorrect usage" collapsible="true
Utilizing Functions
Leverage Spark SQL functions to refine and enhance your conditions.
Correct usage" collapsible="true
Incorrect usage" collapsible="true
Using scan-time variables
To refer to the current dataframe being analyzed, use the reserved dynamic variable {{_qualytics_self}}.
Correct usage" collapsible="true
Incorrect usage" collapsible="true
While subqueries can be useful, their application within filters in our context has limitations. For example, directly referencing other containers or the broader target container in such subqueries is not supported. Attempting to do so will result in an error.
Important Note on {{_qualytics_self}}
The {{_qualytics_self}} keyword refers to the dataframe that's currently under examination. In the context of a full scan, this variable represents the entire target container. However, during incremental scans, it only reflects a subset of the target container, capturing just the incremental data. It's crucial to recognize that in such scenarios, using {{_qualytics_self}} may not encompass all entries from the target container.
Specific Properties
Max Length has two rule-specific properties:
| Name | Description |
|---|---|
Length |
The length the value is compared against. A value's character count must be at most this number. |
Array Element Context |
Measures each element of an array field instead of the array as a whole. Only relevant when the selected field is an array. |
Anomaly Types
| Type | Supported |
|---|---|
| Record Flag inconsistencies at the row level |
|
| Shape Flag inconsistencies in the overall patterns and distributions of a field |
Evaluation Flow
Every Max Length check follows the same four-step evaluation flow:
- Apply the filter clause. If the check has a
filterset, only rows matching the filter expression continue to the next step. Rows outside the filter are ignored and cannot contribute to the violation count. - Measure the value. The platform counts the characters in the row's value, exactly as stored.
- Compare against the length. The row passes when the character count is at most Length (
<=), and fails otherwise. - Apply coverage. At 100% coverage, any failing row causes the check to fail. Below 100% coverage, the check fails only when the passing fraction drops below the threshold (see Coverage and Tolerance).
What Counts as Length
The check counts characters exactly as stored: leading and trailing spaces, punctuation, and accented characters all count. A value that looks short on screen can still fail when it carries padding from a fixed-width source. Trim the value upstream (or with a Computed Field) when the padding is not meaningful.
NULL Handling
The check passes NULL values: a row with NULL in the evaluated field is not counted as a violation. Max Length only measures present values.
An empty string is a present value with length zero, so it is evaluated like any other. If the field must also be populated, pair the check with a Not Null check on the same field.
Array Fields
On an array field the two modes measure different things. With Array Element Context on, each element's own string length is measured and the row fails as soon as one element is longer than Length. With it off, Length caps the number of elements in the array (size(field) <= Length), not the characters in any of them, so a character limit set here silently becomes an element-count limit. NULL elements inside an array are skipped, and empty or NULL arrays pass.
Both array modes run as a field-level check, so they report a Shape Anomaly for the field instead of per-row Record Anomalies, whatever the coverage is set to.
The Filter Clause
The filter clause is a SQL WHERE expression applied before the evaluation. Filtered-out rows are ignored entirely (they cannot trigger a violation and are not counted in the totals).
When a filter is set, both the Record Anomaly and the Shape Anomaly messages end with [filter: <expression>] so the evaluated scope is visible in the alert.
Coverage and Tolerance
Coverage is a fractional value between 0 and 1 that defines the minimum fraction of evaluated rows that must pass:
1.0(100%, default): every row in the filtered set must pass. Any failing row causes the check to fail. This is the strictest setting.< 1.0: the check tolerates a fraction of rows failing. The check fails only when the fraction of passing rows drops below the threshold.
Lower coverage values are useful when a small, known fraction of failures is expected. Use coverage carefully: a 0.5% tolerance can mask a real regression that happens to fall just under the threshold.
See Also
-
Anomaly Reporting
The anomaly messages the check produces, what the numbers mean, Source Records highlighting, and Custom Anomaly Description.
-
Examples
Three production scenarios with sample data, anomaly messages, and the SQL equivalent of what the check evaluates.
-
Best Practices
Guidelines for choosing the length, pairing rules, and keeping the signal clean.
-
Permissions
The team permission each action needs: view, create, edit, archive, restore, and delete.