Skip to content

How Is Credit Card Checks Work

Definition

Asserts that every value in a text field is a credit card number.

Overview

The Is Credit Card rule validates the value of a single text field as a card number. The whole value must be the number: unlike Contains Credit Card, which passes as soon as a card number appears somewhere inside the text, this rule rejects a value that carries extra content around it. Common separators inside the number itself, such as spaces and dashes, are accepted.

Typical use cases:

  • Validate a column that is meant to hold nothing but a card number.
  • Catch tokens, placeholders, or notes written into a card column by a broken integration.
  • Confirm a payment file conforms before it is submitted to an acquirer.

Field Scope

Single: The rule evaluates exactly one field per check.

Accepted Types

Type Supported
String
Array

Array fields are accepted by the form and the API, but the rule has no element-level evaluation; see Array Fields.

General Properties

Name Supported
Filter
Allows the targeting of specific data based on conditions
Coverage Customization
Allows adjusting the percentage of records that must meet the rule's conditions

The filter allows you to define a subset of data upon which the rule will operate.

It requires a valid Spark SQL expression that determines the criteria rows in the DataFrame should meet. This means the expression specifies which rows the DataFrame should include based on those criteria. Since it's applied directly to the Spark DataFrame, traditional SQL constructs like WHERE clauses are not supported.

Examples

Direct Conditions

Simply specify the condition you want to be met.

Correct usage" collapsible="true
O_TOTALPRICE > 1000
C_MKTSEGMENT = 'BUILDING'
Incorrect usage" collapsible="true
WHERE O_TOTALPRICE > 1000
WHERE C_MKTSEGMENT = 'BUILDING'

Combining Conditions

Combine multiple conditions using logical operators like AND and OR.

Correct usage" collapsible="true
O_ORDERPRIORITY = '1-URGENT' AND O_ORDERSTATUS = 'O'
(L_SHIPDATE = '1998-09-02' OR L_RECEIPTDATE = '1998-09-01') AND L_RETURNFLAG = 'R'
Incorrect usage" collapsible="true
WHERE O_ORDERPRIORITY = '1-URGENT' AND O_ORDERSTATUS = 'O'
O_TOTALPRICE > 1000, O_ORDERSTATUS = 'O'

Utilizing Functions

Leverage Spark SQL functions to refine and enhance your conditions.

Correct usage" collapsible="true
RIGHT(
    O_ORDERPRIORITY,
    LENGTH(O_ORDERPRIORITY) - INSTR('-', O_ORDERPRIORITY)
) = 'URGENT'
LEVENSHTEIN(C_NAME, 'Supplier#000000001') < 7
Incorrect usage" collapsible="true
RIGHT(
    O_ORDERPRIORITY,
    LENGTH(O_ORDERPRIORITY) - CHARINDEX('-', O_ORDERPRIORITY)
) = 'URGENT'
EDITDISTANCE(C_NAME, 'Supplier#000000001') < 7

Using scan-time variables

To refer to the current dataframe being analyzed, use the reserved dynamic variable {{_qualytics_self}}.

Correct usage" collapsible="true
O_ORDERSTATUS IN (
    SELECT DISTINCT O_ORDERSTATUS
    FROM {{_qualytics_self}}
    WHERE O_TOTALPRICE > 1000
)
Incorrect usage" collapsible="true
O_ORDERSTATUS IN (
    SELECT DISTINCT O_ORDERSTATUS
    FROM ORDERS
    WHERE O_TOTALPRICE > 1000
)

While subqueries can be useful, their application within filters in our context has limitations. For example, directly referencing other containers or the broader target container in such subqueries is not supported. Attempting to do so will result in an error.

Important Note on {{_qualytics_self}}

The {{_qualytics_self}} keyword refers to the dataframe that's currently under examination. In the context of a full scan, this variable represents the entire target container. However, during incremental scans, it only reflects a subset of the target container, capturing just the incremental data. It's crucial to recognize that in such scenarios, using {{_qualytics_self}} may not encompass all entries from the target container.

Anomaly Types

Type Supported
Record
Flag inconsistencies at the row level
Shape
Flag inconsistencies in the overall patterns and distributions of a field

Evaluation Flow

Every Is Credit Card check follows the same four-step evaluation flow:

  1. Apply the filter clause. If the check has a filter set, only rows matching the filter expression continue to the next step. Rows outside the filter are ignored and cannot contribute to the violation count.
  2. Read the field value as text. The platform reads the row's value for the selected field.
  3. Validate it as a card number. Spaces and dashes are stripped from the value; what remains must be a non-empty string of digits that satisfies the Luhn checksum. Anything else fails, including NULL.
  4. Apply coverage. At 100% coverage, any failing row causes the check to fail. Below 100% coverage, the check fails only when the passing fraction drops below the threshold (see Coverage and Tolerance).

Format, Not Authorization

The check validates the arithmetic of the number, not its provenance. After spaces and dashes are removed, the value must be all digits and its Luhn checksum must come out to zero. Nothing else is asserted: there is no card-number length bound and no issuer-prefix test, so a short all-digit value that happens to satisfy Luhn (0, 18, 0000) passes. It also cannot tell whether the card was ever issued, is still active, or belongs to the person on the row. Treat it as a guard against corrupted or misplaced data, not as a substitute for the validation your payment provider performs. When the column must also have a plausible card length, pair the check with Min Length or Matches Pattern.

Masked and Tokenized Values Fail

A column storing **** **** **** 1111, a vault token, or a truncated last-four no longer contains a full number, so the check reports it. That is correct behavior, and it usually means the check is pointed at the wrong column: masked and tokenized columns are not meant to carry card data.

NULL Handling

The check reports NULL values. The assertion returns false for a NULL input, so a row with NULL in the evaluated field is a violation like any other, and it counts in the evaluated total. Its Record Anomaly renders the value as null.

An empty string fails for the same reason: there is no number to validate.

Because absence is already a violation, there is no need to pair the check with Not Null. If NULL is legitimate in the column and should not be reported, exclude those rows with a filter clause such as card_number IS NOT NULL.

Array Fields

Is Credit Card has no element-level evaluation. The rule routes straight to the scalar card-number assertion on the selected column whatever the field type, so the elements of an array are never validated one by one. Do not rely on an Is Credit Card check to vet the contents of an array column.

To validate the elements of an array of text values, use Contains Credit Card, which does branch on array fields and requires every element to carry a card number.

The Filter Clause

The filter clause is a SQL WHERE expression applied before the evaluation. Filtered-out rows are ignored entirely (they cannot trigger a violation and are not counted in the totals).

When a filter is set, both the Record Anomaly and the Shape Anomaly messages end with [filter: <expression>] so the evaluated scope is visible in the alert.

Coverage and Tolerance

Coverage is a fractional value between 0 and 1 that defines the minimum fraction of evaluated rows that must pass:

  • 1.0 (100%, default): every row in the filtered set must pass. Any failing row causes the check to fail. This is the strictest setting.
  • < 1.0: the check tolerates a fraction of rows failing. The check fails only when the fraction of passing rows drops below the threshold.

Lower coverage values are useful when a small, known fraction of failures is expected. Use coverage carefully: a 0.5% tolerance can mask a real regression that happens to fall just under the threshold.

See Also