Skip to content

API reference

The public Python interface, generated from its docstrings.

Configuration models

Pydantic models for a comparison and its sources. Build them in Python, or read them from YAML with load_config. Baseline and AcceptedChange hold the drift a run accepts; see Accepting drift.

Configuration and result models.

The Pydantic models here define every configuration setting, in YAML or in Python, and the results a comparison returns.

ArtifactFormat = Literal['csv', 'parquet', 'json', 'ndjson', 'arrow'] module-attribute

File format for exported discrepancy artifacts.

A closed set that matches the engine's writers, so a typo fails when the config loads instead of after the comparison runs. Excel can be read but not written: writing a workbook needs another dependency.

BIGQUERY_PROJECT_PATTERN = '^[a-z][a-z0-9-]{4,28}[a-z0-9]$' module-attribute

Google Cloud project id: six to thirty lowercase letters, digits, or hyphens, starting with a letter and not ending with a hyphen. Older domain-scoped ids, such as example.com:project, are refused rather than guessed at.

BIGQUERY_TABLE_PATTERN = '^[A-Za-z_][A-Za-z0-9_]*(\\.[A-Za-z_][A-Za-z0-9_]*)?$' module-attribute

BigQuery table path: a table, or a dataset and table. The project is set separately, since project ids may hold hyphens, which an identifier may not.

CastTarget = Literal['Int64', 'Float64', 'String', 'Boolean', 'Date', 'Datetime'] module-attribute

Polars type a column may be cast to before comparison.

A closed set rather than a free-form name, for two reasons. A typo fails when the config loads instead of skipping the cast. And the SQL compiler renders this field into CAST(x AS <type>), where a type name cannot be quoted or bound, so only a closed set keeps configuration text out of the SQL. Every member maps to a fixed keyword in each dialect.

DATABASE_PUSHDOWN_SCHEMES = POSTGRES_SCHEMES module-attribute

URI schemes whose tables a database source can compare inside the database.

FindingSeverity = Literal['error', 'warning'] module-attribute

How sure a configuration check is: an error stops the run, and a warning stops it only if the tables hold what the setting cannot handle.

POSTGRES_SCHEMES = frozenset({'postgres', 'postgresql'}) module-attribute

The URI schemes that name a Postgres database.

SQL_IDENTIFIER_SEGMENT = re.compile(SQL_IDENTIFIER_SEGMENT_PATTERN) module-attribute

Compiled allowlist applied by the warehouse SQL compiler before quoting.

SQL_IDENTIFIER_SEGMENT_PATTERN = '^[A-Za-z_][A-Za-z0-9_]*$' module-attribute

Unquoted SQL identifier: letter or underscore, then alphanumeric or underscore.

SQL_RELATION_PATTERN = '^[A-Za-z_][A-Za-z0-9_]*(\\.[A-Za-z_][A-Za-z0-9_]*){0,2}$' module-attribute

Warehouse table path: one to three identifier segments joined by dots.

SchemaMode = Literal['exact', 'allow_additions', 'allow_removals', 'intersection'] module-attribute

How strictly the two datasets' columns must agree.

  • exact: both sides have the same columns, in any order.
  • allow_additions: the target may add columns, but keeps every source column.
  • allow_removals: the target may drop columns, but adds none.
  • intersection, the default: compare only the columns on both sides.

SentinelValue = str | int | float | bool module-attribute

One null_values entry, which keeps the type written in the config.

In YAML, -999 is an integer sentinel and "-999" a text one. Each applies only to columns of a matching type.

SourceRef = Annotated[SourceConfig | SnowflakeConfig | DatabricksConfig | BigQueryConfig | DeltaLakeConfig | IcebergConfig | DatabaseConfig | DuckDBConfig, Field(discriminator='type')] module-attribute

YAML/Python source or target: file, warehouse, lakehouse, database, or DuckDB.

SourceType = Literal['csv', 'parquet', 'json', 'ndjson', 'arrow', 'avro', 'excel'] module-attribute

File formats Veridelta can read.

Exactly the set LoaderFactory implements. Delta Lake is not here: it is a table format, read through the delta source type rather than as a file.

SuggestedSetting = float | bool | str | list[SentinelValue] module-attribute

The value of one rule setting a suggestion adds, as DiffRule takes it.

WhitespaceMode = Literal['none', 'left', 'right', 'both'] module-attribute

Which ends of a string to strip whitespace from.

  • none: strip nothing.
  • left: strip leading whitespace.
  • right: strip trailing whitespace.
  • both: strip both ends.

AcceptedChange

Bases: BaseModel

One changed row a baseline accepts, and the columns whose drift it accepts.

Attributes:

Name Type Description
key dict[str, Any]

The row's primary key, by key column.

columns list[str]

Compared columns whose drift on this row is accepted.

Source code in src/veridelta/models.py
class AcceptedChange(BaseModel):
    """One changed row a baseline accepts, and the columns whose drift it accepts.

    Attributes:
        key (dict[str, Any]): The row's primary key, by key column.
        columns (list[str]): Compared columns whose drift on this row is accepted.
    """

    model_config = ConfigDict(extra="forbid", frozen=True)

    key: dict[str, Any] = Field(..., description="The row's primary key, by key column.")
    columns: list[str] = Field(
        ..., min_length=1, description="Compared columns whose drift on this row is accepted."
    )

Baseline

Bases: BaseModel

Drift a run accepts: rows by kind and primary key, and changed columns by row.

veridelta run --baseline FILE reads one and leaves what it lists out of the counts and the verdict, so a run fails only on drift the file does not list. A changed row is accepted only in the columns its entry names: drift in any other column still counts.

Attributes:

Name Type Description
primary_keys list[str]

The key columns every entry names, as the configuration's primary_keys does.

added list[dict[str, Any]]

Keys of accepted rows only in the target.

removed list[dict[str, Any]]

Keys of accepted rows only in the source.

changed list[AcceptedChange]

Accepted changed rows, with their columns.

Source code in src/veridelta/models.py
class Baseline(BaseModel):
    """Drift a run accepts: rows by kind and primary key, and changed columns by row.

    `veridelta run --baseline FILE` reads one and leaves what it lists out of the
    counts and the verdict, so a run fails only on drift the file does not list.
    A changed row is accepted only in the columns its entry names: drift in any
    other column still counts.

    Attributes:
        primary_keys (list[str]): The key columns every entry names, as the
            configuration's `primary_keys` does.
        added (list[dict[str, Any]]): Keys of accepted rows only in the target.
        removed (list[dict[str, Any]]): Keys of accepted rows only in the source.
        changed (list[AcceptedChange]): Accepted changed rows, with their columns.
    """

    model_config = ConfigDict(extra="forbid", frozen=True)

    primary_keys: list[str] = Field(..., min_length=1, description="The key columns.")
    added: list[dict[str, Any]] = Field(
        default_factory=list[dict[str, Any]],
        description="Keys of accepted rows only in the target.",
    )
    removed: list[dict[str, Any]] = Field(
        default_factory=list[dict[str, Any]],
        description="Keys of accepted rows only in the source.",
    )
    changed: list[AcceptedChange] = Field(
        default_factory=list[AcceptedChange],
        description="Accepted changed rows, with their columns.",
    )

    @model_validator(mode="after")
    def _entries_name_the_keys(self) -> "Baseline":
        """Hold every entry to exactly the key columns the baseline names."""
        expected = set(self.primary_keys)
        for key in [*self.added, *self.removed, *(entry.key for entry in self.changed)]:
            if set(key) != expected:
                raise ValueError(
                    f"every entry names the primary keys {sorted(expected)}, "
                    f"but one names {sorted(key)}"
                )
        return self

    @classmethod
    def of_rows(
        cls,
        primary_keys: Sequence[str],
        frames: tuple[pl.DataFrame, pl.DataFrame, pl.DataFrame],
        compared_columns: Sequence[str],
    ) -> "Baseline":
        """Accept the drift in a local run's frames, in key order.

        Args:
            primary_keys (Sequence[str]): The run's key columns.
            frames (tuple[pl.DataFrame, pl.DataFrame, pl.DataFrame]): The added,
                removed, and changed rows, as `DiffResult` holds them.
            compared_columns (Sequence[str]): The columns the run compared.

        Returns:
            Baseline: Added and removed rows by key, and changed rows with the
                columns that differ on each.
        """
        names = list(primary_keys)
        added, removed, changed = frames
        entries: list[AcceptedChange] = []
        if not changed.is_empty():
            differing = pl.concat_list(
                pl.when(pl.col(f"{column}_is_match").not_()).then(pl.lit(column))
                for column in compared_columns
            ).list.drop_nulls()
            rows = changed.sort(names).select(*names, differing.alias(_DIFFERING))
            # A changed row differs in at least one column, so no entry is empty.
            entries = [
                AcceptedChange(key={name: row[name] for name in names}, columns=row[_DIFFERING])
                for row in rows.iter_rows(named=True)
            ]
        return cls(
            primary_keys=names,
            added=added.select(names).sort(names).to_dicts(),
            removed=removed.select(names).sort(names).to_dicts(),
            changed=entries,
        )

    @classmethod
    def of(cls, result: "DiffResult") -> "Baseline":
        """Accept every row of drift a run found, with what its baseline accepted.

        Saved and read back with `run --baseline`, it makes the same run match.
        An entry of the run's own baseline that matched nothing, such as a row
        fixed since, is left out.

        Args:
            result (DiffResult): A local run's result.

        Returns:
            Baseline: The run's drift, by kind and key.

        Raises:
            ConfigError: If the run compared its data where it is stored, which
                returns no differing columns to accept.
        """
        if result.keys_only:
            raise ConfigError(
                "A baseline is saved from a run that compares both sides locally, and this "
                "pair was compared where it is stored. Save one from files exported from it."
            )
        found = cls.of_rows(
            result.primary_keys,
            (result.added, result.removed, result.changed),
            result.compared_columns,
        )
        return found if result.accepted is None else found.merged(result.accepted)

    def _key(self, key: dict[str, Any]) -> tuple[Any, ...]:
        """Turn an entry's key into a tuple in key order, to compare and sort entries."""
        return tuple(key[name] for name in self.primary_keys)

    def _order(self, key: dict[str, Any]) -> tuple[tuple[bool, Any], ...]:
        """Sort keys in key order, with NULL last, as Polars sorts them."""
        return tuple((value is None, "" if value is None else value) for value in self._key(key))

    def without(self, other: "Baseline") -> "Baseline":
        """Return the entries here that `other` does not hold.

        Args:
            other (Baseline): Entries to take away, by key, and by column on a
                changed row.

        Returns:
            Baseline: What is left.
        """
        theirs = {self._key(entry.key): set(entry.columns) for entry in other.changed}
        changed: list[AcceptedChange] = []
        for entry in self.changed:
            left = [c for c in entry.columns if c not in theirs.get(self._key(entry.key), set())]
            if left:
                changed.append(AcceptedChange(key=entry.key, columns=left))
        return Baseline(
            primary_keys=self.primary_keys,
            added=_keys_without(self.added, other.added, self._key),
            removed=_keys_without(self.removed, other.removed, self._key),
            changed=changed,
        )

    def merged(self, other: "Baseline") -> "Baseline":
        """Return the entries of both, with each changed row's columns joined, in key order.

        Args:
            other (Baseline): Entries to add.

        Returns:
            Baseline: Both, without repeats.
        """
        columns: dict[tuple[Any, ...], tuple[dict[str, Any], list[str]]] = {}
        for entry in [*self.changed, *other.changed]:
            _, held = columns.setdefault(self._key(entry.key), (entry.key, []))
            held.extend(column for column in entry.columns if column not in held)
        return Baseline(
            primary_keys=self.primary_keys,
            added=sorted(
                [*self.added, *_keys_without(other.added, self.added, self._key)], key=self._order
            ),
            removed=sorted(
                [*self.removed, *_keys_without(other.removed, self.removed, self._key)],
                key=self._order,
            ),
            changed=[
                AcceptedChange(key=key, columns=held)
                for key, held in sorted(columns.values(), key=lambda pair: self._order(pair[0]))
            ],
        )

    def write(self, path: str) -> Path:
        """Write the baseline as JSON, for `run --baseline` to read.

        Args:
            path (str): The file, such as `accepted.json`.

        Returns:
            Path: The file written.

        Raises:
            ConfigError: If the file cannot be written.
        """
        target = Path(path)
        try:
            target.write_text(self.model_dump_json(indent=2) + "\n", encoding="utf-8")
        except OSError as exc:
            raise ConfigError(f"Cannot write the baseline {path}: {exc.strerror}.") from exc
        return target

    @classmethod
    def read(cls, path: str) -> "Baseline":
        """Read a baseline from a JSON file.

        Args:
            path (str): The file, such as `accepted.json`.

        Returns:
            Baseline: The drift it accepts.

        Raises:
            ConfigError: If the file cannot be read or is not a valid baseline.
        """
        try:
            text = Path(path).read_text(encoding="utf-8")
        except OSError as exc:
            raise ConfigError(f"Cannot read the baseline {path}: {exc.strerror}.") from exc
        try:
            return cls.model_validate_json(text)
        except ValueError as exc:
            raise ConfigError(f"The baseline {path} is not valid: {exc}") from exc

merged(other)

Return the entries of both, with each changed row's columns joined, in key order.

Parameters:

Name Type Description Default
other Baseline

Entries to add.

required

Returns:

Name Type Description
Baseline Baseline

Both, without repeats.

Source code in src/veridelta/models.py
def merged(self, other: "Baseline") -> "Baseline":
    """Return the entries of both, with each changed row's columns joined, in key order.

    Args:
        other (Baseline): Entries to add.

    Returns:
        Baseline: Both, without repeats.
    """
    columns: dict[tuple[Any, ...], tuple[dict[str, Any], list[str]]] = {}
    for entry in [*self.changed, *other.changed]:
        _, held = columns.setdefault(self._key(entry.key), (entry.key, []))
        held.extend(column for column in entry.columns if column not in held)
    return Baseline(
        primary_keys=self.primary_keys,
        added=sorted(
            [*self.added, *_keys_without(other.added, self.added, self._key)], key=self._order
        ),
        removed=sorted(
            [*self.removed, *_keys_without(other.removed, self.removed, self._key)],
            key=self._order,
        ),
        changed=[
            AcceptedChange(key=key, columns=held)
            for key, held in sorted(columns.values(), key=lambda pair: self._order(pair[0]))
        ],
    )

of(result) classmethod

Accept every row of drift a run found, with what its baseline accepted.

Saved and read back with run --baseline, it makes the same run match. An entry of the run's own baseline that matched nothing, such as a row fixed since, is left out.

Parameters:

Name Type Description Default
result DiffResult

A local run's result.

required

Returns:

Name Type Description
Baseline Baseline

The run's drift, by kind and key.

Raises:

Type Description
ConfigError

If the run compared its data where it is stored, which returns no differing columns to accept.

Source code in src/veridelta/models.py
@classmethod
def of(cls, result: "DiffResult") -> "Baseline":
    """Accept every row of drift a run found, with what its baseline accepted.

    Saved and read back with `run --baseline`, it makes the same run match.
    An entry of the run's own baseline that matched nothing, such as a row
    fixed since, is left out.

    Args:
        result (DiffResult): A local run's result.

    Returns:
        Baseline: The run's drift, by kind and key.

    Raises:
        ConfigError: If the run compared its data where it is stored, which
            returns no differing columns to accept.
    """
    if result.keys_only:
        raise ConfigError(
            "A baseline is saved from a run that compares both sides locally, and this "
            "pair was compared where it is stored. Save one from files exported from it."
        )
    found = cls.of_rows(
        result.primary_keys,
        (result.added, result.removed, result.changed),
        result.compared_columns,
    )
    return found if result.accepted is None else found.merged(result.accepted)

of_rows(primary_keys, frames, compared_columns) classmethod

Accept the drift in a local run's frames, in key order.

Parameters:

Name Type Description Default
primary_keys Sequence[str]

The run's key columns.

required
frames tuple[DataFrame, DataFrame, DataFrame]

The added, removed, and changed rows, as DiffResult holds them.

required
compared_columns Sequence[str]

The columns the run compared.

required

Returns:

Name Type Description
Baseline Baseline

Added and removed rows by key, and changed rows with the columns that differ on each.

Source code in src/veridelta/models.py
@classmethod
def of_rows(
    cls,
    primary_keys: Sequence[str],
    frames: tuple[pl.DataFrame, pl.DataFrame, pl.DataFrame],
    compared_columns: Sequence[str],
) -> "Baseline":
    """Accept the drift in a local run's frames, in key order.

    Args:
        primary_keys (Sequence[str]): The run's key columns.
        frames (tuple[pl.DataFrame, pl.DataFrame, pl.DataFrame]): The added,
            removed, and changed rows, as `DiffResult` holds them.
        compared_columns (Sequence[str]): The columns the run compared.

    Returns:
        Baseline: Added and removed rows by key, and changed rows with the
            columns that differ on each.
    """
    names = list(primary_keys)
    added, removed, changed = frames
    entries: list[AcceptedChange] = []
    if not changed.is_empty():
        differing = pl.concat_list(
            pl.when(pl.col(f"{column}_is_match").not_()).then(pl.lit(column))
            for column in compared_columns
        ).list.drop_nulls()
        rows = changed.sort(names).select(*names, differing.alias(_DIFFERING))
        # A changed row differs in at least one column, so no entry is empty.
        entries = [
            AcceptedChange(key={name: row[name] for name in names}, columns=row[_DIFFERING])
            for row in rows.iter_rows(named=True)
        ]
    return cls(
        primary_keys=names,
        added=added.select(names).sort(names).to_dicts(),
        removed=removed.select(names).sort(names).to_dicts(),
        changed=entries,
    )

read(path) classmethod

Read a baseline from a JSON file.

Parameters:

Name Type Description Default
path str

The file, such as accepted.json.

required

Returns:

Name Type Description
Baseline Baseline

The drift it accepts.

Raises:

Type Description
ConfigError

If the file cannot be read or is not a valid baseline.

Source code in src/veridelta/models.py
@classmethod
def read(cls, path: str) -> "Baseline":
    """Read a baseline from a JSON file.

    Args:
        path (str): The file, such as `accepted.json`.

    Returns:
        Baseline: The drift it accepts.

    Raises:
        ConfigError: If the file cannot be read or is not a valid baseline.
    """
    try:
        text = Path(path).read_text(encoding="utf-8")
    except OSError as exc:
        raise ConfigError(f"Cannot read the baseline {path}: {exc.strerror}.") from exc
    try:
        return cls.model_validate_json(text)
    except ValueError as exc:
        raise ConfigError(f"The baseline {path} is not valid: {exc}") from exc

without(other)

Return the entries here that other does not hold.

Parameters:

Name Type Description Default
other Baseline

Entries to take away, by key, and by column on a changed row.

required

Returns:

Name Type Description
Baseline Baseline

What is left.

Source code in src/veridelta/models.py
def without(self, other: "Baseline") -> "Baseline":
    """Return the entries here that `other` does not hold.

    Args:
        other (Baseline): Entries to take away, by key, and by column on a
            changed row.

    Returns:
        Baseline: What is left.
    """
    theirs = {self._key(entry.key): set(entry.columns) for entry in other.changed}
    changed: list[AcceptedChange] = []
    for entry in self.changed:
        left = [c for c in entry.columns if c not in theirs.get(self._key(entry.key), set())]
        if left:
            changed.append(AcceptedChange(key=entry.key, columns=left))
    return Baseline(
        primary_keys=self.primary_keys,
        added=_keys_without(self.added, other.added, self._key),
        removed=_keys_without(self.removed, other.removed, self._key),
        changed=changed,
    )

write(path)

Write the baseline as JSON, for run --baseline to read.

Parameters:

Name Type Description Default
path str

The file, such as accepted.json.

required

Returns:

Name Type Description
Path Path

The file written.

Raises:

Type Description
ConfigError

If the file cannot be written.

Source code in src/veridelta/models.py
def write(self, path: str) -> Path:
    """Write the baseline as JSON, for `run --baseline` to read.

    Args:
        path (str): The file, such as `accepted.json`.

    Returns:
        Path: The file written.

    Raises:
        ConfigError: If the file cannot be written.
    """
    target = Path(path)
    try:
        target.write_text(self.model_dump_json(indent=2) + "\n", encoding="utf-8")
    except OSError as exc:
        raise ConfigError(f"Cannot write the baseline {path}: {exc.strerror}.") from exc
    return target

BigQueryConfig

Bases: BaseModel

Immutable connection settings for BigQuery warehouse pushdown.

Attributes:

Name Type Description
type Literal['bigquery']

Source kind, which selects this model.

table str

Table to compare, as dataset.table, or table when dataset names the default dataset.

project str

Google Cloud project that runs the queries and holds the data. It never appears in SQL.

dataset str | None

Default dataset for a table named without one.

location str | None

Location the queries run in, such as US or europe-west2. BigQuery infers it when unset.

credentials_path str | None

Service account key file. Application Default Credentials are used when unset. Left out when the config is printed, but kept by model_dump(), which the connector needs.

maximum_bytes_billed int | None

Fail any statement that would bill more bytes than this, instead of running it.

Source code in src/veridelta/models.py
class BigQueryConfig(BaseModel):
    """Immutable connection settings for BigQuery warehouse pushdown.

    Attributes:
        type (Literal["bigquery"]): Source kind, which selects this model.
        table (str): Table to compare, as `dataset.table`, or `table` when
            `dataset` names the default dataset.
        project (str): Google Cloud project that runs the queries and holds
            the data. It never appears in SQL.
        dataset (str | None): Default dataset for a table named without one.
        location (str | None): Location the queries run in, such as `US` or
            `europe-west2`. BigQuery infers it when unset.
        credentials_path (str | None): Service account key file. Application
            Default Credentials are used when unset. Left out when the config
            is printed, but kept by `model_dump()`, which the connector needs.
        maximum_bytes_billed (int | None): Fail any statement that would bill
            more bytes than this, instead of running it.
    """

    model_config = ConfigDict(extra="forbid", frozen=True, hide_input_in_errors=True)

    type: Literal["bigquery"] = Field("bigquery", description="Discriminator for BigQuery.")
    table: str = Field(
        ...,
        pattern=BIGQUERY_TABLE_PATTERN,
        description="Table to compare, as dataset.table, or table with a default dataset.",
    )
    project: str = Field(
        ...,
        pattern=BIGQUERY_PROJECT_PATTERN,
        description="Google Cloud project that runs the queries and holds the data.",
    )
    dataset: str | None = Field(
        default=None,
        pattern=SQL_IDENTIFIER_SEGMENT_PATTERN,
        description="Default dataset for a table named without one.",
    )
    location: str | None = Field(
        default=None, description="Location the queries run in, such as US or europe-west2."
    )
    credentials_path: str | None = Field(
        default=None,
        repr=False,
        description="Service account key file; Application Default Credentials when unset.",
    )
    maximum_bytes_billed: int | None = Field(
        default=None,
        ge=1,
        strict=True,
        description="Fail any statement that would bill more bytes than this.",
    )

    @model_validator(mode="after")
    def validate_dataset(self) -> "BigQueryConfig":
        """Require a default dataset when the table is named without one.

        Returns:
            BigQueryConfig: The validated configuration.

        Raises:
            ValueError: If `table` has one segment and `dataset` is unset.
        """
        if "." not in self.table and self.dataset is None:
            raise ValueError(
                "A 'table' without a dataset needs 'dataset'; write it as dataset.table, "
                "or name the default dataset."
            )
        return self

validate_dataset()

Require a default dataset when the table is named without one.

Returns:

Name Type Description
BigQueryConfig BigQueryConfig

The validated configuration.

Raises:

Type Description
ValueError

If table has one segment and dataset is unset.

Source code in src/veridelta/models.py
@model_validator(mode="after")
def validate_dataset(self) -> "BigQueryConfig":
    """Require a default dataset when the table is named without one.

    Returns:
        BigQueryConfig: The validated configuration.

    Raises:
        ValueError: If `table` has one segment and `dataset` is unset.
    """
    if "." not in self.table and self.dataset is None:
        raise ValueError(
            "A 'table' without a dataset needs 'dataset'; write it as dataset.table, "
            "or name the default dataset."
        )
    return self

ConfigFinding

Bases: BaseModel

One problem found by checking a configuration without running it.

Attributes:

Name Type Description
severity FindingSeverity

error when the run would fail, warning when it depends on stored column names or types.

message str

What is wrong and what to do about it.

Source code in src/veridelta/models.py
class ConfigFinding(BaseModel):
    """One problem found by checking a configuration without running it.

    Attributes:
        severity (FindingSeverity): `error` when the run would fail, `warning`
            when it depends on stored column names or types.
        message (str): What is wrong and what to do about it.
    """

    model_config = ConfigDict(extra="forbid", frozen=True)

    severity: FindingSeverity = Field(..., description="error or warning.")
    message: str = Field(..., description="What is wrong and what to do about it.")

DatabaseConfig

Bases: BaseModel

Immutable settings for reading a database table or query into a local comparison.

The rows are read through ConnectorX into Polars and compared by the local engine, so a database pairs with files, lakehouse tables, or another database. Requires the database extra (uv add 'veridelta[database]'). Two Postgres tables on one connection can instead be compared inside the database, without reading their rows, when both sides set pushdown.

Attributes:

Name Type Description
type Literal['database']

Source kind, which selects this model.

uri str

ConnectorX connection URI, such as postgresql://analyst@db.internal:5432/sales or sqlite:///srv/data/legacy.db. A password written inside it is masked when the config is printed, but kept by model_dump(), which the connector needs.

password str | None

Optional password, percent-encoded into the URI's user information when the connector reads. Left out when the config is printed, but kept by model_dump().

table str | None

Table or view to read whole, as one to three unquoted identifier segments.

query str | None

SQL statement to run instead, sent to the database exactly as written. Set exactly one of table and query.

pushdown bool

Whether to compare inside the database instead of reading the rows. Postgres table sources only, and both sides must set it. Defaults to False.

partition_on str | None

Integer column to split a table read on, so ConnectorX reads its ranges over several connections at once. Set it with partitions. The column may hold no NULL, since a range read leaves those rows out, so the read checks it first.

partitions int | None

How many ranges to split the read into: 2 or more. Set it with partition_on.

Source code in src/veridelta/models.py
class DatabaseConfig(BaseModel):
    """Immutable settings for reading a database table or query into a local comparison.

    The rows are read through ConnectorX into Polars and compared by the local
    engine, so a database pairs with files, lakehouse tables, or another
    database. Requires the `database` extra (`uv add 'veridelta[database]'`).
    Two Postgres tables on one connection can instead be compared inside the
    database, without reading their rows, when both sides set `pushdown`.

    Attributes:
        type (Literal["database"]): Source kind, which selects this model.
        uri (str): ConnectorX connection URI, such as
            `postgresql://analyst@db.internal:5432/sales` or
            `sqlite:///srv/data/legacy.db`. A password written inside it is
            masked when the config is printed, but kept by `model_dump()`,
            which the connector needs.
        password (str | None): Optional password, percent-encoded into the
            URI's user information when the connector reads. Left out when the
            config is printed, but kept by `model_dump()`.
        table (str | None): Table or view to read whole, as one to three
            unquoted identifier segments.
        query (str | None): SQL statement to run instead, sent to the database
            exactly as written. Set exactly one of `table` and `query`.
        pushdown (bool): Whether to compare inside the database instead of
            reading the rows. Postgres `table` sources only, and both sides must
            set it. Defaults to False.
        partition_on (str | None): Integer column to split a `table` read on,
            so ConnectorX reads its ranges over several connections at once.
            Set it with `partitions`. The column may hold no NULL, since a
            range read leaves those rows out, so the read checks it first.
        partitions (int | None): How many ranges to split the read into: 2
            or more. Set it with `partition_on`.
    """

    model_config = ConfigDict(extra="forbid", frozen=True, hide_input_in_errors=True)

    type: Literal["database"] = Field("database", description="Discriminator for databases.")
    uri: str = Field(
        ..., description="ConnectorX connection URI, such as postgresql://user@host/db."
    )
    password: str | None = Field(
        default=None,
        repr=False,
        description="Optional password, percent-encoded into the URI when the source is read.",
    )
    table: str | None = Field(
        default=None,
        pattern=SQL_RELATION_PATTERN,
        description="Table or view to read whole.",
    )
    query: str | None = Field(
        default=None,
        min_length=1,
        description="SQL statement to run instead of reading a table, sent as written.",
    )
    pushdown: bool = Field(
        default=False,
        strict=True,
        description=(
            "Compare inside the database instead of reading the rows. Postgres tables "
            "only; set it on both sides."
        ),
    )
    partition_on: str | None = Field(
        default=None,
        pattern=SQL_IDENTIFIER_SEGMENT_PATTERN,
        description=(
            "Integer column to split a table read on, so ConnectorX reads its ranges in "
            "parallel. Set it with partitions."
        ),
    )
    partitions: int | None = Field(
        default=None,
        ge=2,
        strict=True,
        description="How many ranges to split the read into, 2 or more. Set it with partition_on.",
    )

    @model_validator(mode="after")
    def validate_connection(self) -> "DatabaseConfig":
        """Reject a source that is ambiguous about its rows or its password.

        Returns:
            DatabaseConfig: The validated instance.

        Raises:
            ValueError: If both or neither of `table` and `query` are set, the
                URI has no scheme, `password` conflicts with the URI, or
                `pushdown` is set on anything but a Postgres table.
        """
        if (self.table is None) == (self.query is None):
            raise ValueError(
                "A database source reads a 'table' or runs a 'query'; set exactly one."
            )
        parts = urlsplit(self.uri)
        if not parts.scheme:
            raise ValueError("'uri' needs a scheme such as postgresql:// or sqlite://.")
        if self.pushdown and parts.scheme.lower() not in DATABASE_PUSHDOWN_SCHEMES:
            raise ValueError(
                "'pushdown' compares inside the database and works on postgresql:// "
                "connections only."
            )
        if self.pushdown and self.table is None:
            raise ValueError("'pushdown' compares two tables, so set 'table' rather than 'query'.")
        if self.password is not None:
            if parts.password is not None:
                raise ValueError("Set the password in 'password' or inside 'uri', not both.")
            if not parts.username:
                raise ValueError(
                    "'password' needs a user name in 'uri', as in "
                    "postgresql://analyst@db.internal/sales."
                )
        return self

    @model_validator(mode="after")
    def validate_partitions(self) -> "DatabaseConfig":
        """Reject a split read that names half its settings or has nothing to split.

        Returns:
            DatabaseConfig: The validated instance.

        Raises:
            ValueError: If only one of `partition_on` and `partitions` is set,
                or they are set on a `query` or with `pushdown`.
        """
        if (self.partition_on is None) != (self.partitions is None):
            raise ValueError("Set 'partition_on' and 'partitions' together.")
        if self.partition_on is not None and self.table is None:
            raise ValueError(
                "A partitioned read can only split a 'table'; a 'query' runs as written."
            )
        if self.partition_on is not None and self.pushdown:
            raise ValueError("'pushdown' reads no rows, so there is no read to partition.")
        return self

    @property
    def redacted_uri(self) -> str:
        """Return the URI with any password in it replaced by `***`, and no query.

        The query goes whole, since a parameter such as `?password=` can carry
        a credential too.

        Returns:
            str: The URI, safe to print or log.
        """
        parts = urlsplit(self.uri)
        if parts.password is None and not parts.query:
            return self.uri
        netloc = parts.netloc
        if parts.password is not None:
            userinfo, _, hostinfo = netloc.rpartition("@")
            netloc = f"{userinfo.partition(':')[0]}:***@{hostinfo}"
        # Written out, since `urlunsplit` drops the `//` of a URI with no host, such as SQLite's.
        return f"{parts.scheme}://{netloc}{parts.path}"

    def __repr_args__(self) -> Iterable[tuple[str | None, Any]]:
        """Print the URI with its password masked.

        `repr()`, `str()`, and rich displays all read from here, while
        `model_dump()` and the connector get the URI as written.

        Yields:
            tuple[str | None, Any]: Each field name with the value to print.
        """
        for name, value in super().__repr_args__():
            yield (name, self.redacted_uri) if name == "uri" else (name, value)

redacted_uri property

Return the URI with any password in it replaced by ***, and no query.

The query goes whole, since a parameter such as ?password= can carry a credential too.

Returns:

Name Type Description
str str

The URI, safe to print or log.

__repr_args__()

Print the URI with its password masked.

repr(), str(), and rich displays all read from here, while model_dump() and the connector get the URI as written.

Yields:

Type Description
Iterable[tuple[str | None, Any]]

tuple[str | None, Any]: Each field name with the value to print.

Source code in src/veridelta/models.py
def __repr_args__(self) -> Iterable[tuple[str | None, Any]]:
    """Print the URI with its password masked.

    `repr()`, `str()`, and rich displays all read from here, while
    `model_dump()` and the connector get the URI as written.

    Yields:
        tuple[str | None, Any]: Each field name with the value to print.
    """
    for name, value in super().__repr_args__():
        yield (name, self.redacted_uri) if name == "uri" else (name, value)

validate_connection()

Reject a source that is ambiguous about its rows or its password.

Returns:

Name Type Description
DatabaseConfig DatabaseConfig

The validated instance.

Raises:

Type Description
ValueError

If both or neither of table and query are set, the URI has no scheme, password conflicts with the URI, or pushdown is set on anything but a Postgres table.

Source code in src/veridelta/models.py
@model_validator(mode="after")
def validate_connection(self) -> "DatabaseConfig":
    """Reject a source that is ambiguous about its rows or its password.

    Returns:
        DatabaseConfig: The validated instance.

    Raises:
        ValueError: If both or neither of `table` and `query` are set, the
            URI has no scheme, `password` conflicts with the URI, or
            `pushdown` is set on anything but a Postgres table.
    """
    if (self.table is None) == (self.query is None):
        raise ValueError(
            "A database source reads a 'table' or runs a 'query'; set exactly one."
        )
    parts = urlsplit(self.uri)
    if not parts.scheme:
        raise ValueError("'uri' needs a scheme such as postgresql:// or sqlite://.")
    if self.pushdown and parts.scheme.lower() not in DATABASE_PUSHDOWN_SCHEMES:
        raise ValueError(
            "'pushdown' compares inside the database and works on postgresql:// "
            "connections only."
        )
    if self.pushdown and self.table is None:
        raise ValueError("'pushdown' compares two tables, so set 'table' rather than 'query'.")
    if self.password is not None:
        if parts.password is not None:
            raise ValueError("Set the password in 'password' or inside 'uri', not both.")
        if not parts.username:
            raise ValueError(
                "'password' needs a user name in 'uri', as in "
                "postgresql://analyst@db.internal/sales."
            )
    return self

validate_partitions()

Reject a split read that names half its settings or has nothing to split.

Returns:

Name Type Description
DatabaseConfig DatabaseConfig

The validated instance.

Raises:

Type Description
ValueError

If only one of partition_on and partitions is set, or they are set on a query or with pushdown.

Source code in src/veridelta/models.py
@model_validator(mode="after")
def validate_partitions(self) -> "DatabaseConfig":
    """Reject a split read that names half its settings or has nothing to split.

    Returns:
        DatabaseConfig: The validated instance.

    Raises:
        ValueError: If only one of `partition_on` and `partitions` is set,
            or they are set on a `query` or with `pushdown`.
    """
    if (self.partition_on is None) != (self.partitions is None):
        raise ValueError("Set 'partition_on' and 'partitions' together.")
    if self.partition_on is not None and self.table is None:
        raise ValueError(
            "A partitioned read can only split a 'table'; a 'query' runs as written."
        )
    if self.partition_on is not None and self.pushdown:
        raise ValueError("'pushdown' reads no rows, so there is no read to partition.")
    return self

DatabricksConfig

Bases: BaseModel

Immutable connection settings for Databricks SQL warehouse pushdown.

Attributes:

Name Type Description
type Literal['databricks']

Source kind, which selects this model.

table str

Fully qualified table or view to compare.

server_hostname str

Workspace hostname for the SQL warehouse.

http_path str

HTTP path of the SQL warehouse or cluster.

access_token str | None

Optional personal access token. Left out when the config is printed, but kept by model_dump(), which the connector needs.

catalog str | None

Optional Unity Catalog name.

schema_name str | None

Optional default schema name.

Source code in src/veridelta/models.py
class DatabricksConfig(BaseModel):
    """Immutable connection settings for Databricks SQL warehouse pushdown.

    Attributes:
        type (Literal["databricks"]): Source kind, which selects this model.
        table (str): Fully qualified table or view to compare.
        server_hostname (str): Workspace hostname for the SQL warehouse.
        http_path (str): HTTP path of the SQL warehouse or cluster.
        access_token (str | None): Optional personal access token. Left out
            when the config is printed, but kept by `model_dump()`, which the
            connector needs.
        catalog (str | None): Optional Unity Catalog name.
        schema_name (str | None): Optional default schema name.
    """

    model_config = ConfigDict(extra="forbid", frozen=True, hide_input_in_errors=True)

    type: Literal["databricks"] = Field("databricks", description="Discriminator for Databricks.")
    table: str = Field(
        ...,
        pattern=SQL_RELATION_PATTERN,
        description="Fully qualified table or view to compare.",
    )
    server_hostname: str = Field(..., description="Workspace hostname for the SQL warehouse.")
    http_path: str = Field(..., description="HTTP path of the SQL warehouse or cluster.")
    access_token: str | None = Field(
        default=None, repr=False, description="Optional personal access token."
    )
    catalog: str | None = Field(default=None, description="Optional Unity Catalog name.")
    schema_name: str | None = Field(default=None, description="Optional default schema name.")

DeltaLakeConfig

Bases: BaseModel

Immutable settings for a Delta Lake table scan.

Attributes:

Name Type Description
type Literal['delta']

Source kind, which selects this model.

table_uri str

Filesystem path or object-store URI of the table.

version int | None

Optional table version to time-travel.

storage_options dict[str, str]

Object-store credentials and options. Left out when the config is printed, but kept by model_dump(), which the scanner needs.

Source code in src/veridelta/models.py
class DeltaLakeConfig(BaseModel):
    """Immutable settings for a Delta Lake table scan.

    Attributes:
        type (Literal["delta"]): Source kind, which selects this model.
        table_uri (str): Filesystem path or object-store URI of the table.
        version (int | None): Optional table version to time-travel.
        storage_options (dict[str, str]): Object-store credentials and options.
            Left out when the config is printed, but kept by `model_dump()`,
            which the scanner needs.
    """

    model_config = ConfigDict(extra="forbid", frozen=True, hide_input_in_errors=True)

    type: Literal["delta"] = Field("delta", description="Discriminator for Delta Lake.")
    table_uri: str = Field(..., description="Filesystem path or object-store URI of the table.")
    version: int | None = Field(
        default=None,
        ge=0,
        strict=True,
        description="Optional table version to time-travel.",
    )
    storage_options: dict[str, str] = Field(
        default_factory=dict,
        repr=False,
        description="Object-store credentials and options passed to Polars.",
    )

DiffConfig

Bases: BaseModel

Settings and rules for one comparison.

Attributes:

Name Type Description
primary_keys list[str]

Columns that join the datasets. At least one is required, and together they must be unique in each dataset.

schema_mode SchemaMode

How strictly the two sides' columns must agree. Defaults to intersection.

strict_types bool

Whether a column stored as different types on the two sides fails every row. Defaults to False, which compares such a column anyway: two numeric types compare by value, so an integer 10 and a float 10.7 differ, and any other pair casts the target to the source type, with values that do not convert becoming NULL.

normalize_column_names bool

Whether to strip whitespace from column names and lowercase them before alignment. Defaults to False.

default_absolute_tolerance float

Absolute tolerance for numeric columns whose rule sets none. Defaults to 0.

default_relative_tolerance float

Relative tolerance for numeric columns whose rule sets none. Defaults to 0.

default_treat_null_as_equal bool

Whether two NULLs match in columns whose rule does not say. Defaults to True.

default_whitespace_mode WhitespaceMode

Whitespace mode for text columns whose rule sets none. Defaults to none.

default_null_values list[SentinelValue]

Values to read as NULL in every column. Each applies only to columns whose type can hold it, and the rest are skipped without an error.

rules list[DiffRule]

Per-column overrides. A rule naming a column wins over a pattern rule.

threshold float

Largest share of mismatched rows, from 0 to 1, that still counts as a match. Defaults to 0.

report_top_columns_limit int

Most drifted columns to list in report_summary and the Markdown summary. Defaults to 5.

pushdown_sample_rows int

Changed rows a pushdown run fetches with both sides' values, so the HTML report and the result can show values instead of keys. Defaults to 0, which fetches none, so no value leaves the warehouse. A local run holds every row already.

output_path str | None

Folder to write the added, removed, and changed rows to. None, the default, writes nothing.

output_format ArtifactFormat

File format of those rows. Defaults to parquet.

Source code in src/veridelta/models.py
class DiffConfig(BaseModel):
    """Settings and rules for one comparison.

    Attributes:
        primary_keys (list[str]): Columns that join the datasets. At least one is
            required, and together they must be unique in each dataset.
        schema_mode (SchemaMode): How strictly the two sides' columns must agree.
            Defaults to `intersection`.
        strict_types (bool): Whether a column stored as different types on the two
            sides fails every row. Defaults to False, which compares such a column
            anyway: two numeric types compare by value, so an integer `10` and a
            float `10.7` differ, and any other pair casts the target to the source
            type, with values that do not convert becoming NULL.
        normalize_column_names (bool): Whether to strip whitespace from column
            names and lowercase them before alignment. Defaults to False.
        default_absolute_tolerance (float): Absolute tolerance for numeric columns
            whose rule sets none. Defaults to 0.
        default_relative_tolerance (float): Relative tolerance for numeric columns
            whose rule sets none. Defaults to 0.
        default_treat_null_as_equal (bool): Whether two NULLs match in columns whose
            rule does not say. Defaults to True.
        default_whitespace_mode (WhitespaceMode): Whitespace mode for text columns
            whose rule sets none. Defaults to `none`.
        default_null_values (list[SentinelValue]): Values to read as NULL in every
            column. Each applies only to columns whose type can hold it, and the
            rest are skipped without an error.
        rules (list[DiffRule]): Per-column overrides. A rule naming a column wins
            over a `pattern` rule.
        threshold (float): Largest share of mismatched rows, from 0 to 1, that
            still counts as a match. Defaults to 0.
        report_top_columns_limit (int): Most drifted columns to list in
            `report_summary` and the Markdown summary. Defaults to 5.
        pushdown_sample_rows (int): Changed rows a pushdown run fetches with both
            sides' values, so the HTML report and the result can show values
            instead of keys. Defaults to 0, which fetches none, so no value leaves
            the warehouse. A local run holds every row already.
        output_path (str | None): Folder to write the added, removed, and changed
            rows to. None, the default, writes nothing.
        output_format (ArtifactFormat): File format of those rows. Defaults to
            `parquet`.
    """

    model_config = ConfigDict(extra="forbid")

    primary_keys: list[str] = Field(
        ..., min_length=1, description="Columns used to join datasets (at least one)."
    )

    schema_mode: SchemaMode = Field(
        default="intersection",
        description="Schema enforcement mode: 'exact', 'allow_additions', 'allow_removals', or 'intersection'.",
    )
    strict_types: bool = Field(
        default=False,
        description=(
            "Whether differing column types fail every row. If False, numeric types compare "
            "by value and other types cast the target to the source type."
        ),
    )

    normalize_column_names: bool = Field(
        default=False,
        description="Whether to strip whitespace from column names and lowercase them.",
    )

    default_absolute_tolerance: float = Field(
        default=0.0,
        ge=0.0,
        strict=True,
        allow_inf_nan=False,
        description="Global absolute tolerance for numeric columns.",
    )
    default_relative_tolerance: float = Field(
        default=0.0,
        ge=0.0,
        strict=True,
        allow_inf_nan=False,
        description="Global relative tolerance for numeric columns.",
    )
    default_treat_null_as_equal: bool = Field(
        default=True, description="Globally treat NULL == NULL as a match."
    )
    default_whitespace_mode: WhitespaceMode = Field(
        default="none",
        description="Global string whitespace stripping mode: 'none', 'left', 'right', or 'both'.",
    )
    default_null_values: list[SentinelValue] = Field(
        default_factory=list[SentinelValue],
        strict=True,
        description="Global list of values to coerce to NULL, applied per matching dtype.",
    )

    rules: list[DiffRule] = Field(default_factory=list[DiffRule], description="Column overrides.")

    threshold: float = Field(
        default=0.0, ge=0.0, le=1.0, description="Allowed mismatch ratio (0.0 to 1.0)."
    )

    report_top_columns_limit: int = Field(
        default=5,
        ge=0,
        description="Max number of top drifted columns to show in the report summary.",
    )

    pushdown_sample_rows: int = Field(
        default=0,
        ge=0,
        strict=True,
        description=(
            "Pushdown only: fetch up to this many changed rows with both "
            "sides' values. 0 fetches none, so no value leaves the warehouse."
        ),
    )

    output_path: str | None = Field(
        default=None, description="Directory to write discrepancy artifacts to."
    )
    output_format: ArtifactFormat = Field(
        default="parquet",
        description="File format for exported discrepancy artifacts, such as 'parquet'.",
    )

    @field_validator("default_null_values")
    @classmethod
    def validate_default_null_values(cls, v: list[SentinelValue]) -> list[SentinelValue]:
        """Reject a default sentinel that can never match, such as NaN.

        Args:
            v (list[SentinelValue]): Configured default sentinels.

        Returns:
            list[SentinelValue]: The sentinels, unchanged.

        Raises:
            ValueError: If any entry is a non-finite float.
        """
        _reject_non_finite_sentinels(v)
        return v

    @model_validator(mode="after")
    def apply_schema_normalization(self) -> "DiffConfig":
        r"""Lowercase and strip configured column names when normalization is enabled.

        Keys, rule `column_names`, and `rename_to` are normalized the same way
        the engine normalizes headers, so every name still refers to a column.
        A `pattern` is left alone: lowercasing a regex changes what it means
        (`\D` is not `\d`), so patterns are written against the lowercase names.
        Rules are copied rather than edited, since the caller may still hold them.

        Returns:
            DiffConfig: The configuration with normalized names.
        """
        if self.normalize_column_names:
            self.primary_keys = [normalize_column_name(pk) for pk in self.primary_keys]
            self.rules = [
                rule.model_copy(
                    update={
                        "column_names": [normalize_column_name(col) for col in rule.column_names],
                        "rename_to": rule.rename_to and normalize_column_name(rule.rename_to),
                    }
                )
                for rule in self.rules
            ]

        return self

apply_schema_normalization()

Lowercase and strip configured column names when normalization is enabled.

Keys, rule column_names, and rename_to are normalized the same way the engine normalizes headers, so every name still refers to a column. A pattern is left alone: lowercasing a regex changes what it means (\D is not \d), so patterns are written against the lowercase names. Rules are copied rather than edited, since the caller may still hold them.

Returns:

Name Type Description
DiffConfig DiffConfig

The configuration with normalized names.

Source code in src/veridelta/models.py
@model_validator(mode="after")
def apply_schema_normalization(self) -> "DiffConfig":
    r"""Lowercase and strip configured column names when normalization is enabled.

    Keys, rule `column_names`, and `rename_to` are normalized the same way
    the engine normalizes headers, so every name still refers to a column.
    A `pattern` is left alone: lowercasing a regex changes what it means
    (`\D` is not `\d`), so patterns are written against the lowercase names.
    Rules are copied rather than edited, since the caller may still hold them.

    Returns:
        DiffConfig: The configuration with normalized names.
    """
    if self.normalize_column_names:
        self.primary_keys = [normalize_column_name(pk) for pk in self.primary_keys]
        self.rules = [
            rule.model_copy(
                update={
                    "column_names": [normalize_column_name(col) for col in rule.column_names],
                    "rename_to": rule.rename_to and normalize_column_name(rule.rename_to),
                }
            )
            for rule in self.rules
        ]

    return self

validate_default_null_values(v) classmethod

Reject a default sentinel that can never match, such as NaN.

Parameters:

Name Type Description Default
v list[SentinelValue]

Configured default sentinels.

required

Returns:

Type Description
list[SentinelValue]

list[SentinelValue]: The sentinels, unchanged.

Raises:

Type Description
ValueError

If any entry is a non-finite float.

Source code in src/veridelta/models.py
@field_validator("default_null_values")
@classmethod
def validate_default_null_values(cls, v: list[SentinelValue]) -> list[SentinelValue]:
    """Reject a default sentinel that can never match, such as NaN.

    Args:
        v (list[SentinelValue]): Configured default sentinels.

    Returns:
        list[SentinelValue]: The sentinels, unchanged.

    Raises:
        ValueError: If any entry is a non-finite float.
    """
    _reject_non_finite_sentinels(v)
    return v

DiffResult dataclass

A completed comparison: the counts plus the rows behind them.

DiffSummary serializes to JSON, so it cannot carry frames. This class pairs it with the discrepancy rows the engine already materialized.

Attributes:

Name Type Description
summary DiffSummary

Counts, ratios, and the formatted report.

added DataFrame

Rows present only in the target.

removed DataFrame

Rows present only in the source.

changed DataFrame

Rows present in both with at least one differing column. A local run carries {column}_source, {column}_target, and {column}_is_match for every compared column, and a pushdown run carries primary keys alone.

primary_keys tuple[str, ...]

Join keys, in configured order.

compared_columns tuple[str, ...]

Columns compared, after renames and exclusions. Recorded so get_mismatches refuses a mistyped name on both engines, including pushdown, whose frames cannot show which columns were compared.

keys_only bool

Whether the run was pushdown, which returns primary keys instead of rows.

changed_sample DataFrame | None

Pushdown only, when pushdown_sample_rows is set: up to that many changed rows, in key order, with {column}_source, {column}_target, and {column}_is_match for every compared column, as a local run's changed carries them. None when no sample was asked for, nothing changed, or the run was local, where changed holds every row.

accepted Baseline | None

What a baseline accepted in this run: the rows it matched, and on each changed row the columns it accepted. None when the run had no baseline.

Source code in src/veridelta/models.py
@dataclass(frozen=True)
class DiffResult:
    """A completed comparison: the counts plus the rows behind them.

    `DiffSummary` serializes to JSON, so it cannot carry frames. This class pairs it
    with the discrepancy rows the engine already materialized.

    Attributes:
        summary (DiffSummary): Counts, ratios, and the formatted report.
        added (pl.DataFrame): Rows present only in the target.
        removed (pl.DataFrame): Rows present only in the source.
        changed (pl.DataFrame): Rows present in both with at least one differing
            column. A local run carries `{column}_source`, `{column}_target`, and
            `{column}_is_match` for every compared column, and a pushdown run
            carries primary keys alone.
        primary_keys (tuple[str, ...]): Join keys, in configured order.
        compared_columns (tuple[str, ...]): Columns compared, after renames and
            exclusions. Recorded so `get_mismatches` refuses a mistyped name on both
            engines, including pushdown, whose frames cannot show which columns
            were compared.
        keys_only (bool): Whether the run was pushdown, which returns primary keys
            instead of rows.
        changed_sample (pl.DataFrame | None): Pushdown only, when
            `pushdown_sample_rows` is set: up to that many changed rows, in key
            order, with `{column}_source`, `{column}_target`, and
            `{column}_is_match` for every compared column, as a local run's
            `changed` carries them. None when no sample was asked for, nothing
            changed, or the run was local, where `changed` holds every row.
        accepted (Baseline | None): What a baseline accepted in this run: the
            rows it matched, and on each changed row the columns it accepted.
            None when the run had no baseline.
    """

    summary: DiffSummary
    added: pl.DataFrame
    removed: pl.DataFrame
    changed: pl.DataFrame
    primary_keys: tuple[str, ...] = ()
    compared_columns: tuple[str, ...] = ()
    keys_only: bool = False
    changed_sample: pl.DataFrame | None = None
    accepted: "Baseline | None" = None

    def get_mismatches(self, column: str) -> pl.DataFrame:
        """Isolate the rows where one column disagreed.

        Args:
            column (str): Compared column to isolate, named as it appears after
                any `rename_to`.

        Returns:
            pl.DataFrame: Primary keys alongside the source and target values,
            restricted to rows where this column differed. Pushdown runs return
            every changed primary key instead, since the warehouse never
            projected the values and cannot attribute a row to one column.

        Raises:
            ConfigError: If the column was not part of the comparison.

        Examples:
            >>> import polars as pl
            >>> from veridelta.engine import DiffEngine
            >>> source = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 20.0]})
            >>> target = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 21.5]})
            >>> result = DiffEngine(DiffConfig(primary_keys=["id"]), source, target).run()
            >>> result.get_mismatches("amount")["id"].to_list()
            [2]
        """
        if column not in self.compared_columns:
            compared = ", ".join(self.compared_columns) or "none"
            raise ConfigError(
                f"Column '{column}' was not compared, so it has no mismatches to "
                f"report. Compared columns: {compared}."
            )
        if self.keys_only:
            return self.changed
        return self.changed.filter(~pl.col(f"{column}_is_match")).select(
            [*self.primary_keys, f"{column}_source", f"{column}_target"]
        )

    def to_pandas(self) -> "pd.DataFrame":
        """Convert the changed rows to pandas for notebook use.

        Returns:
            pd.DataFrame: `changed` as a pandas DataFrame. The other frames
            convert the same way through Polars' own `to_pandas`.

        Raises:
            ConfigError: If pandas or pyarrow is not installed.
        """
        try:
            return self.changed.to_pandas()
        except ImportError as exc:
            raise ConfigError(
                "Converting to pandas requires pandas and pyarrow, which Veridelta "
                f"does not depend on. Install them to use this method. ({exc})"
            ) from exc

get_mismatches(column)

Isolate the rows where one column disagreed.

Parameters:

Name Type Description Default
column str

Compared column to isolate, named as it appears after any rename_to.

required

Returns:

Type Description
DataFrame

pl.DataFrame: Primary keys alongside the source and target values,

DataFrame

restricted to rows where this column differed. Pushdown runs return

DataFrame

every changed primary key instead, since the warehouse never

DataFrame

projected the values and cannot attribute a row to one column.

Raises:

Type Description
ConfigError

If the column was not part of the comparison.

Examples:

>>> import polars as pl
>>> from veridelta.engine import DiffEngine
>>> source = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 20.0]})
>>> target = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 21.5]})
>>> result = DiffEngine(DiffConfig(primary_keys=["id"]), source, target).run()
>>> result.get_mismatches("amount")["id"].to_list()
[2]
Source code in src/veridelta/models.py
def get_mismatches(self, column: str) -> pl.DataFrame:
    """Isolate the rows where one column disagreed.

    Args:
        column (str): Compared column to isolate, named as it appears after
            any `rename_to`.

    Returns:
        pl.DataFrame: Primary keys alongside the source and target values,
        restricted to rows where this column differed. Pushdown runs return
        every changed primary key instead, since the warehouse never
        projected the values and cannot attribute a row to one column.

    Raises:
        ConfigError: If the column was not part of the comparison.

    Examples:
        >>> import polars as pl
        >>> from veridelta.engine import DiffEngine
        >>> source = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 20.0]})
        >>> target = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 21.5]})
        >>> result = DiffEngine(DiffConfig(primary_keys=["id"]), source, target).run()
        >>> result.get_mismatches("amount")["id"].to_list()
        [2]
    """
    if column not in self.compared_columns:
        compared = ", ".join(self.compared_columns) or "none"
        raise ConfigError(
            f"Column '{column}' was not compared, so it has no mismatches to "
            f"report. Compared columns: {compared}."
        )
    if self.keys_only:
        return self.changed
    return self.changed.filter(~pl.col(f"{column}_is_match")).select(
        [*self.primary_keys, f"{column}_source", f"{column}_target"]
    )

to_pandas()

Convert the changed rows to pandas for notebook use.

Returns:

Type Description
DataFrame

pd.DataFrame: changed as a pandas DataFrame. The other frames

DataFrame

convert the same way through Polars' own to_pandas.

Raises:

Type Description
ConfigError

If pandas or pyarrow is not installed.

Source code in src/veridelta/models.py
def to_pandas(self) -> "pd.DataFrame":
    """Convert the changed rows to pandas for notebook use.

    Returns:
        pd.DataFrame: `changed` as a pandas DataFrame. The other frames
        convert the same way through Polars' own `to_pandas`.

    Raises:
        ConfigError: If pandas or pyarrow is not installed.
    """
    try:
        return self.changed.to_pandas()
    except ImportError as exc:
        raise ConfigError(
            "Converting to pandas requires pandas and pyarrow, which Veridelta "
            f"does not depend on. Install them to use this method. ({exc})"
        ) from exc

DiffRule

Bases: BaseModel

Overrides for one or more columns, chosen by exact name or by pattern.

Each setting runs at a fixed stage of the transform order, the same in a local run and in a warehouse.

Attributes:

Name Type Description
column_names list[str]

Exact source column names the rule governs.

pattern str | None

Regular expression that selects columns by name, such as ^AMT_.*.

absolute_tolerance float | None

Largest absolute difference that still matches, for numeric columns. Must be finite.

relative_tolerance float | None

Largest difference relative to the source value that still matches, such as 0.01 for 1%. Must be finite. Neither tolerance forgives a non-finite value: NaN matches only NaN, and an infinity only itself.

max_levenshtein_distance int | None

Most single-character insertions, deletions, and substitutions that still match, for columns compared as text. Needs the fuzzy extra locally.

min_jaro_winkler_similarity float | None

Lowest Jaro-Winkler similarity, above 0 and at most 1, that still matches, for columns compared as text. Needs the fuzzy extra, and runs locally only: warehouse pushdown refuses it before any query runs. A rule sets at most one of the two limits.

case_insensitive bool | None

Whether to ignore case in text.

whitespace_mode WhitespaceMode | None

Which ends of text to strip whitespace from.

regex_replace dict[str, str] | None

{pattern: replacement} pairs applied to text columns.

pad_zeros int | None

Width to left-pad values to with zeros, such as 5 for 00123. Other types become text first, so a numeric 123 matches a text '00123'.

value_map dict[str, str] | None

Source values to translate to target values before comparison, such as {'M': 'Male'}.

null_values list[SentinelValue] | None

Values to read as NULL, such as ['N/A', -999, false]. Each applies only to columns whose type can hold it: text reaches string, categorical, and enum columns, numbers reach numeric columns, decimals included, and booleans reach boolean columns. An explicit rule whose sentinels fit no column type raises ConfigError.

treat_null_as_equal bool | None

Whether two NULLs match.

datetime_format str | None

Format for strptime, such as %Y-%m-%d %H:%M:%S, that parses text columns into datetimes. A value that does not match becomes NULL. Pushdown translates the directives from a fixed table and raises ConfigError for any directive outside it.

timezone str | None

Timezone to convert timestamps to, such as UTC. The column must be timezone-aware once parsed: a naive timestamp raises ConfigError instead of being assigned a guessed zone.

cast_to CastTarget | None

Polars type to cast the column to, such as Float64. A closed set, so a typo fails when the config loads.

ignore bool

Whether to leave the column out of the comparison. Defaults to False.

rename_to str | None

Target column name, when it differs from the source. Valid only when column_names holds exactly one name.

Source code in src/veridelta/models.py
class DiffRule(BaseModel):
    """Overrides for one or more columns, chosen by exact name or by pattern.

    Each setting runs at a fixed stage of the
    [transform order](https://veridelta.github.io/veridelta/rules/#transform-order),
    the same in a local run and in a warehouse.

    Attributes:
        column_names (list[str]): Exact source column names the rule governs.
        pattern (str | None): Regular expression that selects columns by name, such
            as `^AMT_.*`.
        absolute_tolerance (float | None): Largest absolute difference that still
            matches, for numeric columns. Must be finite.
        relative_tolerance (float | None): Largest difference relative to the source
            value that still matches, such as `0.01` for 1%. Must be finite. Neither
            tolerance forgives a non-finite value: NaN matches only NaN, and an
            infinity only itself.
        max_levenshtein_distance (int | None): Most single-character insertions,
            deletions, and substitutions that still match, for columns compared as
            text. Needs the `fuzzy` extra locally.
        min_jaro_winkler_similarity (float | None): Lowest Jaro-Winkler similarity,
            above 0 and at most 1, that still matches, for columns compared as text.
            Needs the `fuzzy` extra, and runs locally only: warehouse pushdown refuses
            it before any query runs. A rule sets at most one of the two limits.
        case_insensitive (bool | None): Whether to ignore case in text.
        whitespace_mode (WhitespaceMode | None): Which ends of text to strip
            whitespace from.
        regex_replace (dict[str, str] | None): `{pattern: replacement}` pairs applied
            to text columns.
        pad_zeros (int | None): Width to left-pad values to with zeros, such as `5`
            for `00123`. Other types become text first, so a numeric `123` matches a
            text `'00123'`.
        value_map (dict[str, str] | None): Source values to translate to target
            values before comparison, such as `{'M': 'Male'}`.
        null_values (list[SentinelValue] | None): Values to read as NULL, such as
            `['N/A', -999, false]`. Each applies only to columns whose type can hold
            it: text reaches string, categorical, and enum columns, numbers reach
            numeric columns, decimals included, and booleans reach boolean columns.
            An explicit rule whose sentinels fit no column type raises `ConfigError`.
        treat_null_as_equal (bool | None): Whether two NULLs match.
        datetime_format (str | None): Format for `strptime`, such as
            `%Y-%m-%d %H:%M:%S`, that parses text columns into datetimes. A value that does not match
            becomes NULL. Pushdown translates the directives from a fixed table and
            raises `ConfigError` for any directive outside it.
        timezone (str | None): Timezone to convert timestamps to, such as `UTC`. The
            column must be timezone-aware once parsed: a naive timestamp raises
            `ConfigError` instead of being assigned a guessed zone.
        cast_to (CastTarget | None): Polars type to cast the column to, such as
            `Float64`. A closed set, so a typo fails when the config loads.
        ignore (bool): Whether to leave the column out of the comparison. Defaults
            to False.
        rename_to (str | None): Target column name, when it differs from the source.
            Valid only when `column_names` holds exactly one name.
    """

    model_config = ConfigDict(extra="forbid")

    column_names: list[str] = Field(
        default_factory=list, description="Exact names of the columns in the source."
    )
    pattern: str | None = Field(
        default=None, description="Regex pattern that selects columns by name, such as '^AMT_.*'."
    )

    absolute_tolerance: float | None = Field(
        default=None,
        ge=0.0,
        strict=True,
        allow_inf_nan=False,
        description="Absolute tolerance for numeric differences.",
    )
    relative_tolerance: float | None = Field(
        default=None,
        ge=0.0,
        strict=True,
        allow_inf_nan=False,
        description="Relative tolerance, such as 0.01 for 1%.",
    )
    max_levenshtein_distance: int | None = Field(
        default=None,
        ge=1,
        strict=True,
        description="Most character edits between two text values that still match.",
    )
    min_jaro_winkler_similarity: float | None = Field(
        default=None,
        gt=0.0,
        le=1.0,
        strict=True,
        allow_inf_nan=False,
        description="Lowest Jaro-Winkler similarity between two text values that still matches.",
    )

    case_insensitive: bool | None = Field(
        default=None, description="Ignore case for string comparisons."
    )
    whitespace_mode: WhitespaceMode | None = Field(
        default=None,
        description="Whitespace stripping mode: 'none', 'left', 'right', or 'both'.",
    )
    regex_replace: dict[str, str] | None = Field(
        default=None,
        description="Dictionary of {regex_pattern: replacement_string} to sanitize text.",
    )
    pad_zeros: int | None = Field(
        default=None,
        ge=0,
        strict=True,
        description="Left-pad numeric strings to this length, such as 5 for '00123'.",
    )

    value_map: dict[str, str] | None = Field(
        default=None,
        description="Source values to translate to target values, such as {'M': 'Male'}.",
    )
    null_values: list[SentinelValue] | None = Field(
        default=None,
        strict=True,
        description="Values to read as NULL, such as ['N/A', -999, false].",
    )
    treat_null_as_equal: bool | None = Field(
        default=None, description="Treat missing values (NULL/None) in both sources as a match."
    )

    datetime_format: str | None = Field(
        default=None, description="Format for strptime, such as '%Y-%m-%d %H:%M:%S'."
    )
    timezone: str | None = Field(
        default=None, description="Target timezone to normalize dates to before comparison."
    )

    cast_to: CastTarget | None = Field(
        default=None,
        description="Polars type to cast the column to, such as 'Float64'.",
    )
    ignore: bool = Field(
        default=False, description="Whether to leave this column out of the comparison."
    )
    rename_to: str | None = Field(
        default=None,
        description="Name in target dataset if different (use only for single columns).",
    )

    @field_validator("pattern")
    @classmethod
    def validate_pattern(cls, v: str | None) -> str | None:
        """Reject a `pattern` that is not a valid regular expression.

        Args:
            v (str | None): Configured pattern.

        Returns:
            str | None: The pattern, unchanged.

        Raises:
            ValueError: If the pattern does not compile.
        """
        if v is not None:
            try:
                re.compile(v)
            except re.error as err:
                raise ValueError(f"Invalid regex pattern '{v}': {err}") from err
        return v

    @field_validator("null_values")
    @classmethod
    def validate_null_values(cls, v: list[SentinelValue] | None) -> list[SentinelValue] | None:
        """Reject a sentinel that can never match, such as NaN.

        Args:
            v (list[SentinelValue] | None): Configured sentinels.

        Returns:
            list[SentinelValue] | None: The sentinels, unchanged.

        Raises:
            ValueError: If any entry is a non-finite float.
        """
        _reject_non_finite_sentinels(v)
        return v

    @field_validator("regex_replace")
    @classmethod
    def validate_regex_replace(cls, v: dict[str, str] | None) -> dict[str, str] | None:
        """Reject a `regex_replace` key that is not a valid regular expression.

        Args:
            v (dict[str, str] | None): Patterns and their replacements.

        Returns:
            dict[str, str] | None: The mapping, unchanged.

        Raises:
            ValueError: If a pattern does not compile.
        """
        if v is not None:
            for pattern in v:
                try:
                    re.compile(pattern)
                except re.error as err:
                    raise ValueError(f"Invalid regex replace pattern '{pattern}': {err}") from err
        return v

    @model_validator(mode="after")
    def validate_similarity_measure(self) -> "DiffRule":
        """Reject a rule that sets both text similarity limits.

        Returns:
            DiffRule: The validated rule.

        Raises:
            ValueError: If both `max_levenshtein_distance` and
                `min_jaro_winkler_similarity` are set.
        """
        if (
            self.max_levenshtein_distance is not None
            and self.min_jaro_winkler_similarity is not None
        ):
            raise ValueError(
                "Set max_levenshtein_distance or min_jaro_winkler_similarity, not both: "
                "a rule compares text with one similarity measure."
            )
        return self

validate_null_values(v) classmethod

Reject a sentinel that can never match, such as NaN.

Parameters:

Name Type Description Default
v list[SentinelValue] | None

Configured sentinels.

required

Returns:

Type Description
list[SentinelValue] | None

list[SentinelValue] | None: The sentinels, unchanged.

Raises:

Type Description
ValueError

If any entry is a non-finite float.

Source code in src/veridelta/models.py
@field_validator("null_values")
@classmethod
def validate_null_values(cls, v: list[SentinelValue] | None) -> list[SentinelValue] | None:
    """Reject a sentinel that can never match, such as NaN.

    Args:
        v (list[SentinelValue] | None): Configured sentinels.

    Returns:
        list[SentinelValue] | None: The sentinels, unchanged.

    Raises:
        ValueError: If any entry is a non-finite float.
    """
    _reject_non_finite_sentinels(v)
    return v

validate_pattern(v) classmethod

Reject a pattern that is not a valid regular expression.

Parameters:

Name Type Description Default
v str | None

Configured pattern.

required

Returns:

Type Description
str | None

str | None: The pattern, unchanged.

Raises:

Type Description
ValueError

If the pattern does not compile.

Source code in src/veridelta/models.py
@field_validator("pattern")
@classmethod
def validate_pattern(cls, v: str | None) -> str | None:
    """Reject a `pattern` that is not a valid regular expression.

    Args:
        v (str | None): Configured pattern.

    Returns:
        str | None: The pattern, unchanged.

    Raises:
        ValueError: If the pattern does not compile.
    """
    if v is not None:
        try:
            re.compile(v)
        except re.error as err:
            raise ValueError(f"Invalid regex pattern '{v}': {err}") from err
    return v

validate_regex_replace(v) classmethod

Reject a regex_replace key that is not a valid regular expression.

Parameters:

Name Type Description Default
v dict[str, str] | None

Patterns and their replacements.

required

Returns:

Type Description
dict[str, str] | None

dict[str, str] | None: The mapping, unchanged.

Raises:

Type Description
ValueError

If a pattern does not compile.

Source code in src/veridelta/models.py
@field_validator("regex_replace")
@classmethod
def validate_regex_replace(cls, v: dict[str, str] | None) -> dict[str, str] | None:
    """Reject a `regex_replace` key that is not a valid regular expression.

    Args:
        v (dict[str, str] | None): Patterns and their replacements.

    Returns:
        dict[str, str] | None: The mapping, unchanged.

    Raises:
        ValueError: If a pattern does not compile.
    """
    if v is not None:
        for pattern in v:
            try:
                re.compile(pattern)
            except re.error as err:
                raise ValueError(f"Invalid regex replace pattern '{pattern}': {err}") from err
    return v

validate_similarity_measure()

Reject a rule that sets both text similarity limits.

Returns:

Name Type Description
DiffRule DiffRule

The validated rule.

Raises:

Type Description
ValueError

If both max_levenshtein_distance and min_jaro_winkler_similarity are set.

Source code in src/veridelta/models.py
@model_validator(mode="after")
def validate_similarity_measure(self) -> "DiffRule":
    """Reject a rule that sets both text similarity limits.

    Returns:
        DiffRule: The validated rule.

    Raises:
        ValueError: If both `max_levenshtein_distance` and
            `min_jaro_winkler_similarity` are set.
    """
    if (
        self.max_levenshtein_distance is not None
        and self.min_jaro_winkler_similarity is not None
    ):
        raise ValueError(
            "Set max_levenshtein_distance or min_jaro_winkler_similarity, not both: "
            "a rule compares text with one similarity measure."
        )
    return self

DiffSummary

Bases: BaseModel

Counts and verdict for one comparison.

Attributes:

Name Type Description
total_rows_source int

Rows in the source dataset.

total_rows_target int

Rows in the target dataset.

added_count int

Rows only in the target.

removed_count int

Rows only in the source.

changed_count int

Rows in both datasets with at least one differing column.

column_mismatches dict[str, int]

Mismatched rows per compared column.

is_match bool

Whether the mismatch ratio is within threshold.

accepted_count int

Rows of drift a baseline accepted, which the counts above leave out. Zero without --baseline.

total_mismatches int

Added, removed, and changed rows together.

mismatch_ratio float

total_mismatches divided by the source row count.

match_rate_percentage float

Match rate as a percentage, such as 99.98.

is_perfect_match bool

Whether nothing mismatched.

volume_shift int

Target rows minus source rows.

report_summary str

Plain-text report for CI logs.

report_limit int

Most columns report_summary lists. Left out of JSON.

artifacts_written bool

Whether artifacts were written to output_path. Pushdown writes primary keys only, under _pks_only file names. Left out of JSON.

Source code in src/veridelta/models.py
class DiffSummary(BaseModel):
    """Counts and verdict for one comparison.

    Attributes:
        total_rows_source (int): Rows in the source dataset.
        total_rows_target (int): Rows in the target dataset.
        added_count (int): Rows only in the target.
        removed_count (int): Rows only in the source.
        changed_count (int): Rows in both datasets with at least one differing
            column.
        column_mismatches (dict[str, int]): Mismatched rows per compared column.
        is_match (bool): Whether the mismatch ratio is within `threshold`.
        accepted_count (int): Rows of drift a baseline accepted, which the counts
            above leave out. Zero without `--baseline`.
        total_mismatches (int): Added, removed, and changed rows together.
        mismatch_ratio (float): `total_mismatches` divided by the source row count.
        match_rate_percentage (float): Match rate as a percentage, such as `99.98`.
        is_perfect_match (bool): Whether nothing mismatched.
        volume_shift (int): Target rows minus source rows.
        report_summary (str): Plain-text report for CI logs.
        report_limit (int): Most columns `report_summary` lists. Left out of JSON.
        artifacts_written (bool): Whether artifacts were written to `output_path`.
            Pushdown writes primary keys only, under `_pks_only` file names. Left
            out of JSON.
    """

    model_config = ConfigDict(extra="forbid")

    total_rows_source: int
    total_rows_target: int
    added_count: int
    removed_count: int
    changed_count: int
    column_mismatches: dict[str, int] = Field(default_factory=dict)
    is_match: bool
    accepted_count: int = Field(default=0, ge=0)

    report_limit: int = Field(default=5, exclude=True)
    artifacts_written: bool = Field(default=False, exclude=True)

    @computed_field
    @property
    def total_mismatches(self) -> int:
        """Count added, removed, and changed rows together.

        Returns:
            int: The total.
        """
        return self.added_count + self.removed_count + self.changed_count

    @computed_field
    @property
    def mismatch_ratio(self) -> float:
        """Divide the mismatches by the source row count.

        Returns:
            float: The ratio, which can exceed 1.
        """
        return mismatch_ratio_of(self.total_mismatches, self.total_rows_source)

    @computed_field
    @property
    def match_rate_percentage(self) -> float:
        """Express the match rate as a percentage.

        Returns:
            float: The percentage, rounded to two decimal places, such as `99.98`.
        """
        return round((1.0 - self.mismatch_ratio) * 100.0, 2)

    @computed_field
    @property
    def is_perfect_match(self) -> bool:
        """Return whether nothing mismatched under the configured rules.

        Returns:
            bool: Whether `total_mismatches` is zero.
        """
        return self.total_mismatches == 0

    @computed_field
    @property
    def volume_shift(self) -> int:
        """Subtract the source row count from the target's.

        Returns:
            int: Target rows minus source rows.
        """
        return self.total_rows_target - self.total_rows_source

    @computed_field
    @property
    def report_summary(self) -> str:
        """Format the counts and the most drifted columns as plain text for a CI log.

        Returns:
            str: The report.
        """
        status_icon = "PASSED" if self.is_match else "FAILED"
        perfect_tag = " (Perfect Match)" if self.is_perfect_match else ""

        base_report = (
            f"Veridelta Execution Summary\n"
            f"===========================\n"
            f"Status:        {status_icon}{perfect_tag}\n"
            f"Match Rate:    {self.match_rate_percentage}%\n"
            f"Source Rows:   {self.total_rows_source:,}\n"
            f"Target Rows:   {self.total_rows_target:,}\n"
            f"Volume Shift:  {self.volume_shift:+,} {_noun(self.volume_shift, 'row', 'rows')}\n"
            f"\nRow-Level Discrepancies:\n"
            f"---------------------------\n"
            f"Added:         {self.added_count:,}\n"
            f"Removed:       {self.removed_count:,}\n"
            f"Changed:       {self.changed_count:,}\n"
            f"Total Issues:  {self.total_mismatches:,}\n"
        )
        if self.accepted_count:
            base_report += f"Accepted:      {self.accepted_count:,}\n"

        if not self.column_mismatches or self.report_limit == 0:
            return base_report

        top_cols = sorted(self.column_mismatches.items(), key=lambda x: x[1], reverse=True)[
            : self.report_limit
        ]

        col_report = "".join(
            f"- {col}: {count:,} {_noun(count, 'mismatch', 'mismatches')}\n"
            for col, count in top_cols
        )
        return f"{base_report}\nTop Column-Level Drifts:\n---------------------------\n{col_report}"

is_perfect_match property

Return whether nothing mismatched under the configured rules.

Returns:

Name Type Description
bool bool

Whether total_mismatches is zero.

match_rate_percentage property

Express the match rate as a percentage.

Returns:

Name Type Description
float float

The percentage, rounded to two decimal places, such as 99.98.

mismatch_ratio property

Divide the mismatches by the source row count.

Returns:

Name Type Description
float float

The ratio, which can exceed 1.

report_summary property

Format the counts and the most drifted columns as plain text for a CI log.

Returns:

Name Type Description
str str

The report.

total_mismatches property

Count added, removed, and changed rows together.

Returns:

Name Type Description
int int

The total.

volume_shift property

Subtract the source row count from the target's.

Returns:

Name Type Description
int int

Target rows minus source rows.

DuckDBConfig

Bases: BaseModel

Immutable settings for reading a DuckDB or MotherDuck table or query into a local comparison.

The rows are read through DuckDB into Polars and compared by the local engine, so a DuckDB source pairs with files, lakehouse tables, databases, or another DuckDB source. Requires the duckdb extra (uv add 'veridelta[duckdb]'). Two tables in one database can instead be compared inside DuckDB, without reading their rows, when both sides set pushdown.

Attributes:

Name Type Description
type Literal['duckdb']

Source kind, which selects this model.

database str

Path to a DuckDB file, which opens read-only, or a MotherDuck database written as md:name or motherduck:name.

table str | None

Table or view to read whole, as one to three unquoted identifier segments, such as main.orders.

query str | None

SQL statement to run instead, sent to DuckDB exactly as written. Set exactly one of table and query.

motherduck_token str | None

Token for a MotherDuck database. When unset, the MOTHERDUCK_TOKEN environment variable supplies it, then motherduck_token. Left out when the config is printed, but kept by model_dump().

pushdown bool

Whether to compare inside DuckDB instead of reading the rows. table sources only, and both sides must set it. Defaults to False.

Source code in src/veridelta/models.py
class DuckDBConfig(BaseModel):
    """Immutable settings for reading a DuckDB or MotherDuck table or query into a local comparison.

    The rows are read through DuckDB into Polars and compared by the local
    engine, so a DuckDB source pairs with files, lakehouse tables, databases,
    or another DuckDB source. Requires the `duckdb` extra
    (`uv add 'veridelta[duckdb]'`). Two tables in one database can instead be
    compared inside DuckDB, without reading their rows, when both sides set
    `pushdown`.

    Attributes:
        type (Literal["duckdb"]): Source kind, which selects this model.
        database (str): Path to a DuckDB file, which opens read-only, or a
            MotherDuck database written as `md:name` or `motherduck:name`.
        table (str | None): Table or view to read whole, as one to three
            unquoted identifier segments, such as `main.orders`.
        query (str | None): SQL statement to run instead, sent to DuckDB
            exactly as written. Set exactly one of `table` and `query`.
        motherduck_token (str | None): Token for a MotherDuck database.
            When unset, the `MOTHERDUCK_TOKEN` environment variable supplies
            it, then `motherduck_token`. Left out when the config is printed,
            but kept by `model_dump()`.
        pushdown (bool): Whether to compare inside DuckDB instead of reading
            the rows. `table` sources only, and both sides must set it.
            Defaults to False.
    """

    model_config = ConfigDict(extra="forbid", frozen=True, hide_input_in_errors=True)

    type: Literal["duckdb"] = Field(
        "duckdb", description="Discriminator for DuckDB and MotherDuck."
    )
    database: str = Field(
        ...,
        min_length=1,
        description="DuckDB file path, or md:name or motherduck:name for a MotherDuck database.",
    )
    table: str | None = Field(
        default=None,
        pattern=SQL_RELATION_PATTERN,
        description="Table or view to read whole.",
    )
    query: str | None = Field(
        default=None,
        min_length=1,
        description="SQL statement to run instead of reading a table, sent as written.",
    )
    motherduck_token: str | None = Field(
        default=None,
        repr=False,
        description="Token for a MotherDuck database. Defaults to MOTHERDUCK_TOKEN.",
    )
    pushdown: bool = Field(
        default=False,
        strict=True,
        description=(
            "Compare inside DuckDB instead of reading the rows. Tables only; set it on both sides."
        ),
    )

    @property
    def is_motherduck(self) -> bool:
        """Return whether `database` names a MotherDuck database rather than a file.

        Returns:
            bool: Whether `database` starts with `md:` or `motherduck:`, in any case.
        """
        return self.database.lower().startswith(_MOTHERDUCK_PREFIXES)

    @model_validator(mode="after")
    def validate_connection(self) -> "DuckDBConfig":
        """Reject a source that is ambiguous about its rows or would expose its token.

        Returns:
            DuckDBConfig: The validated instance.

        Raises:
            ValueError: If both or neither of `table` and `query` are set,
                `database` is in memory or holds a token, a token is set for a
                file, or `pushdown` is set on a `query`.
        """
        if (self.table is None) == (self.query is None):
            raise ValueError("A DuckDB source reads a 'table' or runs a 'query'; set exactly one.")
        if self.pushdown and self.table is None:
            raise ValueError("'pushdown' compares two tables, so set 'table' rather than 'query'.")
        # Messages never repeat `database`, which may hold the token they refuse.
        if self.database.lower().startswith(":memory:"):
            raise ValueError(
                "':memory:' opens a new, empty database. Point 'database' at a DuckDB "
                "file or an md: database."
            )
        if self.is_motherduck and "token" in self.database.partition("?")[2].lower():
            raise ValueError(
                "Set the MotherDuck token in 'motherduck_token' or MOTHERDUCK_TOKEN rather "
                "than in 'database', which is printed and logged."
            )
        if self.motherduck_token is not None and not self.is_motherduck:
            raise ValueError("'motherduck_token' is for an md: database; a DuckDB file has none.")
        return self

is_motherduck property

Return whether database names a MotherDuck database rather than a file.

Returns:

Name Type Description
bool bool

Whether database starts with md: or motherduck:, in any case.

validate_connection()

Reject a source that is ambiguous about its rows or would expose its token.

Returns:

Name Type Description
DuckDBConfig DuckDBConfig

The validated instance.

Raises:

Type Description
ValueError

If both or neither of table and query are set, database is in memory or holds a token, a token is set for a file, or pushdown is set on a query.

Source code in src/veridelta/models.py
@model_validator(mode="after")
def validate_connection(self) -> "DuckDBConfig":
    """Reject a source that is ambiguous about its rows or would expose its token.

    Returns:
        DuckDBConfig: The validated instance.

    Raises:
        ValueError: If both or neither of `table` and `query` are set,
            `database` is in memory or holds a token, a token is set for a
            file, or `pushdown` is set on a `query`.
    """
    if (self.table is None) == (self.query is None):
        raise ValueError("A DuckDB source reads a 'table' or runs a 'query'; set exactly one.")
    if self.pushdown and self.table is None:
        raise ValueError("'pushdown' compares two tables, so set 'table' rather than 'query'.")
    # Messages never repeat `database`, which may hold the token they refuse.
    if self.database.lower().startswith(":memory:"):
        raise ValueError(
            "':memory:' opens a new, empty database. Point 'database' at a DuckDB "
            "file or an md: database."
        )
    if self.is_motherduck and "token" in self.database.partition("?")[2].lower():
        raise ValueError(
            "Set the MotherDuck token in 'motherduck_token' or MOTHERDUCK_TOKEN rather "
            "than in 'database', which is printed and logged."
        )
    if self.motherduck_token is not None and not self.is_motherduck:
        raise ValueError("'motherduck_token' is for an md: database; a DuckDB file has none.")
    return self

IcebergConfig

Bases: BaseModel

Immutable settings for an Apache Iceberg table scan.

Attributes:

Name Type Description
type Literal['iceberg']

Source kind, which selects this model.

table_uri str

Catalog identifier or filesystem URI of the table.

snapshot_id int | None

Optional snapshot to time-travel.

storage_options dict[str, str]

Object-store credentials and options. Left out when the config is printed, but kept by model_dump(), which the scanner needs.

Source code in src/veridelta/models.py
class IcebergConfig(BaseModel):
    """Immutable settings for an Apache Iceberg table scan.

    Attributes:
        type (Literal["iceberg"]): Source kind, which selects this model.
        table_uri (str): Catalog identifier or filesystem URI of the table.
        snapshot_id (int | None): Optional snapshot to time-travel.
        storage_options (dict[str, str]): Object-store credentials and options.
            Left out when the config is printed, but kept by `model_dump()`,
            which the scanner needs.
    """

    model_config = ConfigDict(extra="forbid", frozen=True, hide_input_in_errors=True)

    type: Literal["iceberg"] = Field("iceberg", description="Discriminator for Apache Iceberg.")
    table_uri: str = Field(..., description="Catalog identifier or filesystem URI of the table.")
    snapshot_id: int | None = Field(
        default=None,
        ge=0,
        strict=True,
        description="Optional snapshot identifier to time-travel.",
    )
    storage_options: dict[str, str] = Field(
        default_factory=dict,
        repr=False,
        description="Object-store credentials and options passed to Polars.",
    )

RuleSuggestion

Bases: BaseModel

A rule suggest proposes for one column, with the evidence for it.

Attributes:

Name Type Description
column str

Compared column, named as it appears after any rename_to.

settings dict[str, SuggestedSetting]

The DiffRule settings the rule adds, such as {"absolute_tolerance": 0.005}.

differing int

Joined rows whose values in the column differ under the declared rules.

explained int

Of those, the rows that match once the rule is added, counted by running the comparison again with it.

largest_gap float | None

For a tolerance, the largest absolute difference among the rows it explains, and None for other settings.

examples tuple[dict[str, Any], ...]

The primary keys of up to three rows the rule explains, lowest first.

rule DiffRule

The rule to add: the settings of the rule that governs the column today, if any, with settings added, naming the column alone.

governing_rule_index int | None

Position in DiffConfig.rules of the rule that governs the column today, or None when no rule does.

Source code in src/veridelta/models.py
class RuleSuggestion(BaseModel):
    """A rule `suggest` proposes for one column, with the evidence for it.

    Attributes:
        column (str): Compared column, named as it appears after any `rename_to`.
        settings (dict[str, SuggestedSetting]): The `DiffRule` settings the rule
            adds, such as `{"absolute_tolerance": 0.005}`.
        differing (int): Joined rows whose values in the column differ under the
            declared rules.
        explained (int): Of those, the rows that match once the rule is added,
            counted by running the comparison again with it.
        largest_gap (float | None): For a tolerance, the largest absolute
            difference among the rows it explains, and None for other settings.
        examples (tuple[dict[str, Any], ...]): The primary keys of up to three
            rows the rule explains, lowest first.
        rule (DiffRule): The rule to add: the settings of the rule that governs
            the column today, if any, with `settings` added, naming the column
            alone.
        governing_rule_index (int | None): Position in `DiffConfig.rules` of the
            rule that governs the column today, or None when no rule does.
    """

    model_config = ConfigDict(extra="forbid", frozen=True)

    column: str = Field(..., description="Compared column, after any rename_to.")
    settings: dict[str, SuggestedSetting] = Field(
        ..., min_length=1, description="The rule settings the suggestion adds."
    )
    differing: int = Field(..., ge=1, description="Rows that differ in the column today.")
    explained: int = Field(..., ge=1, description="Of those, rows the rule makes match.")
    largest_gap: float | None = Field(
        default=None, ge=0, description="For a tolerance, the largest gap it explains."
    )
    examples: tuple[dict[str, Any], ...] = Field(
        ..., description="Primary keys of up to three explained rows, lowest first."
    )
    rule: DiffRule = Field(..., description="The rule to add, first in rules.")
    governing_rule_index: int | None = Field(
        default=None, description="Index of the rule that governs the column today."
    )

SnowflakeConfig

Bases: BaseModel

Immutable connection settings for Snowflake warehouse pushdown.

Attributes:

Name Type Description
type Literal['snowflake']

Source kind, which selects this model.

table str

Fully qualified table or view to compare.

account str

Snowflake account identifier.

user str

Login name used to authenticate the session.

warehouse str

Virtual warehouse that executes pushdown SQL.

database str

Default database for unqualified object names.

schema_name str

Default schema for unqualified object names.

password str | None

Optional password or programmatic access token, unset for a key pair or SSO. Left out when the config is printed, but kept by model_dump(), which the connector needs.

private_key_path str | None

Optional path to a PEM private key, for key-pair sign-in instead of a password. Left out when printed.

private_key_passphrase str | None

Passphrase of an encrypted private_key_path. Left out when printed.

role str | None

Optional role assumed after authentication.

Source code in src/veridelta/models.py
class SnowflakeConfig(BaseModel):
    """Immutable connection settings for Snowflake warehouse pushdown.

    Attributes:
        type (Literal["snowflake"]): Source kind, which selects this model.
        table (str): Fully qualified table or view to compare.
        account (str): Snowflake account identifier.
        user (str): Login name used to authenticate the session.
        warehouse (str): Virtual warehouse that executes pushdown SQL.
        database (str): Default database for unqualified object names.
        schema_name (str): Default schema for unqualified object names.
        password (str | None): Optional password or programmatic access
            token, unset for a key pair or SSO. Left out when the config is
            printed, but kept by `model_dump()`, which the connector needs.
        private_key_path (str | None): Optional path to a PEM private key, for
            key-pair sign-in instead of a password. Left out when printed.
        private_key_passphrase (str | None): Passphrase of an encrypted
            `private_key_path`. Left out when printed.
        role (str | None): Optional role assumed after authentication.
    """

    # Credentials pass through here, and Pydantic quotes raw input in its errors.
    model_config = ConfigDict(extra="forbid", frozen=True, hide_input_in_errors=True)

    type: Literal["snowflake"] = Field("snowflake", description="Discriminator for Snowflake.")
    table: str = Field(
        ...,
        pattern=SQL_RELATION_PATTERN,
        description="Fully qualified table or view to compare.",
    )
    account: str = Field(..., description="Snowflake account identifier.")
    user: str = Field(..., description="Login name used to authenticate the session.")
    warehouse: str = Field(..., description="Virtual warehouse that executes pushdown SQL.")
    database: str = Field(..., description="Default database for unqualified object names.")
    schema_name: str = Field(..., description="Default schema for unqualified object names.")
    password: str | None = Field(
        default=None,
        repr=False,
        description="Password or programmatic access token; omitted for a key pair or SSO.",
    )
    private_key_path: str | None = Field(
        default=None,
        repr=False,
        description="Path to a PEM private key, for key-pair sign-in instead of a password.",
    )
    private_key_passphrase: str | None = Field(
        default=None, repr=False, description="Passphrase of an encrypted private_key_path."
    )
    role: str | None = Field(
        default=None, description="Optional role assumed after authentication."
    )

    @model_validator(mode="after")
    def validate_sign_in(self) -> "SnowflakeConfig":
        """Reject credentials that leave unclear how the session signs in.

        Returns:
            SnowflakeConfig: The validated instance.

        Raises:
            ValueError: If both `password` and `private_key_path` are set, or
                `private_key_passphrase` is set without `private_key_path`.
        """
        if self.password is not None and self.private_key_path is not None:
            raise ValueError(
                "Snowflake signs in with a 'password' or a 'private_key_path', not both."
            )
        if self.private_key_passphrase is not None and self.private_key_path is None:
            raise ValueError(
                "'private_key_passphrase' decrypts 'private_key_path', which is not set."
            )
        return self

validate_sign_in()

Reject credentials that leave unclear how the session signs in.

Returns:

Name Type Description
SnowflakeConfig SnowflakeConfig

The validated instance.

Raises:

Type Description
ValueError

If both password and private_key_path are set, or private_key_passphrase is set without private_key_path.

Source code in src/veridelta/models.py
@model_validator(mode="after")
def validate_sign_in(self) -> "SnowflakeConfig":
    """Reject credentials that leave unclear how the session signs in.

    Returns:
        SnowflakeConfig: The validated instance.

    Raises:
        ValueError: If both `password` and `private_key_path` are set, or
            `private_key_passphrase` is set without `private_key_path`.
    """
    if self.password is not None and self.private_key_path is not None:
        raise ValueError(
            "Snowflake signs in with a 'password' or a 'private_key_path', not both."
        )
    if self.private_key_passphrase is not None and self.private_key_path is None:
        raise ValueError(
            "'private_key_passphrase' decrypts 'private_key_path', which is not set."
        )
    return self

SourceConfig

Bases: BaseModel

Settings for a file source.

Attributes:

Name Type Description
type Literal['file']

Source kind. A YAML file source may omit it.

path str

Local path or URI of the file.

format SourceType

File format, such as parquet. When absent, the path's suffix decides it, and csv is the default for a suffix Veridelta does not know.

options dict[str, Any]

Keyword arguments for the Polars reader, such as {'separator': ';'}. A nested storage_options map is left out when the config is printed, but kept by model_dump(), which the reader needs.

Source code in src/veridelta/models.py
class SourceConfig(BaseModel):
    """Settings for a file source.

    Attributes:
        type (Literal["file"]): Source kind. A YAML file source may omit it.
        path (str): Local path or URI of the file.
        format (SourceType): File format, such as `parquet`. When absent, the path's
            suffix decides it, and `csv` is the default for a suffix Veridelta does
            not know.
        options (dict[str, Any]): Keyword arguments for the Polars reader, such as
            `{'separator': ';'}`. A nested `storage_options` map is left out when the
            config is printed, but kept by `model_dump()`, which the reader needs.
    """

    # Reader options can carry object-store credentials, which Pydantic would
    # otherwise quote in its errors.
    model_config = ConfigDict(extra="forbid", hide_input_in_errors=True)

    type: Literal["file"] = Field("file", description="Discriminator for file-backed sources.")
    path: str = Field(..., description="File system path or URI to the data.")
    format: SourceType = Field(
        "csv",
        description=(
            "The format of the file. When absent, the path's suffix decides it, and csv "
            "is the default for a suffix Veridelta does not know."
        ),
    )
    options: dict[str, Any] = Field(
        default_factory=dict,
        description="Options for the Polars reader, such as {'separator': ';'}.",
    )

    @model_validator(mode="before")
    @classmethod
    def _format_from_suffix(cls, data: Any) -> Any:
        """Fill an absent `format` from the path's suffix.

        A `format` the user wrote always wins. A suffix Veridelta does not know
        leaves the key absent, so the default applies and the error for a
        missing primary key can say that `format` is not set.
        """
        if not isinstance(data, dict):
            return data
        values = cast("dict[str, Any]", data)
        if "format" in values:
            return values
        inferred = _infer_format(str(values.get("path", "")))
        return values if inferred is None else {**values, "format": inferred}

    def __repr_args__(self) -> Iterable[tuple[str | None, Any]]:
        """Leave object-store credentials out of the printed reader options.

        A cloud path's credentials travel in a nested `storage_options` map.
        Printing drops that one key and keeps the other options, which are the
        useful part when debugging. `repr()`, `str()`, and rich displays all
        read from here, while `model_dump()` and the reader get the full map.

        Yields:
            tuple[str | None, Any]: Each field name with the value to print.
        """
        for name, value in super().__repr_args__():
            if name == "options" and "storage_options" in self.options:
                yield (
                    name,
                    {key: item for key, item in self.options.items() if key != "storage_options"},
                )
            else:
                yield name, value

__repr_args__()

Leave object-store credentials out of the printed reader options.

A cloud path's credentials travel in a nested storage_options map. Printing drops that one key and keeps the other options, which are the useful part when debugging. repr(), str(), and rich displays all read from here, while model_dump() and the reader get the full map.

Yields:

Type Description
Iterable[tuple[str | None, Any]]

tuple[str | None, Any]: Each field name with the value to print.

Source code in src/veridelta/models.py
def __repr_args__(self) -> Iterable[tuple[str | None, Any]]:
    """Leave object-store credentials out of the printed reader options.

    A cloud path's credentials travel in a nested `storage_options` map.
    Printing drops that one key and keeps the other options, which are the
    useful part when debugging. `repr()`, `str()`, and rich displays all
    read from here, while `model_dump()` and the reader get the full map.

    Yields:
        tuple[str | None, Any]: Each field name with the value to print.
    """
    for name, value in super().__repr_args__():
        if name == "options" and "storage_options" in self.options:
            yield (
                name,
                {key: item for key, item in self.options.items() if key != "storage_options"},
            )
        else:
            yield name, value

ValueMapEntry

Bases: BaseModel

One proposed value_map entry and the rows that support it.

Attributes:

Name Type Description
source_value str

Source text as the value_map stage sees it, after null sentinels, regex replacement, whitespace, and case folding.

target_value str

The target value those rows compare against, as text.

rows int

Joined rows whose source holds source_value, whatever their target, NULL included.

agreeing_rows int

Those rows whose target is target_value.

confidence float

agreeing_rows / rows, the share of the source value's rows the entry would make match.

Source code in src/veridelta/models.py
class ValueMapEntry(BaseModel):
    """One proposed `value_map` entry and the rows that support it.

    Attributes:
        source_value (str): Source text as the `value_map` stage sees it, after
            null sentinels, regex replacement, whitespace, and case folding.
        target_value (str): The target value those rows compare against, as text.
        rows (int): Joined rows whose source holds `source_value`, whatever
            their target, NULL included.
        agreeing_rows (int): Those rows whose target is `target_value`.
        confidence (float): `agreeing_rows / rows`, the share of the source
            value's rows the entry would make match.
    """

    model_config = ConfigDict(extra="forbid", frozen=True)

    source_value: str = Field(..., description="Source text as the value_map stage sees it.")
    target_value: str = Field(..., description="Target value the rows compare against.")
    rows: int = Field(..., ge=1, description="Joined rows holding the source value.")
    agreeing_rows: int = Field(
        ..., ge=1, description="Rows among them whose target is the target value."
    )

    @computed_field
    @property
    def confidence(self) -> float:
        """Return the share of the source value's rows that agree.

        Returns:
            float: `agreeing_rows / rows`.
        """
        return self.agreeing_rows / self.rows

    @model_validator(mode="after")
    def validate_counts(self) -> "ValueMapEntry":
        """Reject more agreeing rows than rows.

        Returns:
            ValueMapEntry: The validated entry.

        Raises:
            ValueError: If `agreeing_rows` exceeds `rows`.
        """
        if self.agreeing_rows > self.rows:
            raise ValueError(
                f"agreeing_rows ({self.agreeing_rows}) cannot exceed rows ({self.rows})."
            )
        return self

confidence property

Return the share of the source value's rows that agree.

Returns:

Name Type Description
float float

agreeing_rows / rows.

validate_counts()

Reject more agreeing rows than rows.

Returns:

Name Type Description
ValueMapEntry ValueMapEntry

The validated entry.

Raises:

Type Description
ValueError

If agreeing_rows exceeds rows.

Source code in src/veridelta/models.py
@model_validator(mode="after")
def validate_counts(self) -> "ValueMapEntry":
    """Reject more agreeing rows than rows.

    Returns:
        ValueMapEntry: The validated entry.

    Raises:
        ValueError: If `agreeing_rows` exceeds `rows`.
    """
    if self.agreeing_rows > self.rows:
        raise ValueError(
            f"agreeing_rows ({self.agreeing_rows}) cannot exceed rows ({self.rows})."
        )
    return self

ValueMapProposal

Bases: BaseModel

A proposed value_map for one column, with the evidence for each new entry.

Attributes:

Name Type Description
column str

Compared column, named as it appears after any rename_to.

value_map dict[str, str]

The governing rule's existing entries plus the proposed ones.

entries tuple[ValueMapEntry, ...]

The proposed entries alone, most agreeing rows first.

governing_rule_index int | None

Position in DiffConfig.rules of the rule that governs the column today, or None when no rule does. Only one rule governs a column, so new entries belong in that rule.

Source code in src/veridelta/models.py
class ValueMapProposal(BaseModel):
    """A proposed `value_map` for one column, with the evidence for each new entry.

    Attributes:
        column (str): Compared column, named as it appears after any `rename_to`.
        value_map (dict[str, str]): The governing rule's existing entries plus
            the proposed ones.
        entries (tuple[ValueMapEntry, ...]): The proposed entries alone, most
            agreeing rows first.
        governing_rule_index (int | None): Position in `DiffConfig.rules` of the
            rule that governs the column today, or None when no rule does. Only
            one rule governs a column, so new entries belong in that rule.
    """

    model_config = ConfigDict(extra="forbid", frozen=True)

    column: str = Field(..., description="Compared column, after any rename_to.")
    value_map: dict[str, str] = Field(..., description="Existing plus proposed entries.")
    entries: tuple[ValueMapEntry, ...] = Field(..., description="Proposed entries alone.")
    governing_rule_index: int | None = Field(
        default=None, description="Index of the rule that governs the column today."
    )

    def to_rule(self) -> DiffRule:
        """Build a rule for the column carrying the proposed map.

        When `governing_rule_index` is set, merge `value_map` into that rule
        instead, since a second rule for the column would not apply.

        Returns:
            DiffRule: Rule naming the column, with the full proposed map.
        """
        return DiffRule(column_names=[self.column], value_map=self.value_map)

to_rule()

Build a rule for the column carrying the proposed map.

When governing_rule_index is set, merge value_map into that rule instead, since a second rule for the column would not apply.

Returns:

Name Type Description
DiffRule DiffRule

Rule naming the column, with the full proposed map.

Source code in src/veridelta/models.py
def to_rule(self) -> DiffRule:
    """Build a rule for the column carrying the proposed map.

    When `governing_rule_index` is set, merge `value_map` into that rule
    instead, since a second rule for the column would not apply.

    Returns:
        DiffRule: Rule naming the column, with the full proposed map.
    """
    return DiffRule(column_names=[self.column], value_map=self.value_map)

mismatch_ratio_of(mismatches, source_rows)

Divide the mismatches by the source row count, which the threshold applies to.

Parameters:

Name Type Description Default
mismatches int

Added, removed, and changed rows together.

required
source_rows int

The rows the source holds. An empty source counts as one.

required

Returns:

Name Type Description
float float

The ratio, which can exceed 1.

Source code in src/veridelta/models.py
def mismatch_ratio_of(mismatches: int, source_rows: int) -> float:
    """Divide the mismatches by the source row count, which the threshold applies to.

    Args:
        mismatches (int): Added, removed, and changed rows together.
        source_rows (int): The rows the source holds. An empty source counts as one.

    Returns:
        float: The ratio, which can exceed 1.
    """
    return float(mismatches) / float(max(source_rows, 1))

normalize_column_name(name)

Strip and lowercase a column name, as normalize_column_names asks.

Parameters:

Name Type Description Default
name str

A header, or a name a configuration gives.

required

Returns:

Name Type Description
str str

The name the engine compares by.

Source code in src/veridelta/models.py
def normalize_column_name(name: str) -> str:
    """Strip and lowercase a column name, as `normalize_column_names` asks.

    Args:
        name (str): A header, or a name a configuration gives.

    Returns:
        str: The name the engine compares by.
    """
    return name.strip().lower()

redacted_location(location)

Return a path or URL with anything secret left out, safe to print or log.

A URL keeps its scheme, host, and path. Its query goes, since it can hold a token or the signature of a pre-signed link, and so does its user part, which can hold a login, unless an Azure scheme names a container there. A path on this machine comes back as it is.

Parameters:

Name Type Description Default
location str

A file path, or the URL of a file or a table.

required

Returns:

Type Description
str | None

str | None: The location without its secrets, or None when it does not parse as a URL.

Examples:

>>> redacted_location("s3://bucket/events.parquet?X-Amz-Signature=abc")
's3://bucket/events.parquet'
>>> redacted_location("abfss://lake@account.dfs.core.windows.net/events")
'abfss://lake@account.dfs.core.windows.net/events'
Source code in src/veridelta/models.py
def redacted_location(location: str) -> str | None:
    """Return a path or URL with anything secret left out, safe to print or log.

    A URL keeps its scheme, host, and path. Its query goes, since it can hold a
    token or the signature of a pre-signed link, and so does its user part,
    which can hold a login, unless an Azure scheme names a container there. A
    path on this machine comes back as it is.

    Args:
        location (str): A file path, or the URL of a file or a table.

    Returns:
        str | None: The location without its secrets, or None when it does not
            parse as a URL.

    Examples:
        >>> redacted_location("s3://bucket/events.parquet?X-Amz-Signature=abc")
        's3://bucket/events.parquet'
        >>> redacted_location("abfss://lake@account.dfs.core.windows.net/events")
        'abfss://lake@account.dfs.core.windows.net/events'
    """
    try:
        parts = urlsplit(location)
    except ValueError:
        return None
    if not (parts.scheme and parts.netloc):
        return location
    user, _, host = parts.netloc.rpartition("@")
    if user and parts.scheme in _CONTAINER_SCHEMES and ":" not in user:
        host = f"{user}@{host}"
    return urlunsplit((parts.scheme, host, parts.path, "", ""))

Engine

DiffEngine loads, aligns, and compares two datasets, on Polars or inside a warehouse.

Load, align, and compare datasets.

Holds the DiffEngine that loads, aligns, and compares the two sides with Polars. Its helpers live in private modules beside it, and __all__ lists the public names, the file loaders from _reading.py among them.

DEFAULT_MAX_SHARE = 0.01 module-attribute

Largest gap a suggested tolerance may explain, as a share of the larger of its two values. A gap past it is a change, not noise, so the column gets no suggestion.

DEFAULT_MIN_CONFIDENCE = 0.95 module-attribute

Share of a source value's rows that must agree on one target value before it is proposed. Above one half, at most one target can qualify, and 5% leaves room for noise in legacy data.

DEFAULT_MIN_SUPPORT = 5 module-attribute

Agreeing rows a proposal needs, so a one-off coincidence is never offered.

ArrowLoader

Bases: BaseLoader

Streaming loader for Arrow IPC (Feather v2) files over pl.scan_ipc.

IPC files carry their schema, so no type inference runs and the dtypes the engine compares are exactly the ones the writer stored.

Source code in src/veridelta/_reading.py
class ArrowLoader(BaseLoader):
    """Streaming loader for Arrow IPC (Feather v2) files over `pl.scan_ipc`.

    IPC files carry their schema, so no type inference runs and the dtypes the
    engine compares are exactly the ones the writer stored.
    """

    def load(self, config: SourceConfig) -> pl.LazyFrame:
        """Scan an Arrow IPC file.

        Args:
            config (SourceConfig): Source configuration, whose options go to `pl.scan_ipc`.

        Returns:
            pl.LazyFrame: The unevaluated rows.
        """
        return pl.scan_ipc(config.path, **config.options)

load(config)

Scan an Arrow IPC file.

Parameters:

Name Type Description Default
config SourceConfig

Source configuration, whose options go to pl.scan_ipc.

required

Returns:

Type Description
LazyFrame

pl.LazyFrame: The unevaluated rows.

Source code in src/veridelta/_reading.py
def load(self, config: SourceConfig) -> pl.LazyFrame:
    """Scan an Arrow IPC file.

    Args:
        config (SourceConfig): Source configuration, whose options go to `pl.scan_ipc`.

    Returns:
        pl.LazyFrame: The unevaluated rows.
    """
    return pl.scan_ipc(config.path, **config.options)

AvroLoader

Bases: BaseLoader

Loader for Avro object container files, read eagerly with pl.read_avro.

Polars has no lazy Avro reader, so the file is read whole and wrapped, like JSON and Excel. Avro carries its schema, so the dtypes compared are the writer's, with no inference. The reader takes a local path only.

Source code in src/veridelta/_reading.py
class AvroLoader(BaseLoader):
    """Loader for Avro object container files, read eagerly with `pl.read_avro`.

    Polars has no lazy Avro reader, so the file is read whole and wrapped, like
    JSON and Excel. Avro carries its schema, so the dtypes compared are the
    writer's, with no inference. The reader takes a local path only.
    """

    def load(self, config: SourceConfig) -> pl.LazyFrame:
        """Read an Avro file.

        Args:
            config (SourceConfig): Source configuration, whose `columns` and `n_rows`
                options go to `pl.read_avro`.

        Returns:
            pl.LazyFrame: A lazy wrapper over the rows read.
        """
        return pl.read_avro(config.path, **config.options).lazy()

load(config)

Read an Avro file.

Parameters:

Name Type Description Default
config SourceConfig

Source configuration, whose columns and n_rows options go to pl.read_avro.

required

Returns:

Type Description
LazyFrame

pl.LazyFrame: A lazy wrapper over the rows read.

Source code in src/veridelta/_reading.py
def load(self, config: SourceConfig) -> pl.LazyFrame:
    """Read an Avro file.

    Args:
        config (SourceConfig): Source configuration, whose `columns` and `n_rows`
            options go to `pl.read_avro`.

    Returns:
        pl.LazyFrame: A lazy wrapper over the rows read.
    """
    return pl.read_avro(config.path, **config.options).lazy()

BaseLoader

Bases: ABC

Base class for the file loaders, each turning a SourceConfig into a LazyFrame.

Each file format has a subclass registered in LoaderFactory._loaders, keyed by its SourceType. A loader prefers a Polars scan_* reader, so the comparison stays lazy end to end; the eager loaders say why in their own docstrings. SourceConfig.options reach the reader unchanged, so any keyword the Polars function accepts is valid.

Source code in src/veridelta/_reading.py
class BaseLoader(ABC):
    """Base class for the file loaders, each turning a `SourceConfig` into a LazyFrame.

    Each file format has a subclass registered in `LoaderFactory._loaders`, keyed by
    its `SourceType`. A loader prefers a Polars `scan_*` reader, so the comparison
    stays lazy end to end; the eager loaders say why in their own docstrings.
    `SourceConfig.options` reach the reader unchanged, so any keyword the Polars
    function accepts is valid.
    """

    @abstractmethod
    def load(self, config: SourceConfig) -> pl.LazyFrame:
        """Load a source into a LazyFrame.

        Args:
            config (SourceConfig): Path, format, and reader options.

        Returns:
            pl.LazyFrame: The unevaluated rows.
        """

load(config) abstractmethod

Load a source into a LazyFrame.

Parameters:

Name Type Description Default
config SourceConfig

Path, format, and reader options.

required

Returns:

Type Description
LazyFrame

pl.LazyFrame: The unevaluated rows.

Source code in src/veridelta/_reading.py
@abstractmethod
def load(self, config: SourceConfig) -> pl.LazyFrame:
    """Load a source into a LazyFrame.

    Args:
        config (SourceConfig): Path, format, and reader options.

    Returns:
        pl.LazyFrame: The unevaluated rows.
    """

CSVLoader

Bases: BaseLoader

Streaming CSV loader over pl.scan_csv.

Delimiters, encodings, and header handling are all controlled through SourceConfig.options, for example {"separator": ";"}.

Source code in src/veridelta/_reading.py
class CSVLoader(BaseLoader):
    """Streaming CSV loader over `pl.scan_csv`.

    Delimiters, encodings, and header handling are all controlled through
    `SourceConfig.options`, for example `{"separator": ";"}`.
    """

    def load(self, config: SourceConfig) -> pl.LazyFrame:
        """Scan a CSV file.

        Args:
            config (SourceConfig): Source configuration, whose options go to `pl.scan_csv`.

        Returns:
            pl.LazyFrame: The unevaluated rows.
        """
        return pl.scan_csv(config.path, **config.options)

load(config)

Scan a CSV file.

Parameters:

Name Type Description Default
config SourceConfig

Source configuration, whose options go to pl.scan_csv.

required

Returns:

Type Description
LazyFrame

pl.LazyFrame: The unevaluated rows.

Source code in src/veridelta/_reading.py
def load(self, config: SourceConfig) -> pl.LazyFrame:
    """Scan a CSV file.

    Args:
        config (SourceConfig): Source configuration, whose options go to `pl.scan_csv`.

    Returns:
        pl.LazyFrame: The unevaluated rows.
    """
    return pl.scan_csv(config.path, **config.options)

DiffEngine

Compare two datasets on their primary keys and report what differs.

The engine takes two Polars LazyFrame inputs and a DiffConfig, applies the nine-stage DiffRule pipeline to each side, then joins on the primary keys to classify rows as added (target-only), removed (source-only), or changed (present on both sides with at least one compared column differing). Nothing is materialized until run() collects the joins, so the source frames can be scan_* graphs over files far larger than memory.

These entry points cover the usual situations:

  • DiffEngine(config, source, target).run() for frames you already hold. In-memory DataFrame inputs must be wrapped with .lazy() first.
  • DiffEngine.run_from_configs(diff, source, target) for SourceRef pairs from YAML. File, lakehouse, database, and DuckDB pairs load through LoaderFactory and run locally. A same-warehouse pair, or a Postgres or DuckDB pair that sets pushdown, compiles to SQL and runs where it is stored.
  • DiffEngine.validate_schemas(...) to enforce schema_mode and primary key presence on metadata alone, before any rows are read.
  • DiffEngine.validate_rules(...) to also resolve every rule and build each column's comparison, still on metadata alone.
  • DiffEngine(config, source, target).propose_value_maps(), or DiffEngine.propose_value_maps_from_configs(...) for any SourceRef pair, counted in the warehouse for same-warehouse pairs, to suggest value_map entries from how the two sides' values line up, without running the comparison.

Attributes:

Name Type Description
config DiffConfig

Primary keys, rules, defaults, and threshold.

source LazyFrame

Source side; mutated in place as alignment and normalization stages run.

target LazyFrame

Target side, treated as the authoritative schema.

Raises:

Type Description
ConfigError

From run() when primary keys are missing, schema_mode is violated, or a rule cannot apply to the column it names.

DataIntegrityError

From run() when primary keys are not unique on either side after normalization.

Source code in src/veridelta/engine.py
 134
 135
 136
 137
 138
 139
 140
 141
 142
 143
 144
 145
 146
 147
 148
 149
 150
 151
 152
 153
 154
 155
 156
 157
 158
 159
 160
 161
 162
 163
 164
 165
 166
 167
 168
 169
 170
 171
 172
 173
 174
 175
 176
 177
 178
 179
 180
 181
 182
 183
 184
 185
 186
 187
 188
 189
 190
 191
 192
 193
 194
 195
 196
 197
 198
 199
 200
 201
 202
 203
 204
 205
 206
 207
 208
 209
 210
 211
 212
 213
 214
 215
 216
 217
 218
 219
 220
 221
 222
 223
 224
 225
 226
 227
 228
 229
 230
 231
 232
 233
 234
 235
 236
 237
 238
 239
 240
 241
 242
 243
 244
 245
 246
 247
 248
 249
 250
 251
 252
 253
 254
 255
 256
 257
 258
 259
 260
 261
 262
 263
 264
 265
 266
 267
 268
 269
 270
 271
 272
 273
 274
 275
 276
 277
 278
 279
 280
 281
 282
 283
 284
 285
 286
 287
 288
 289
 290
 291
 292
 293
 294
 295
 296
 297
 298
 299
 300
 301
 302
 303
 304
 305
 306
 307
 308
 309
 310
 311
 312
 313
 314
 315
 316
 317
 318
 319
 320
 321
 322
 323
 324
 325
 326
 327
 328
 329
 330
 331
 332
 333
 334
 335
 336
 337
 338
 339
 340
 341
 342
 343
 344
 345
 346
 347
 348
 349
 350
 351
 352
 353
 354
 355
 356
 357
 358
 359
 360
 361
 362
 363
 364
 365
 366
 367
 368
 369
 370
 371
 372
 373
 374
 375
 376
 377
 378
 379
 380
 381
 382
 383
 384
 385
 386
 387
 388
 389
 390
 391
 392
 393
 394
 395
 396
 397
 398
 399
 400
 401
 402
 403
 404
 405
 406
 407
 408
 409
 410
 411
 412
 413
 414
 415
 416
 417
 418
 419
 420
 421
 422
 423
 424
 425
 426
 427
 428
 429
 430
 431
 432
 433
 434
 435
 436
 437
 438
 439
 440
 441
 442
 443
 444
 445
 446
 447
 448
 449
 450
 451
 452
 453
 454
 455
 456
 457
 458
 459
 460
 461
 462
 463
 464
 465
 466
 467
 468
 469
 470
 471
 472
 473
 474
 475
 476
 477
 478
 479
 480
 481
 482
 483
 484
 485
 486
 487
 488
 489
 490
 491
 492
 493
 494
 495
 496
 497
 498
 499
 500
 501
 502
 503
 504
 505
 506
 507
 508
 509
 510
 511
 512
 513
 514
 515
 516
 517
 518
 519
 520
 521
 522
 523
 524
 525
 526
 527
 528
 529
 530
 531
 532
 533
 534
 535
 536
 537
 538
 539
 540
 541
 542
 543
 544
 545
 546
 547
 548
 549
 550
 551
 552
 553
 554
 555
 556
 557
 558
 559
 560
 561
 562
 563
 564
 565
 566
 567
 568
 569
 570
 571
 572
 573
 574
 575
 576
 577
 578
 579
 580
 581
 582
 583
 584
 585
 586
 587
 588
 589
 590
 591
 592
 593
 594
 595
 596
 597
 598
 599
 600
 601
 602
 603
 604
 605
 606
 607
 608
 609
 610
 611
 612
 613
 614
 615
 616
 617
 618
 619
 620
 621
 622
 623
 624
 625
 626
 627
 628
 629
 630
 631
 632
 633
 634
 635
 636
 637
 638
 639
 640
 641
 642
 643
 644
 645
 646
 647
 648
 649
 650
 651
 652
 653
 654
 655
 656
 657
 658
 659
 660
 661
 662
 663
 664
 665
 666
 667
 668
 669
 670
 671
 672
 673
 674
 675
 676
 677
 678
 679
 680
 681
 682
 683
 684
 685
 686
 687
 688
 689
 690
 691
 692
 693
 694
 695
 696
 697
 698
 699
 700
 701
 702
 703
 704
 705
 706
 707
 708
 709
 710
 711
 712
 713
 714
 715
 716
 717
 718
 719
 720
 721
 722
 723
 724
 725
 726
 727
 728
 729
 730
 731
 732
 733
 734
 735
 736
 737
 738
 739
 740
 741
 742
 743
 744
 745
 746
 747
 748
 749
 750
 751
 752
 753
 754
 755
 756
 757
 758
 759
 760
 761
 762
 763
 764
 765
 766
 767
 768
 769
 770
 771
 772
 773
 774
 775
 776
 777
 778
 779
 780
 781
 782
 783
 784
 785
 786
 787
 788
 789
 790
 791
 792
 793
 794
 795
 796
 797
 798
 799
 800
 801
 802
 803
 804
 805
 806
 807
 808
 809
 810
 811
 812
 813
 814
 815
 816
 817
 818
 819
 820
 821
 822
 823
 824
 825
 826
 827
 828
 829
 830
 831
 832
 833
 834
 835
 836
 837
 838
 839
 840
 841
 842
 843
 844
 845
 846
 847
 848
 849
 850
 851
 852
 853
 854
 855
 856
 857
 858
 859
 860
 861
 862
 863
 864
 865
 866
 867
 868
 869
 870
 871
 872
 873
 874
 875
 876
 877
 878
 879
 880
 881
 882
 883
 884
 885
 886
 887
 888
 889
 890
 891
 892
 893
 894
 895
 896
 897
 898
 899
 900
 901
 902
 903
 904
 905
 906
 907
 908
 909
 910
 911
 912
 913
 914
 915
 916
 917
 918
 919
 920
 921
 922
 923
 924
 925
 926
 927
 928
 929
 930
 931
 932
 933
 934
 935
 936
 937
 938
 939
 940
 941
 942
 943
 944
 945
 946
 947
 948
 949
 950
 951
 952
 953
 954
 955
 956
 957
 958
 959
 960
 961
 962
 963
 964
 965
 966
 967
 968
 969
 970
 971
 972
 973
 974
 975
 976
 977
 978
 979
 980
 981
 982
 983
 984
 985
 986
 987
 988
 989
 990
 991
 992
 993
 994
 995
 996
 997
 998
 999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
1030
1031
1032
1033
1034
1035
1036
1037
1038
1039
1040
1041
1042
1043
1044
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
1075
1076
1077
1078
1079
1080
1081
1082
1083
1084
1085
1086
1087
1088
1089
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
1100
1101
1102
class DiffEngine:
    """Compare two datasets on their primary keys and report what differs.

    The engine takes two Polars `LazyFrame` inputs and a `DiffConfig`, applies
    the nine-stage `DiffRule` pipeline to each side, then joins on the primary
    keys to classify rows as added (target-only), removed (source-only), or
    changed (present on both sides with at least one compared column
    differing). Nothing is materialized until `run()` collects the joins, so
    the source frames can be `scan_*` graphs over files far larger than memory.

    These entry points cover the usual situations:

    - `DiffEngine(config, source, target).run()` for frames you already hold.
      In-memory `DataFrame` inputs must be wrapped with `.lazy()` first.
    - `DiffEngine.run_from_configs(diff, source, target)` for `SourceRef`
      pairs from YAML. File, lakehouse, database, and DuckDB pairs load
      through `LoaderFactory` and run locally. A same-warehouse pair, or a
      Postgres or DuckDB pair that sets `pushdown`, compiles to SQL and runs
      where it is stored.
    - `DiffEngine.validate_schemas(...)` to enforce `schema_mode` and primary
      key presence on metadata alone, before any rows are read.
    - `DiffEngine.validate_rules(...)` to also resolve every rule and build
      each column's comparison, still on metadata alone.
    - `DiffEngine(config, source, target).propose_value_maps()`, or
      `DiffEngine.propose_value_maps_from_configs(...)` for any `SourceRef`
      pair, counted in the warehouse for same-warehouse pairs, to suggest
      `value_map` entries from how the two sides' values line up, without
      running the comparison.

    Attributes:
        config (DiffConfig): Primary keys, rules, defaults, and threshold.
        source (pl.LazyFrame): Source side; mutated in place as alignment and
            normalization stages run.
        target (pl.LazyFrame): Target side, treated as the authoritative schema.

    Raises:
        ConfigError: From `run()` when primary keys are missing, `schema_mode`
            is violated, or a rule cannot apply to the column it names.
        DataIntegrityError: From `run()` when primary keys are not unique on
            either side after normalization.
    """

    def __init__(
        self, config: DiffConfig, source_df: pl.LazyFrame, target_df: pl.LazyFrame
    ) -> None:
        """Hold the configuration and the two datasets to compare.

        `run()` aligns the datasets and compares them.

        Args:
            config (DiffConfig): Keys, rules, and settings for the comparison.
            source_df (pl.LazyFrame): Source rows, as stored.
            target_df (pl.LazyFrame): Target rows, as stored.
        """
        self.config = config
        self.source = source_df
        self.target = target_df
        # How each side was read, set by the entry points that load a `SourceRef`,
        # so an error can name the file and its format rather than only a column.
        self._sides: dict[str, str] | None = None

    @classmethod
    def _on_sources(cls, diff: DiffConfig, source: SourceRef, target: SourceRef) -> "DiffEngine":
        """Load both sides and build an engine whose errors name them."""
        engine = cls(diff, LoaderFactory.load(source), LoaderFactory.load(target))
        engine._sides = describe_sides(source, target)
        return engine

    @classmethod
    def run_from_configs(
        cls,
        diff: DiffConfig,
        source: SourceRef,
        target: SourceRef,
        *,
        baseline: Baseline | None = None,
    ) -> DiffResult:
        """Route a comparison to pushdown or local Polars evaluation.

        Args:
            diff (DiffConfig): Comparison settings and rules.
            source (SourceRef): Source file, lakehouse, database, DuckDB, or warehouse config.
            target (SourceRef): Target file, lakehouse, database, DuckDB, or warehouse config.
            baseline (Baseline | None): Drift to accept, which a local run leaves out
                of the counts and the verdict.

        Returns:
            DiffResult: The result. A pushdown pair returns counts and keys, and a
                local pair also returns the differing rows.

        Raises:
            ConfigError: If primary keys are missing, `schema_mode` is violated, or a
                baseline is given for a pair compared where it is stored.
            DataIntegrityError: If either dataset repeats a normalized primary key.
            ConnectorError: If warehouse backends are mixed or connections differ.

        Examples:
            >>> from veridelta.models import DiffConfig, SourceConfig
            >>> diff = DiffConfig(primary_keys=["order_id"])
            >>> source = SourceConfig(path="legacy/orders.parquet", format="parquet")
            >>> target = SourceConfig(path="modern/orders.parquet", format="parquet")
            >>> result = DiffEngine.run_from_configs(diff, source, target)  # doctest: +SKIP
        """
        pair = check_backend_pairing(source, target)
        if pair is not None:
            if baseline is not None:
                raise ConfigError(
                    "A baseline applies to a run that compares both sides locally, and this "
                    "pair is compared where it is stored. Leave out --baseline, or compare "
                    "files exported from it."
                )
            return pair.with_session(
                lambda session, source_table, target_table: collect_pushdown_summary(
                    session, source_table, target_table, diff
                )
            )

        # `run()` normalizes headers and applies renames exactly once, so the
        # frames go to it straight from the loaders: aligning them first and
        # renaming again would undo a swap and collapse a chain.
        return cls._on_sources(diff, source, target).run(baseline=baseline)

    @classmethod
    def validate_schemas(
        cls, config: DiffConfig, source_df: pl.LazyFrame, target_df: pl.LazyFrame
    ) -> None:
        """Align structure and enforce `SchemaMode` without comparing any rows.

        Operates on schema metadata only, so callers may pass zero-row frames.

        Args:
            config (DiffConfig): Comparison settings and rules.
            source_df (pl.LazyFrame): Source frame or column probe.
            target_df (pl.LazyFrame): Target frame or column probe.

        Raises:
            ConfigError: If primary keys are missing or schema constraints are violated.
        """
        engine = cls(config, source_df, target_df)
        engine._align_structure()
        engine._validate_schema()

    @classmethod
    def validate_rules(
        cls, config: DiffConfig, source_df: pl.LazyFrame, target_df: pl.LazyFrame
    ) -> list[str]:
        """Check everything a run checks before it reads a row.

        Goes past `validate_schemas`: every rule is resolved against the aligned
        columns, both schemas are normalized, and each column's comparison is built. A
        rule the run could not honor fails here, such as a null sentinel its column's
        type cannot hold, or a similarity limit without the `fuzzy` extra. Operates on
        schema metadata only, so callers may pass zero-row frames. Repeated keys and
        invalid regular expressions surface only when rows are read.

        Args:
            config (DiffConfig): Comparison settings and rules.
            source_df (pl.LazyFrame): Source frame or column probe.
            target_df (pl.LazyFrame): Target frame or column probe.

        Returns:
            list[str]: The columns a run would compare, in source order, under their
                target names.

        Raises:
            ConfigError: If primary keys are missing or hold two types a join
                cannot pair, schema constraints are violated, or a rule cannot
                apply as configured.
        """
        return cls(config, source_df, target_df)._plan()[0]

    @staticmethod
    def check_configs(
        diff: DiffConfig, source: SourceRef, target: SourceRef, *, schemas: bool = False
    ) -> list[ConfigFinding]:
        """Check a loaded configuration for what would stop a run.

        By default nothing connects and no rows are read, so the checks need
        only the configuration and the installed extras:

        - the pair is one an engine can compare: both local, or two tables on
          one warehouse connection;
        - every extra a side reads through, or a local run scores with, is
          installed;
        - a database `table` uses a URI scheme Veridelta can quote for;
        - each `regex_replace` pattern compiles in Polars;
        - on a warehouse pair, settings the warehouse refuses for some stored
          names or types, reported as warnings.

        With `schemas`, and no errors so far, each side's columns are read too,
        but never its rows. Local sides are checked with `validate_rules`; a
        database `table` is read with a zero-row probe, and a `query` is not
        run at all. A warehouse pair runs the schema probes a run starts with,
        then compiles every comparison statement without executing it, which
        settles the warnings above one way or the other. Either way, a name in a
        rule's `column_names` that neither side has is a warning, since that rule
        does nothing.

        Args:
            diff (DiffConfig): Comparison settings and rules.
            source (SourceRef): Source configuration.
            target (SourceRef): Target configuration.
            schemas (bool): Whether to also connect and check the rules against
                the stored columns.

        Returns:
            list[ConfigFinding]: Errors and warnings, empty when nothing is
                wrong. A configuration with no errors is expected to start.
        """
        findings: list[ConfigFinding] = []
        pair: WarehousePair | None = None
        try:
            pair = check_backend_pairing(source, target)
        except (ConfigError, ConnectorError) as exc:
            findings.append(error_finding(str(exc)))
        findings += missing_extra_findings(source, target)
        findings += database_findings(source, target)
        findings += regex_findings(diff, pushdown=pair is not None)
        if pair is None:
            findings += fuzzy_extra_findings(diff)
        if not schemas:
            return findings if pair is None else findings + pushdown_findings(diff, pair)
        if any(finding.severity == "error" for finding in findings):
            return [
                *findings,
                warning_finding("Schemas were not checked, because of the errors above."),
            ]
        if pair is None:
            return findings + DiffEngine._local_schema_findings(diff, source, target)
        return findings + pushdown_schema_findings(diff, pair)

    @staticmethod
    def check_config_file(
        path: str | Path, *, schemas: bool = False, allow_missing_env: bool = False
    ) -> list[ConfigFinding]:
        """Load a configuration file and check it for what would stop a run.

        A file that does not load is one error finding, with the loader's
        message, so every problem is reported the same way. A file that loads
        gets the checks of `check_configs`. `veridelta validate` and the MCP
        server's `validate_config` tool both report through this method.

        Args:
            path (str | Path): The configuration file.
            schemas (bool): Whether to also connect and check the rules against
                the stored columns.
            allow_missing_env (bool): Whether to read an unset `${NAME}` as the
                text `NAME` and warn, instead of failing, so a file can be
                checked without its secrets.

        Returns:
            list[ConfigFinding]: A warning for each unset variable first, then
                the findings of `check_configs`, or the error that stopped the
                load.
        """
        unset: list[str] | None = [] if allow_missing_env else None
        try:
            diff, source, target = load_config(path, unset_env=unset)
            findings = DiffEngine.check_configs(diff, source, target, schemas=schemas)
        except ConfigError as exc:
            findings = [error_finding(str(exc).strip())]
        unset_findings = [
            warning_finding(
                f"Environment variable '{name}' is not set, so its references were checked "
                f"as the text '{name}'."
            )
            for name in unset or []
        ]
        return [*unset_findings, *findings]

    @staticmethod
    def read_schema(config: SourceRef) -> pl.Schema:
        """Read one side's columns and their types, and return none of its rows.

        Each side is read as a run reads it before its first row. A file is
        read by its loader, and a CSV file's types come from its first rows; a
        lakehouse table gives its schema; a database or DuckDB `table` is read
        with a probe that returns no rows. A warehouse table, or a table with
        `pushdown`, gets the probe a pushdown run starts with, in a session of
        its own that is closed after.

        Args:
            config (SourceRef): One side of a configuration.

        Returns:
            pl.Schema: The columns in their stored order and with their stored
                names, before `normalize_column_names` or a `rename_to`.

        Raises:
            ConfigError: If the side reads a database or DuckDB `query`, which a
                probe would have to run in full.
            ConnectorError: If the side cannot be reached or read.
        """
        if not is_warehouse(config):
            return schema_frame(config).collect_schema()
        with warehouse_session(config) as session:
            return probe_relation(session, table_name(config))[1]

    @classmethod
    def _local_schema_findings(
        cls, diff: DiffConfig, source: SourceRef, target: SourceRef
    ) -> list[ConfigFinding]:
        """Check the rules against two local sides' stored columns."""
        queries = [
            warning_finding(
                f"The {label} reads a query, which validate does not run, so the rules were "
                "not checked against stored columns."
            )
            for label, config in (("source", source), ("target", target))
            if isinstance(config, (DatabaseConfig, DuckDBConfig)) and config.query is not None
        ]
        if queries:
            return queries
        try:
            source_frame, target_frame = schema_frame(source), schema_frame(target)
        except (VerideltaError, OSError, pl.exceptions.PolarsError) as exc:
            return [error_finding(f"Could not read the schemas: {exc}")]
        engine = cls(diff, source_frame, target_frame)
        engine._sides = describe_sides(source, target)
        try:
            engine._plan()
        except ConfigError as exc:
            return [error_finding(str(exc))]
        return unknown_column_findings(
            diff, source_frame.collect_schema().names(), target_frame.collect_schema().names()
        )

    @classmethod
    def propose_value_maps_from_configs(
        cls,
        diff: DiffConfig,
        source: SourceRef,
        target: SourceRef,
        *,
        min_confidence: float = DEFAULT_MIN_CONFIDENCE,
        min_support: int = DEFAULT_MIN_SUPPORT,
        sample_fraction: float = 1.0,
    ) -> list[ValueMapProposal]:
        """Propose `value_map` entries for a `SourceRef` pair, wherever it lives.

        File, lakehouse, and database pairs are loaded and proposed locally. A
        pair of tables on one warehouse connection is counted in the warehouse
        instead, in one statement, for columns stored as text on both sides;
        rows never leave it. A sampled warehouse run reads a different, though
        equally repeatable, set of keys than a local one.

        Args:
            diff (DiffConfig): Comparison settings and rules.
            source (SourceRef): Source configuration.
            target (SourceRef): Target configuration.
            min_confidence (float): Share of rows that must agree, above 0.5.
            min_support (int): Agreeing rows a proposal needs.
            sample_fraction (float): Share of source rows to read.

        Returns:
            list[ValueMapProposal]: One proposal per column with new entries.

        Raises:
            ConnectorError: If only one side is a warehouse table, the sides use
                different warehouses or connections, or a warehouse returns a
                malformed result.
            ConfigError: If a threshold is out of range, or the configuration
                fails as it would in a run.
            DataIntegrityError: If either dataset repeats a normalized primary key.
        """
        # Bad thresholds fail before a session opens or a file is read.
        check_value_map_thresholds(min_confidence, min_support, sample_fraction)
        pair = check_backend_pairing(source, target)
        if pair is not None:
            return pair.with_session(
                lambda session, source_table, target_table: collect_value_map_proposals(
                    session,
                    source_table,
                    target_table,
                    diff,
                    min_confidence=min_confidence,
                    min_support=min_support,
                    sample_fraction=sample_fraction,
                )
            )
        engine = cls._on_sources(diff, source, target)
        return engine.propose_value_maps(
            min_confidence=min_confidence,
            min_support=min_support,
            sample_fraction=sample_fraction,
        )

    def propose_value_maps(
        self,
        *,
        min_confidence: float = DEFAULT_MIN_CONFIDENCE,
        min_support: int = DEFAULT_MIN_SUPPORT,
        sample_fraction: float = 1.0,
    ) -> list[ValueMapProposal]:
        """Propose `value_map` entries from how source and target values line up.

        Rows are aligned, normalized, and joined as `run()` does, on a copy, so this
        engine can still run afterward. For each compared text column that stage 4 can
        map, a source value is proposed for the target value it lines up with in at
        least `min_confidence` of its joined rows, provided at least `min_support` rows
        agree. Values are read as the `value_map` stage sees them, so a
        `case_insensitive` column gets lowercase keys. Rows a column's existing map
        already translates are left out, so a raw value equal to one of that map's
        outputs cannot receive an entry.

        Args:
            min_confidence (float): Share of rows that must agree, above 0.5.
            min_support (int): Agreeing rows a proposal needs.
            sample_fraction (float): Share of source rows to read, picked by a hash of
                the primary keys, so the same data samples the same rows under one
                Polars version.

        Returns:
            list[ValueMapProposal]: One proposal per column with new entries, in source
                column order.

        Raises:
            ConfigError: If a threshold is out of range, or the configuration fails as
                it would in a run.
            DataIntegrityError: If either dataset repeats a normalized primary key.

        Examples:
            >>> import polars as pl
            >>> from veridelta.models import DiffConfig
            >>> source = pl.LazyFrame({"id": range(6), "sex": ["M"] * 6})
            >>> target = pl.LazyFrame({"id": range(6), "sex": ["Male"] * 6})
            >>> engine = DiffEngine(DiffConfig(primary_keys=["id"]), source, target)
            >>> proposals = engine.propose_value_maps()
            >>> proposals[0].to_rule().value_map
            {'M': 'Male'}
        """
        check_value_map_thresholds(min_confidence, min_support, sample_fraction)
        prepared = type(self)(self.config, self.source, self.target)
        prepared._sides = self._sides
        prepared._align_structure()
        prepared._validate_schema()
        stored = prepared.source.collect_schema()
        prepared._normalize_sides()
        prepared._check_uniqueness()

        columns = prepared._value_map_columns(stored)
        if not columns:
            return []
        joined = prepared._value_map_join(columns, sample_fraction)
        frames = pl.collect_all(
            value_map_query(
                joined,
                column,
                mapped_values=list(
                    (prepared._get_effective_rule(column)["value_map"] or {}).values()
                ),
                min_confidence=min_confidence,
                min_support=min_support,
            )
            for column in columns
        )
        proposals = (
            value_map_proposal(prepared.config, column, frame)
            for column, frame in zip(columns, frames, strict=True)
        )
        return [proposal for proposal in proposals if proposal is not None]

    @classmethod
    def suggest_rules_from_configs(
        cls,
        diff: DiffConfig,
        source: SourceRef,
        target: SourceRef,
        *,
        max_share: float = DEFAULT_MAX_SHARE,
    ) -> list[RuleSuggestion]:
        """Suggest rules for a `SourceRef` pair, read locally.

        Args:
            diff (DiffConfig): Comparison settings and rules.
            source (SourceRef): Source configuration.
            target (SourceRef): Target configuration.
            max_share (float): Largest gap a tolerance may explain, as a share of
                the larger of its two values, above 0 and at most 1.

        Returns:
            list[RuleSuggestion]: One suggestion per column it can explain.

        Raises:
            ConfigError: If `max_share` is out of range, the pair is compared where
                it is stored, or the configuration fails as it would in a run.
            DataIntegrityError: If either dataset repeats a normalized primary key.
        """
        check_max_share(max_share)
        if check_backend_pairing(source, target) is not None:
            raise ConfigError(
                "veridelta suggest reads both sides locally, and this pair is compared "
                "where it is stored. Suggest rules on files exported from it, or on a "
                "database pair that does not set pushdown."
            )
        engine = cls._on_sources(diff, source, target)
        return engine.suggest_rules(max_share=max_share)

    def suggest_rules(self, *, max_share: float = DEFAULT_MAX_SHARE) -> list[RuleSuggestion]:
        """Suggest rules that would explain the differences in each compared column.

        The comparison runs as `run()` runs it, on a copy, without writing artifacts,
        so this engine can still run afterward. A numeric column gets a tolerance when
        every gap between its differing values is at most `max_share` of the larger
        of the two: an absolute tolerance when the gaps stay about one size, and a
        relative one when they grow with the values. Each tolerance is the round value
        just above the largest gap, and the comparison runs again with it to count
        the rows it explains. No model is called.

        Args:
            max_share (float): Largest gap a tolerance may explain, as a share of the
                larger of its two values, above 0 and at most 1.

        Returns:
            list[RuleSuggestion]: One suggestion per column it can explain, in
                compared column order.

        Raises:
            ConfigError: If `max_share` is out of range, or the configuration fails
                as it would in a run.
            DataIntegrityError: If either dataset repeats a normalized primary key.

        Examples:
            >>> import polars as pl
            >>> from veridelta.models import DiffConfig
            >>> source = pl.LazyFrame({"id": [1, 2, 3], "fare": [10.0, 20.0, 30.0]})
            >>> target = pl.LazyFrame({"id": [1, 2, 3], "fare": [10.004, 20.004, 30.003]})
            >>> engine = DiffEngine(DiffConfig(primary_keys=["id"]), source, target)
            >>> [(s.column, s.settings) for s in engine.suggest_rules()]
            [('fare', {'absolute_tolerance': 0.005})]
        """
        check_max_share(max_share)
        result = self._run_copy(self.config)
        keys = list(result.primary_keys)
        suggestions: list[RuleSuggestion] = []
        for column in result.compared_columns:
            differing = result.summary.column_mismatches.get(column, 0)
            if differing == 0:
                continue
            proposal = column_proposal(result.changed, column, max_share)
            if proposal is None:
                continue
            governing = match_rule(self.config.rules, column)
            rule = suggested_rule(
                governing, column, proposal.settings, self.config.default_null_values
            )
            try:
                tried = self._run_copy(
                    self.config.model_copy(update={"rules": [rule, *self.config.rules]})
                )
            except ConfigError:
                # The configuration refuses the rule, such as a sentinel that one side's
                # type cannot hold, so it explains nothing.
                continue
            explained = differing_only_in(result.changed, tried.changed, keys, column)
            # A rule that makes a matching row differ, such as case folding ahead of a
            # `value_map` written in capitals, is no explanation.
            broken = differing_only_in(tried.changed, result.changed, keys, column)
            if explained.is_empty() or not broken.is_empty():
                continue
            suggestions.append(
                RuleSuggestion(
                    column=column,
                    settings=proposal.settings,
                    differing=differing,
                    explained=explained.height,
                    largest_gap=proposal.largest_gap,
                    examples=tuple(explained.head(SUGGESTED_EXAMPLES).to_dicts()),
                    rule=rule,
                    governing_rule_index=(
                        None if governing is None else self.config.rules.index(governing)
                    ),
                )
            )
        return suggestions

    def _run_copy(self, config: DiffConfig) -> DiffResult:
        """Run a comparison of this engine's data under `config`, writing no artifacts."""
        engine = type(self)(
            config.model_copy(update={"output_path": None}), self.source, self.target
        )
        engine._sides = self._sides
        return engine.run()

    def _value_map_columns(self, stored: pl.Schema) -> list[str]:
        """Pick the compared columns a `value_map` proposal can apply to."""
        keys = set(self.config.primary_keys)
        target_schema = self.target.collect_schema()
        columns: list[str] = []
        for column, dtype in stored.items():
            if column in keys or column not in target_schema:
                continue
            rule = self._get_effective_rule(column)
            if rule["ignore"] or not compares_mapped_text(rule, dtype):
                continue
            if self.config.strict_types and not isinstance(target_schema[column], pl.String):
                # Types that differ always mismatch under strict_types; no map helps.
                continue
            columns.append(column)
        return columns

    def _value_map_join(self, columns: list[str], sample_fraction: float) -> pl.LazyFrame:
        """Pair each candidate column's normalized values on the primary keys."""
        keys = self.config.primary_keys
        source = self.source.select(*keys, pl.col(columns).name.suffix("_source"))
        target = self.target.select(*keys, pl.col(columns).name.suffix("_target"))
        if sample_fraction < 1:
            cutoff = round(sample_fraction * SAMPLE_BUCKETS)
            source = source.filter(pl.struct(keys).hash(seed=0) % SAMPLE_BUCKETS < cutoff)
        return source.join(target, on=keys, how="inner")

    def _get_effective_rule(self, col_name: str) -> EffectiveRule:
        """Resolve all rules (Specific > Pattern > Global) into a unified dictionary."""
        return fold_rule_defaults(match_rule(self.config.rules, col_name), self.config)

    def _check_uniqueness(self) -> None:
        """Verify that primary keys are unique in both datasets."""
        pks = self.config.primary_keys
        for side, frame in (("SOURCE", self.source), ("TARGET", self.target)):
            keys = frame.select(pks).collect()
            duplicated = keys.is_duplicated()
            if duplicated.any():
                raise duplicate_keys_error(pks, side, keys.filter(duplicated).height)

    def _normalize_sides(self) -> None:
        """Normalize both frames, then check that their primary keys can pair rows."""
        self.source = self._normalize_frame(self.source, is_source=True)
        self.target = self._normalize_frame(self.target, is_source=False)
        self._check_key_types()

    def _check_key_types(self) -> None:
        """Refuse a primary key whose two sides hold types a join cannot pair.

        A join pairs keys of one type, of two integer types, or of two float
        types, and fails on any other pair with an error that names no side.
        A rule with `cast_to` on the key brings both sides to one type before
        rows are paired.

        Raises:
            ConfigError: If a key holds two such types after normalization.
        """
        source = self.source.collect_schema()
        target = self.target.collect_schema()
        mixed = [
            f"'{key}' ({source[key]} in the source, {target[key]} in the target)"
            for key in self.config.primary_keys
            if not pairable(source[key], target[key])
        ]
        if mixed:
            raise ConfigError(
                "Rows pair only on keys of one type, or of two integer or two float types, "
                f"and these primary keys hold two other types: {', '.join(mixed)}. Give each "
                "a rule with cast_to, such as cast_to: Int64, so both sides hold one type."
            )

    def _normalize_frame(self, frame: pl.LazyFrame, *, is_source: bool) -> pl.LazyFrame:
        """Apply stages 1 through 7 of the canonical transform order to one dataset."""
        schema = frame.collect_schema()
        rules = {
            column: rule
            for column in schema.names()
            if not (rule := self._get_effective_rule(column))["ignore"]
        }

        value_exprs: list[pl.Expr] = []
        for column, rule in rules.items():
            value_expr = self._normalize_value_expr(
                column, rule, schema[column], is_source=is_source
            )
            if value_expr is not None:
                value_exprs.append(value_expr)
        if value_exprs:
            frame = frame.with_columns(value_exprs)

        temporal = [column for column, rule in rules.items() if rule["timezone"] or rule["cast_to"]]
        if not temporal:
            return frame
        # A second pass, since every expression in one `with_columns` reads the
        # input frame. Stage 6a can turn a String column into a Datetime, so the
        # timezone guard inspects the post-parse schema.
        parsed_schema = frame.collect_schema()
        return frame.with_columns(
            self._normalize_temporal_expr(column, rules[column], parsed_schema[column])
            for column in temporal
        )

    def _normalize_value_expr(
        self, column: str, rule: EffectiveRule, dtype: pl.DataType, *, is_source: bool
    ) -> pl.Expr | None:
        """Build stages 1 through 6a for one column: sentinels through datetime parsing."""
        expr = pl.col(column)
        applied = False
        # Deliberately narrower than `is_text_dtype`: `is_in` accepts Categorical
        # and Enum, but the `.str` namespace used below rejects both.
        is_text = isinstance(dtype, pl.String)

        sentinels = usable_sentinels(rule["null_values"], dtype)
        if sentinels:
            expr = pl.when(expr.is_in(sentinels)).then(None).otherwise(expr)
            applied = True
        # Only an explicit rule raises: a global default is expected to span a mixed schema.
        elif rule["null_values"] and rule["null_values_explicit"]:
            raise unusable_sentinel_error(column, dtype, rule["null_values"])

        if is_text:
            expr, text_applied = self._normalize_text_expr(expr, rule, is_source=is_source)
            applied = applied or text_applied

        if rule["pad_zeros"] is not None:
            # Stringify first so a numeric 123 and a text '00123' converge.
            expr = expr.cast(pl.String).str.zfill(rule["pad_zeros"])
            applied = True
            is_text = True

        if rule["datetime_format"] and is_text:
            expr = expr.str.strptime(
                pl.Datetime,
                format=polars_datetime_format(rule["datetime_format"]),
                strict=False,
            )
            applied = True

        return expr.alias(column) if applied else None

    @staticmethod
    def _normalize_text_expr(
        expr: pl.Expr, rule: EffectiveRule, *, is_source: bool
    ) -> tuple[pl.Expr, bool]:
        """Build stages 2 through 4 for a text column: regex, whitespace, case, value map."""
        applied = False
        if rule["regex_replace"]:
            for pattern, replacement in rule["regex_replace"].items():
                expr = expr.str.replace_all(pattern, replacement)
            applied = True

        mode = rule["whitespace"]
        if mode == "left":
            expr = expr.str.strip_chars_start()
            applied = True
        elif mode == "right":
            expr = expr.str.strip_chars_end()
            applied = True
        elif mode == "both":
            expr = expr.str.strip_chars()
            applied = True

        if rule["case_insensitive"]:
            expr = expr.str.to_lowercase()
            applied = True

        if is_source and rule["value_map"]:
            expr = expr.replace(rule["value_map"])
            applied = True

        return expr, applied

    def _normalize_temporal_expr(
        self, column: str, rule: EffectiveRule, dtype: pl.DataType
    ) -> pl.Expr:
        """Build stages 6b and 7 for one column: timezone conversion, then cast."""
        expr = pl.col(column)
        if rule["timezone"]:
            expr = self._convert_time_zone(column, expr, dtype, rule["timezone"])
        if rule["cast_to"]:
            target = rule["cast_to"]
            if target in UNCASTABLE.get(type(dtype), frozenset()):
                raise ConfigError(
                    f"Column '{column}' sets cast_to='{target}', but holds {dtype}, which "
                    f"cannot be cast to {target}. Convert it where it is read, such as in "
                    "a database query."
                )
            expr = expr.cast(CAST_TARGETS[target])
        return expr.alias(column)

    def _convert_time_zone(
        self, column: str, expr: pl.Expr, dtype: pl.DataType, zone: str
    ) -> pl.Expr:
        """Convert a timezone-aware column to `zone`, refusing to guess for naive data."""
        # `convert_time_zone` treats a naive timestamp as UTC instead of refusing it.
        reject_unzoned_timezone(column, dtype, zone)
        try:
            return expr.dt.convert_time_zone(zone)
        except pl.exceptions.ComputeError as exc:
            raise ConfigError(f"Column '{column}' sets an unusable timezone. {exc}") from exc

    def _build_match_expr(self, col_name: str, rule: EffectiveRule, dtype: pl.DataType) -> pl.Expr:
        """Build stages 8 and 9 of the transform order: comparison and null equality."""
        src = pl.col(f"{col_name}_source")
        tgt = pl.col(f"{col_name}_target")

        tgt_dtype = self.target.collect_schema().get(col_name)
        # The type the target is compared as, once any soft cast below applies.
        compared_tgt_dtype = tgt_dtype

        if dtype != tgt_dtype:
            if self.config.strict_types:
                return null_equality(pl.lit(False), src, tgt, rule)
            # Numbers skip the cast, which would truncate a Float64 `10.7` to an Int64 `10`.
            elif not (dtype.is_numeric() and tgt_dtype is not None and tgt_dtype.is_numeric()):
                tgt = tgt.cast(dtype, strict=False)
                compared_tgt_dtype = dtype

        similar = similarity_test(rule) if isinstance(dtype, pl.String) else None
        if dtype.is_numeric() and (rule["abs_tol"] != 0.0 or rule["rel_tol"] != 0.0):
            val_match = tolerance_match(src, tgt, rule, dtype, compared_tgt_dtype)
        elif similar is not None:
            # Equal text matches outright, so only a differing pair is scored.
            val_match = (src == tgt) | similarity_expr(src, tgt, similar)
        else:
            val_match = src == tgt
        return null_equality(val_match, src, tgt, rule)

    def _align_structure(self) -> None:
        """Perform structural normalization to reconcile asymmetrical schemas."""
        if self.config.normalize_column_names:
            self.source = normalize_header_names(self.source)
            self.target = normalize_header_names(self.target)
        self.source, self.target = drop_and_rename(self.config.rules, self.source, self.target)

    def _validate_schema(self) -> None:
        """Enforce the configured `SchemaMode` before comparison."""
        check_schema(self.config, self.source, self.target, self._sides)

    def run(self, *, baseline: Baseline | None = None) -> DiffResult:
        """Compare the two datasets and return the result.

        The comparison stays lazy until it collects the joins, so the inputs can be
        scans over files larger than memory. A run takes these steps, in order:

        1. Align the columns: apply renames, drop ignored columns, and treat the
           target as the authoritative schema.
        2. Check that the primary keys exist and that `schema_mode` holds.
        3. Apply stages 1 to 7 of the `DiffRule` transform order to each side, and
           build each column's stage 8 and 9 comparison, so a rule the run cannot
           honor fails before any rows move.
        4. Check that the normalized primary keys are unique on each side.
        5. Find the added, removed, and changed rows, leave out the drift
           `baseline` accepts, and count the mismatches.
        6. Write the artifacts, when `output_path` is set.

        Args:
            baseline (Baseline | None): Drift to accept: rows by kind and key, and
                changed columns by row. The counts, the verdict, the artifacts,
                and the reports leave it out, and `accepted_count` counts it.

        Returns:
            DiffResult: Counts, column-level drift, and the differing rows.

        Raises:
            ConfigError: If a primary key is missing or holds two types a join
                cannot pair, `schema_mode` is violated, a similarity limit needs
                the missing `fuzzy` extra, or the artifact format has no writer.
            DataIntegrityError: If either dataset repeats a normalized primary key.

        Examples:
            >>> import polars as pl
            >>> from veridelta.models import DiffConfig
            >>> source = pl.LazyFrame({"id": [1, 2, 3], "amount": [10.0, 20.0, 30.0]})
            >>> target = pl.LazyFrame({"id": [2, 3, 4], "amount": [20.0, 31.0, 40.0]})
            >>> result = DiffEngine(DiffConfig(primary_keys=["id"]), source, target).run()
            >>> result.summary.added_count, result.summary.removed_count
            (1, 1)
            >>> result.summary.column_mismatches
            {'amount': 1}
        """
        compared_columns, match_expressions = self._plan()

        # Runs after normalization: case folding or sentinel coercion on a key
        # column can collapse distinct rows into duplicates, and that must fail
        # here rather than silently exploding the joins below.
        self._check_uniqueness()

        keys = self.config.primary_keys
        added_df = self.target.join(self.source, on=keys, how="anti").collect()
        removed_df = self.source.join(self.target, on=keys, how="anti").collect()
        changed_df = self._collect_changed_rows(compared_columns, match_expressions)
        if baseline is None:
            return self._build_result(added_df, removed_df, changed_df, compared_columns)
        kept = accept_baseline(
            baseline, list(keys), (added_df, removed_df, changed_df), compared_columns
        )
        return self._build_result(
            kept.added, kept.removed, kept.changed, compared_columns, kept.rows, kept.accepted
        )

    def _plan(self) -> tuple[list[str], list[pl.Expr]]:
        """Align, validate, and normalize both frames, then build the comparisons."""
        self._align_structure()
        self._validate_schema()
        self._normalize_sides()
        return self._match_expressions()

    def _match_expressions(self) -> tuple[list[str], list[pl.Expr]]:
        """Build one boolean match expression per compared column."""
        # Re-read the schema here: normalization may have retyped columns.
        source_schema = self.source.collect_schema()
        target_columns = set(self.target.collect_schema().names())
        keys = set(self.config.primary_keys)

        compared: list[str] = []
        expressions: list[pl.Expr] = []
        for column in source_schema.names():
            if column in keys or column not in target_columns:
                continue
            rule = self._get_effective_rule(column)
            if rule["ignore"]:
                continue
            expr = self._build_match_expr(column, rule, source_schema[column])
            compared.append(column)
            expressions.append(expr.alias(f"{column}_is_match"))
        return compared, expressions

    def _collect_changed_rows(
        self, compared_columns: list[str], match_expressions: list[pl.Expr]
    ) -> pl.DataFrame:
        """Inner-join both sides and keep the rows where any compared column differs."""
        if not match_expressions:
            return pl.DataFrame()

        keys = self.config.primary_keys
        source = self.source.rename(lambda col: col if col in keys else f"{col}_source")
        target = self.target.rename(lambda col: col if col in keys else f"{col}_target")
        common_lazy = source.join(target, on=keys, how="inner")

        all_matched = pl.all_horizontal([f"{col}_is_match" for col in compared_columns])
        return common_lazy.with_columns(match_expressions).filter(~all_matched).collect()

    def _build_result(
        self,
        added_df: pl.DataFrame,
        removed_df: pl.DataFrame,
        changed_df: pl.DataFrame,
        compared_columns: list[str],
        accepted_count: int = 0,
        accepted: Baseline | None = None,
    ) -> DiffResult:
        """Count totals, apply the threshold, export artifacts, and assemble the result."""
        column_mismatches = local_column_mismatches(changed_df, compared_columns)

        # Count rows, as the warehouse's COUNT(*) does. Counting a key column
        # would skip rows whose key is null, which still count as removed or
        # added and so belong in the threshold's denominator.
        src_total = self.source.select(pl.len()).collect().item()
        tgt_total = self.target.select(pl.len()).collect().item()

        artifacts_written = export_artifacts(
            {"added_rows": added_df, "removed_rows": removed_df, "changed_rows": changed_df},
            self.config.output_path,
            self.config.output_format,
        )
        return DiffResult(
            summary=build_summary(
                self.config,
                changed_df,
                added_df,
                removed_df,
                src_total,
                tgt_total,
                column_mismatches,
                artifacts_written,
                accepted_count,
            ),
            added=added_df,
            removed=removed_df,
            changed=changed_df,
            primary_keys=tuple(self.config.primary_keys),
            compared_columns=tuple(compared_columns),
            accepted=accepted,
        )

__init__(config, source_df, target_df)

Hold the configuration and the two datasets to compare.

run() aligns the datasets and compares them.

Parameters:

Name Type Description Default
config DiffConfig

Keys, rules, and settings for the comparison.

required
source_df LazyFrame

Source rows, as stored.

required
target_df LazyFrame

Target rows, as stored.

required
Source code in src/veridelta/engine.py
def __init__(
    self, config: DiffConfig, source_df: pl.LazyFrame, target_df: pl.LazyFrame
) -> None:
    """Hold the configuration and the two datasets to compare.

    `run()` aligns the datasets and compares them.

    Args:
        config (DiffConfig): Keys, rules, and settings for the comparison.
        source_df (pl.LazyFrame): Source rows, as stored.
        target_df (pl.LazyFrame): Target rows, as stored.
    """
    self.config = config
    self.source = source_df
    self.target = target_df
    # How each side was read, set by the entry points that load a `SourceRef`,
    # so an error can name the file and its format rather than only a column.
    self._sides: dict[str, str] | None = None

check_config_file(path, *, schemas=False, allow_missing_env=False) staticmethod

Load a configuration file and check it for what would stop a run.

A file that does not load is one error finding, with the loader's message, so every problem is reported the same way. A file that loads gets the checks of check_configs. veridelta validate and the MCP server's validate_config tool both report through this method.

Parameters:

Name Type Description Default
path str | Path

The configuration file.

required
schemas bool

Whether to also connect and check the rules against the stored columns.

False
allow_missing_env bool

Whether to read an unset ${NAME} as the text NAME and warn, instead of failing, so a file can be checked without its secrets.

False

Returns:

Type Description
list[ConfigFinding]

list[ConfigFinding]: A warning for each unset variable first, then the findings of check_configs, or the error that stopped the load.

Source code in src/veridelta/engine.py
@staticmethod
def check_config_file(
    path: str | Path, *, schemas: bool = False, allow_missing_env: bool = False
) -> list[ConfigFinding]:
    """Load a configuration file and check it for what would stop a run.

    A file that does not load is one error finding, with the loader's
    message, so every problem is reported the same way. A file that loads
    gets the checks of `check_configs`. `veridelta validate` and the MCP
    server's `validate_config` tool both report through this method.

    Args:
        path (str | Path): The configuration file.
        schemas (bool): Whether to also connect and check the rules against
            the stored columns.
        allow_missing_env (bool): Whether to read an unset `${NAME}` as the
            text `NAME` and warn, instead of failing, so a file can be
            checked without its secrets.

    Returns:
        list[ConfigFinding]: A warning for each unset variable first, then
            the findings of `check_configs`, or the error that stopped the
            load.
    """
    unset: list[str] | None = [] if allow_missing_env else None
    try:
        diff, source, target = load_config(path, unset_env=unset)
        findings = DiffEngine.check_configs(diff, source, target, schemas=schemas)
    except ConfigError as exc:
        findings = [error_finding(str(exc).strip())]
    unset_findings = [
        warning_finding(
            f"Environment variable '{name}' is not set, so its references were checked "
            f"as the text '{name}'."
        )
        for name in unset or []
    ]
    return [*unset_findings, *findings]

check_configs(diff, source, target, *, schemas=False) staticmethod

Check a loaded configuration for what would stop a run.

By default nothing connects and no rows are read, so the checks need only the configuration and the installed extras:

  • the pair is one an engine can compare: both local, or two tables on one warehouse connection;
  • every extra a side reads through, or a local run scores with, is installed;
  • a database table uses a URI scheme Veridelta can quote for;
  • each regex_replace pattern compiles in Polars;
  • on a warehouse pair, settings the warehouse refuses for some stored names or types, reported as warnings.

With schemas, and no errors so far, each side's columns are read too, but never its rows. Local sides are checked with validate_rules; a database table is read with a zero-row probe, and a query is not run at all. A warehouse pair runs the schema probes a run starts with, then compiles every comparison statement without executing it, which settles the warnings above one way or the other. Either way, a name in a rule's column_names that neither side has is a warning, since that rule does nothing.

Parameters:

Name Type Description Default
diff DiffConfig

Comparison settings and rules.

required
source SourceRef

Source configuration.

required
target SourceRef

Target configuration.

required
schemas bool

Whether to also connect and check the rules against the stored columns.

False

Returns:

Type Description
list[ConfigFinding]

list[ConfigFinding]: Errors and warnings, empty when nothing is wrong. A configuration with no errors is expected to start.

Source code in src/veridelta/engine.py
@staticmethod
def check_configs(
    diff: DiffConfig, source: SourceRef, target: SourceRef, *, schemas: bool = False
) -> list[ConfigFinding]:
    """Check a loaded configuration for what would stop a run.

    By default nothing connects and no rows are read, so the checks need
    only the configuration and the installed extras:

    - the pair is one an engine can compare: both local, or two tables on
      one warehouse connection;
    - every extra a side reads through, or a local run scores with, is
      installed;
    - a database `table` uses a URI scheme Veridelta can quote for;
    - each `regex_replace` pattern compiles in Polars;
    - on a warehouse pair, settings the warehouse refuses for some stored
      names or types, reported as warnings.

    With `schemas`, and no errors so far, each side's columns are read too,
    but never its rows. Local sides are checked with `validate_rules`; a
    database `table` is read with a zero-row probe, and a `query` is not
    run at all. A warehouse pair runs the schema probes a run starts with,
    then compiles every comparison statement without executing it, which
    settles the warnings above one way or the other. Either way, a name in a
    rule's `column_names` that neither side has is a warning, since that rule
    does nothing.

    Args:
        diff (DiffConfig): Comparison settings and rules.
        source (SourceRef): Source configuration.
        target (SourceRef): Target configuration.
        schemas (bool): Whether to also connect and check the rules against
            the stored columns.

    Returns:
        list[ConfigFinding]: Errors and warnings, empty when nothing is
            wrong. A configuration with no errors is expected to start.
    """
    findings: list[ConfigFinding] = []
    pair: WarehousePair | None = None
    try:
        pair = check_backend_pairing(source, target)
    except (ConfigError, ConnectorError) as exc:
        findings.append(error_finding(str(exc)))
    findings += missing_extra_findings(source, target)
    findings += database_findings(source, target)
    findings += regex_findings(diff, pushdown=pair is not None)
    if pair is None:
        findings += fuzzy_extra_findings(diff)
    if not schemas:
        return findings if pair is None else findings + pushdown_findings(diff, pair)
    if any(finding.severity == "error" for finding in findings):
        return [
            *findings,
            warning_finding("Schemas were not checked, because of the errors above."),
        ]
    if pair is None:
        return findings + DiffEngine._local_schema_findings(diff, source, target)
    return findings + pushdown_schema_findings(diff, pair)

propose_value_maps(*, min_confidence=DEFAULT_MIN_CONFIDENCE, min_support=DEFAULT_MIN_SUPPORT, sample_fraction=1.0)

Propose value_map entries from how source and target values line up.

Rows are aligned, normalized, and joined as run() does, on a copy, so this engine can still run afterward. For each compared text column that stage 4 can map, a source value is proposed for the target value it lines up with in at least min_confidence of its joined rows, provided at least min_support rows agree. Values are read as the value_map stage sees them, so a case_insensitive column gets lowercase keys. Rows a column's existing map already translates are left out, so a raw value equal to one of that map's outputs cannot receive an entry.

Parameters:

Name Type Description Default
min_confidence float

Share of rows that must agree, above 0.5.

DEFAULT_MIN_CONFIDENCE
min_support int

Agreeing rows a proposal needs.

DEFAULT_MIN_SUPPORT
sample_fraction float

Share of source rows to read, picked by a hash of the primary keys, so the same data samples the same rows under one Polars version.

1.0

Returns:

Type Description
list[ValueMapProposal]

list[ValueMapProposal]: One proposal per column with new entries, in source column order.

Raises:

Type Description
ConfigError

If a threshold is out of range, or the configuration fails as it would in a run.

DataIntegrityError

If either dataset repeats a normalized primary key.

Examples:

>>> import polars as pl
>>> from veridelta.models import DiffConfig
>>> source = pl.LazyFrame({"id": range(6), "sex": ["M"] * 6})
>>> target = pl.LazyFrame({"id": range(6), "sex": ["Male"] * 6})
>>> engine = DiffEngine(DiffConfig(primary_keys=["id"]), source, target)
>>> proposals = engine.propose_value_maps()
>>> proposals[0].to_rule().value_map
{'M': 'Male'}
Source code in src/veridelta/engine.py
def propose_value_maps(
    self,
    *,
    min_confidence: float = DEFAULT_MIN_CONFIDENCE,
    min_support: int = DEFAULT_MIN_SUPPORT,
    sample_fraction: float = 1.0,
) -> list[ValueMapProposal]:
    """Propose `value_map` entries from how source and target values line up.

    Rows are aligned, normalized, and joined as `run()` does, on a copy, so this
    engine can still run afterward. For each compared text column that stage 4 can
    map, a source value is proposed for the target value it lines up with in at
    least `min_confidence` of its joined rows, provided at least `min_support` rows
    agree. Values are read as the `value_map` stage sees them, so a
    `case_insensitive` column gets lowercase keys. Rows a column's existing map
    already translates are left out, so a raw value equal to one of that map's
    outputs cannot receive an entry.

    Args:
        min_confidence (float): Share of rows that must agree, above 0.5.
        min_support (int): Agreeing rows a proposal needs.
        sample_fraction (float): Share of source rows to read, picked by a hash of
            the primary keys, so the same data samples the same rows under one
            Polars version.

    Returns:
        list[ValueMapProposal]: One proposal per column with new entries, in source
            column order.

    Raises:
        ConfigError: If a threshold is out of range, or the configuration fails as
            it would in a run.
        DataIntegrityError: If either dataset repeats a normalized primary key.

    Examples:
        >>> import polars as pl
        >>> from veridelta.models import DiffConfig
        >>> source = pl.LazyFrame({"id": range(6), "sex": ["M"] * 6})
        >>> target = pl.LazyFrame({"id": range(6), "sex": ["Male"] * 6})
        >>> engine = DiffEngine(DiffConfig(primary_keys=["id"]), source, target)
        >>> proposals = engine.propose_value_maps()
        >>> proposals[0].to_rule().value_map
        {'M': 'Male'}
    """
    check_value_map_thresholds(min_confidence, min_support, sample_fraction)
    prepared = type(self)(self.config, self.source, self.target)
    prepared._sides = self._sides
    prepared._align_structure()
    prepared._validate_schema()
    stored = prepared.source.collect_schema()
    prepared._normalize_sides()
    prepared._check_uniqueness()

    columns = prepared._value_map_columns(stored)
    if not columns:
        return []
    joined = prepared._value_map_join(columns, sample_fraction)
    frames = pl.collect_all(
        value_map_query(
            joined,
            column,
            mapped_values=list(
                (prepared._get_effective_rule(column)["value_map"] or {}).values()
            ),
            min_confidence=min_confidence,
            min_support=min_support,
        )
        for column in columns
    )
    proposals = (
        value_map_proposal(prepared.config, column, frame)
        for column, frame in zip(columns, frames, strict=True)
    )
    return [proposal for proposal in proposals if proposal is not None]

propose_value_maps_from_configs(diff, source, target, *, min_confidence=DEFAULT_MIN_CONFIDENCE, min_support=DEFAULT_MIN_SUPPORT, sample_fraction=1.0) classmethod

Propose value_map entries for a SourceRef pair, wherever it lives.

File, lakehouse, and database pairs are loaded and proposed locally. A pair of tables on one warehouse connection is counted in the warehouse instead, in one statement, for columns stored as text on both sides; rows never leave it. A sampled warehouse run reads a different, though equally repeatable, set of keys than a local one.

Parameters:

Name Type Description Default
diff DiffConfig

Comparison settings and rules.

required
source SourceRef

Source configuration.

required
target SourceRef

Target configuration.

required
min_confidence float

Share of rows that must agree, above 0.5.

DEFAULT_MIN_CONFIDENCE
min_support int

Agreeing rows a proposal needs.

DEFAULT_MIN_SUPPORT
sample_fraction float

Share of source rows to read.

1.0

Returns:

Type Description
list[ValueMapProposal]

list[ValueMapProposal]: One proposal per column with new entries.

Raises:

Type Description
ConnectorError

If only one side is a warehouse table, the sides use different warehouses or connections, or a warehouse returns a malformed result.

ConfigError

If a threshold is out of range, or the configuration fails as it would in a run.

DataIntegrityError

If either dataset repeats a normalized primary key.

Source code in src/veridelta/engine.py
@classmethod
def propose_value_maps_from_configs(
    cls,
    diff: DiffConfig,
    source: SourceRef,
    target: SourceRef,
    *,
    min_confidence: float = DEFAULT_MIN_CONFIDENCE,
    min_support: int = DEFAULT_MIN_SUPPORT,
    sample_fraction: float = 1.0,
) -> list[ValueMapProposal]:
    """Propose `value_map` entries for a `SourceRef` pair, wherever it lives.

    File, lakehouse, and database pairs are loaded and proposed locally. A
    pair of tables on one warehouse connection is counted in the warehouse
    instead, in one statement, for columns stored as text on both sides;
    rows never leave it. A sampled warehouse run reads a different, though
    equally repeatable, set of keys than a local one.

    Args:
        diff (DiffConfig): Comparison settings and rules.
        source (SourceRef): Source configuration.
        target (SourceRef): Target configuration.
        min_confidence (float): Share of rows that must agree, above 0.5.
        min_support (int): Agreeing rows a proposal needs.
        sample_fraction (float): Share of source rows to read.

    Returns:
        list[ValueMapProposal]: One proposal per column with new entries.

    Raises:
        ConnectorError: If only one side is a warehouse table, the sides use
            different warehouses or connections, or a warehouse returns a
            malformed result.
        ConfigError: If a threshold is out of range, or the configuration
            fails as it would in a run.
        DataIntegrityError: If either dataset repeats a normalized primary key.
    """
    # Bad thresholds fail before a session opens or a file is read.
    check_value_map_thresholds(min_confidence, min_support, sample_fraction)
    pair = check_backend_pairing(source, target)
    if pair is not None:
        return pair.with_session(
            lambda session, source_table, target_table: collect_value_map_proposals(
                session,
                source_table,
                target_table,
                diff,
                min_confidence=min_confidence,
                min_support=min_support,
                sample_fraction=sample_fraction,
            )
        )
    engine = cls._on_sources(diff, source, target)
    return engine.propose_value_maps(
        min_confidence=min_confidence,
        min_support=min_support,
        sample_fraction=sample_fraction,
    )

read_schema(config) staticmethod

Read one side's columns and their types, and return none of its rows.

Each side is read as a run reads it before its first row. A file is read by its loader, and a CSV file's types come from its first rows; a lakehouse table gives its schema; a database or DuckDB table is read with a probe that returns no rows. A warehouse table, or a table with pushdown, gets the probe a pushdown run starts with, in a session of its own that is closed after.

Parameters:

Name Type Description Default
config SourceRef

One side of a configuration.

required

Returns:

Type Description
Schema

pl.Schema: The columns in their stored order and with their stored names, before normalize_column_names or a rename_to.

Raises:

Type Description
ConfigError

If the side reads a database or DuckDB query, which a probe would have to run in full.

ConnectorError

If the side cannot be reached or read.

Source code in src/veridelta/engine.py
@staticmethod
def read_schema(config: SourceRef) -> pl.Schema:
    """Read one side's columns and their types, and return none of its rows.

    Each side is read as a run reads it before its first row. A file is
    read by its loader, and a CSV file's types come from its first rows; a
    lakehouse table gives its schema; a database or DuckDB `table` is read
    with a probe that returns no rows. A warehouse table, or a table with
    `pushdown`, gets the probe a pushdown run starts with, in a session of
    its own that is closed after.

    Args:
        config (SourceRef): One side of a configuration.

    Returns:
        pl.Schema: The columns in their stored order and with their stored
            names, before `normalize_column_names` or a `rename_to`.

    Raises:
        ConfigError: If the side reads a database or DuckDB `query`, which a
            probe would have to run in full.
        ConnectorError: If the side cannot be reached or read.
    """
    if not is_warehouse(config):
        return schema_frame(config).collect_schema()
    with warehouse_session(config) as session:
        return probe_relation(session, table_name(config))[1]

run(*, baseline=None)

Compare the two datasets and return the result.

The comparison stays lazy until it collects the joins, so the inputs can be scans over files larger than memory. A run takes these steps, in order:

  1. Align the columns: apply renames, drop ignored columns, and treat the target as the authoritative schema.
  2. Check that the primary keys exist and that schema_mode holds.
  3. Apply stages 1 to 7 of the DiffRule transform order to each side, and build each column's stage 8 and 9 comparison, so a rule the run cannot honor fails before any rows move.
  4. Check that the normalized primary keys are unique on each side.
  5. Find the added, removed, and changed rows, leave out the drift baseline accepts, and count the mismatches.
  6. Write the artifacts, when output_path is set.

Parameters:

Name Type Description Default
baseline Baseline | None

Drift to accept: rows by kind and key, and changed columns by row. The counts, the verdict, the artifacts, and the reports leave it out, and accepted_count counts it.

None

Returns:

Name Type Description
DiffResult DiffResult

Counts, column-level drift, and the differing rows.

Raises:

Type Description
ConfigError

If a primary key is missing or holds two types a join cannot pair, schema_mode is violated, a similarity limit needs the missing fuzzy extra, or the artifact format has no writer.

DataIntegrityError

If either dataset repeats a normalized primary key.

Examples:

>>> import polars as pl
>>> from veridelta.models import DiffConfig
>>> source = pl.LazyFrame({"id": [1, 2, 3], "amount": [10.0, 20.0, 30.0]})
>>> target = pl.LazyFrame({"id": [2, 3, 4], "amount": [20.0, 31.0, 40.0]})
>>> result = DiffEngine(DiffConfig(primary_keys=["id"]), source, target).run()
>>> result.summary.added_count, result.summary.removed_count
(1, 1)
>>> result.summary.column_mismatches
{'amount': 1}
Source code in src/veridelta/engine.py
def run(self, *, baseline: Baseline | None = None) -> DiffResult:
    """Compare the two datasets and return the result.

    The comparison stays lazy until it collects the joins, so the inputs can be
    scans over files larger than memory. A run takes these steps, in order:

    1. Align the columns: apply renames, drop ignored columns, and treat the
       target as the authoritative schema.
    2. Check that the primary keys exist and that `schema_mode` holds.
    3. Apply stages 1 to 7 of the `DiffRule` transform order to each side, and
       build each column's stage 8 and 9 comparison, so a rule the run cannot
       honor fails before any rows move.
    4. Check that the normalized primary keys are unique on each side.
    5. Find the added, removed, and changed rows, leave out the drift
       `baseline` accepts, and count the mismatches.
    6. Write the artifacts, when `output_path` is set.

    Args:
        baseline (Baseline | None): Drift to accept: rows by kind and key, and
            changed columns by row. The counts, the verdict, the artifacts,
            and the reports leave it out, and `accepted_count` counts it.

    Returns:
        DiffResult: Counts, column-level drift, and the differing rows.

    Raises:
        ConfigError: If a primary key is missing or holds two types a join
            cannot pair, `schema_mode` is violated, a similarity limit needs
            the missing `fuzzy` extra, or the artifact format has no writer.
        DataIntegrityError: If either dataset repeats a normalized primary key.

    Examples:
        >>> import polars as pl
        >>> from veridelta.models import DiffConfig
        >>> source = pl.LazyFrame({"id": [1, 2, 3], "amount": [10.0, 20.0, 30.0]})
        >>> target = pl.LazyFrame({"id": [2, 3, 4], "amount": [20.0, 31.0, 40.0]})
        >>> result = DiffEngine(DiffConfig(primary_keys=["id"]), source, target).run()
        >>> result.summary.added_count, result.summary.removed_count
        (1, 1)
        >>> result.summary.column_mismatches
        {'amount': 1}
    """
    compared_columns, match_expressions = self._plan()

    # Runs after normalization: case folding or sentinel coercion on a key
    # column can collapse distinct rows into duplicates, and that must fail
    # here rather than silently exploding the joins below.
    self._check_uniqueness()

    keys = self.config.primary_keys
    added_df = self.target.join(self.source, on=keys, how="anti").collect()
    removed_df = self.source.join(self.target, on=keys, how="anti").collect()
    changed_df = self._collect_changed_rows(compared_columns, match_expressions)
    if baseline is None:
        return self._build_result(added_df, removed_df, changed_df, compared_columns)
    kept = accept_baseline(
        baseline, list(keys), (added_df, removed_df, changed_df), compared_columns
    )
    return self._build_result(
        kept.added, kept.removed, kept.changed, compared_columns, kept.rows, kept.accepted
    )

run_from_configs(diff, source, target, *, baseline=None) classmethod

Route a comparison to pushdown or local Polars evaluation.

Parameters:

Name Type Description Default
diff DiffConfig

Comparison settings and rules.

required
source SourceRef

Source file, lakehouse, database, DuckDB, or warehouse config.

required
target SourceRef

Target file, lakehouse, database, DuckDB, or warehouse config.

required
baseline Baseline | None

Drift to accept, which a local run leaves out of the counts and the verdict.

None

Returns:

Name Type Description
DiffResult DiffResult

The result. A pushdown pair returns counts and keys, and a local pair also returns the differing rows.

Raises:

Type Description
ConfigError

If primary keys are missing, schema_mode is violated, or a baseline is given for a pair compared where it is stored.

DataIntegrityError

If either dataset repeats a normalized primary key.

ConnectorError

If warehouse backends are mixed or connections differ.

Examples:

>>> from veridelta.models import DiffConfig, SourceConfig
>>> diff = DiffConfig(primary_keys=["order_id"])
>>> source = SourceConfig(path="legacy/orders.parquet", format="parquet")
>>> target = SourceConfig(path="modern/orders.parquet", format="parquet")
>>> result = DiffEngine.run_from_configs(diff, source, target)
Source code in src/veridelta/engine.py
@classmethod
def run_from_configs(
    cls,
    diff: DiffConfig,
    source: SourceRef,
    target: SourceRef,
    *,
    baseline: Baseline | None = None,
) -> DiffResult:
    """Route a comparison to pushdown or local Polars evaluation.

    Args:
        diff (DiffConfig): Comparison settings and rules.
        source (SourceRef): Source file, lakehouse, database, DuckDB, or warehouse config.
        target (SourceRef): Target file, lakehouse, database, DuckDB, or warehouse config.
        baseline (Baseline | None): Drift to accept, which a local run leaves out
            of the counts and the verdict.

    Returns:
        DiffResult: The result. A pushdown pair returns counts and keys, and a
            local pair also returns the differing rows.

    Raises:
        ConfigError: If primary keys are missing, `schema_mode` is violated, or a
            baseline is given for a pair compared where it is stored.
        DataIntegrityError: If either dataset repeats a normalized primary key.
        ConnectorError: If warehouse backends are mixed or connections differ.

    Examples:
        >>> from veridelta.models import DiffConfig, SourceConfig
        >>> diff = DiffConfig(primary_keys=["order_id"])
        >>> source = SourceConfig(path="legacy/orders.parquet", format="parquet")
        >>> target = SourceConfig(path="modern/orders.parquet", format="parquet")
        >>> result = DiffEngine.run_from_configs(diff, source, target)  # doctest: +SKIP
    """
    pair = check_backend_pairing(source, target)
    if pair is not None:
        if baseline is not None:
            raise ConfigError(
                "A baseline applies to a run that compares both sides locally, and this "
                "pair is compared where it is stored. Leave out --baseline, or compare "
                "files exported from it."
            )
        return pair.with_session(
            lambda session, source_table, target_table: collect_pushdown_summary(
                session, source_table, target_table, diff
            )
        )

    # `run()` normalizes headers and applies renames exactly once, so the
    # frames go to it straight from the loaders: aligning them first and
    # renaming again would undo a swap and collapse a chain.
    return cls._on_sources(diff, source, target).run(baseline=baseline)

suggest_rules(*, max_share=DEFAULT_MAX_SHARE)

Suggest rules that would explain the differences in each compared column.

The comparison runs as run() runs it, on a copy, without writing artifacts, so this engine can still run afterward. A numeric column gets a tolerance when every gap between its differing values is at most max_share of the larger of the two: an absolute tolerance when the gaps stay about one size, and a relative one when they grow with the values. Each tolerance is the round value just above the largest gap, and the comparison runs again with it to count the rows it explains. No model is called.

Parameters:

Name Type Description Default
max_share float

Largest gap a tolerance may explain, as a share of the larger of its two values, above 0 and at most 1.

DEFAULT_MAX_SHARE

Returns:

Type Description
list[RuleSuggestion]

list[RuleSuggestion]: One suggestion per column it can explain, in compared column order.

Raises:

Type Description
ConfigError

If max_share is out of range, or the configuration fails as it would in a run.

DataIntegrityError

If either dataset repeats a normalized primary key.

Examples:

>>> import polars as pl
>>> from veridelta.models import DiffConfig
>>> source = pl.LazyFrame({"id": [1, 2, 3], "fare": [10.0, 20.0, 30.0]})
>>> target = pl.LazyFrame({"id": [1, 2, 3], "fare": [10.004, 20.004, 30.003]})
>>> engine = DiffEngine(DiffConfig(primary_keys=["id"]), source, target)
>>> [(s.column, s.settings) for s in engine.suggest_rules()]
[('fare', {'absolute_tolerance': 0.005})]
Source code in src/veridelta/engine.py
def suggest_rules(self, *, max_share: float = DEFAULT_MAX_SHARE) -> list[RuleSuggestion]:
    """Suggest rules that would explain the differences in each compared column.

    The comparison runs as `run()` runs it, on a copy, without writing artifacts,
    so this engine can still run afterward. A numeric column gets a tolerance when
    every gap between its differing values is at most `max_share` of the larger
    of the two: an absolute tolerance when the gaps stay about one size, and a
    relative one when they grow with the values. Each tolerance is the round value
    just above the largest gap, and the comparison runs again with it to count
    the rows it explains. No model is called.

    Args:
        max_share (float): Largest gap a tolerance may explain, as a share of the
            larger of its two values, above 0 and at most 1.

    Returns:
        list[RuleSuggestion]: One suggestion per column it can explain, in
            compared column order.

    Raises:
        ConfigError: If `max_share` is out of range, or the configuration fails
            as it would in a run.
        DataIntegrityError: If either dataset repeats a normalized primary key.

    Examples:
        >>> import polars as pl
        >>> from veridelta.models import DiffConfig
        >>> source = pl.LazyFrame({"id": [1, 2, 3], "fare": [10.0, 20.0, 30.0]})
        >>> target = pl.LazyFrame({"id": [1, 2, 3], "fare": [10.004, 20.004, 30.003]})
        >>> engine = DiffEngine(DiffConfig(primary_keys=["id"]), source, target)
        >>> [(s.column, s.settings) for s in engine.suggest_rules()]
        [('fare', {'absolute_tolerance': 0.005})]
    """
    check_max_share(max_share)
    result = self._run_copy(self.config)
    keys = list(result.primary_keys)
    suggestions: list[RuleSuggestion] = []
    for column in result.compared_columns:
        differing = result.summary.column_mismatches.get(column, 0)
        if differing == 0:
            continue
        proposal = column_proposal(result.changed, column, max_share)
        if proposal is None:
            continue
        governing = match_rule(self.config.rules, column)
        rule = suggested_rule(
            governing, column, proposal.settings, self.config.default_null_values
        )
        try:
            tried = self._run_copy(
                self.config.model_copy(update={"rules": [rule, *self.config.rules]})
            )
        except ConfigError:
            # The configuration refuses the rule, such as a sentinel that one side's
            # type cannot hold, so it explains nothing.
            continue
        explained = differing_only_in(result.changed, tried.changed, keys, column)
        # A rule that makes a matching row differ, such as case folding ahead of a
        # `value_map` written in capitals, is no explanation.
        broken = differing_only_in(tried.changed, result.changed, keys, column)
        if explained.is_empty() or not broken.is_empty():
            continue
        suggestions.append(
            RuleSuggestion(
                column=column,
                settings=proposal.settings,
                differing=differing,
                explained=explained.height,
                largest_gap=proposal.largest_gap,
                examples=tuple(explained.head(SUGGESTED_EXAMPLES).to_dicts()),
                rule=rule,
                governing_rule_index=(
                    None if governing is None else self.config.rules.index(governing)
                ),
            )
        )
    return suggestions

suggest_rules_from_configs(diff, source, target, *, max_share=DEFAULT_MAX_SHARE) classmethod

Suggest rules for a SourceRef pair, read locally.

Parameters:

Name Type Description Default
diff DiffConfig

Comparison settings and rules.

required
source SourceRef

Source configuration.

required
target SourceRef

Target configuration.

required
max_share float

Largest gap a tolerance may explain, as a share of the larger of its two values, above 0 and at most 1.

DEFAULT_MAX_SHARE

Returns:

Type Description
list[RuleSuggestion]

list[RuleSuggestion]: One suggestion per column it can explain.

Raises:

Type Description
ConfigError

If max_share is out of range, the pair is compared where it is stored, or the configuration fails as it would in a run.

DataIntegrityError

If either dataset repeats a normalized primary key.

Source code in src/veridelta/engine.py
@classmethod
def suggest_rules_from_configs(
    cls,
    diff: DiffConfig,
    source: SourceRef,
    target: SourceRef,
    *,
    max_share: float = DEFAULT_MAX_SHARE,
) -> list[RuleSuggestion]:
    """Suggest rules for a `SourceRef` pair, read locally.

    Args:
        diff (DiffConfig): Comparison settings and rules.
        source (SourceRef): Source configuration.
        target (SourceRef): Target configuration.
        max_share (float): Largest gap a tolerance may explain, as a share of
            the larger of its two values, above 0 and at most 1.

    Returns:
        list[RuleSuggestion]: One suggestion per column it can explain.

    Raises:
        ConfigError: If `max_share` is out of range, the pair is compared where
            it is stored, or the configuration fails as it would in a run.
        DataIntegrityError: If either dataset repeats a normalized primary key.
    """
    check_max_share(max_share)
    if check_backend_pairing(source, target) is not None:
        raise ConfigError(
            "veridelta suggest reads both sides locally, and this pair is compared "
            "where it is stored. Suggest rules on files exported from it, or on a "
            "database pair that does not set pushdown."
        )
    engine = cls._on_sources(diff, source, target)
    return engine.suggest_rules(max_share=max_share)

validate_rules(config, source_df, target_df) classmethod

Check everything a run checks before it reads a row.

Goes past validate_schemas: every rule is resolved against the aligned columns, both schemas are normalized, and each column's comparison is built. A rule the run could not honor fails here, such as a null sentinel its column's type cannot hold, or a similarity limit without the fuzzy extra. Operates on schema metadata only, so callers may pass zero-row frames. Repeated keys and invalid regular expressions surface only when rows are read.

Parameters:

Name Type Description Default
config DiffConfig

Comparison settings and rules.

required
source_df LazyFrame

Source frame or column probe.

required
target_df LazyFrame

Target frame or column probe.

required

Returns:

Type Description
list[str]

list[str]: The columns a run would compare, in source order, under their target names.

Raises:

Type Description
ConfigError

If primary keys are missing or hold two types a join cannot pair, schema constraints are violated, or a rule cannot apply as configured.

Source code in src/veridelta/engine.py
@classmethod
def validate_rules(
    cls, config: DiffConfig, source_df: pl.LazyFrame, target_df: pl.LazyFrame
) -> list[str]:
    """Check everything a run checks before it reads a row.

    Goes past `validate_schemas`: every rule is resolved against the aligned
    columns, both schemas are normalized, and each column's comparison is built. A
    rule the run could not honor fails here, such as a null sentinel its column's
    type cannot hold, or a similarity limit without the `fuzzy` extra. Operates on
    schema metadata only, so callers may pass zero-row frames. Repeated keys and
    invalid regular expressions surface only when rows are read.

    Args:
        config (DiffConfig): Comparison settings and rules.
        source_df (pl.LazyFrame): Source frame or column probe.
        target_df (pl.LazyFrame): Target frame or column probe.

    Returns:
        list[str]: The columns a run would compare, in source order, under their
            target names.

    Raises:
        ConfigError: If primary keys are missing or hold two types a join
            cannot pair, schema constraints are violated, or a rule cannot
            apply as configured.
    """
    return cls(config, source_df, target_df)._plan()[0]

validate_schemas(config, source_df, target_df) classmethod

Align structure and enforce SchemaMode without comparing any rows.

Operates on schema metadata only, so callers may pass zero-row frames.

Parameters:

Name Type Description Default
config DiffConfig

Comparison settings and rules.

required
source_df LazyFrame

Source frame or column probe.

required
target_df LazyFrame

Target frame or column probe.

required

Raises:

Type Description
ConfigError

If primary keys are missing or schema constraints are violated.

Source code in src/veridelta/engine.py
@classmethod
def validate_schemas(
    cls, config: DiffConfig, source_df: pl.LazyFrame, target_df: pl.LazyFrame
) -> None:
    """Align structure and enforce `SchemaMode` without comparing any rows.

    Operates on schema metadata only, so callers may pass zero-row frames.

    Args:
        config (DiffConfig): Comparison settings and rules.
        source_df (pl.LazyFrame): Source frame or column probe.
        target_df (pl.LazyFrame): Target frame or column probe.

    Raises:
        ConfigError: If primary keys are missing or schema constraints are violated.
    """
    engine = cls(config, source_df, target_df)
    engine._align_structure()
    engine._validate_schema()

EffectiveRule

Bases: TypedDict

Per-column settings after specific, pattern, and global rules are merged.

Source code in src/veridelta/_resolution.py
class EffectiveRule(TypedDict):
    """Per-column settings after specific, pattern, and global rules are merged."""

    abs_tol: float
    rel_tol: float
    treat_null: bool
    whitespace: WhitespaceMode
    null_values: list[SentinelValue]
    null_values_explicit: bool
    case_insensitive: bool
    regex_replace: dict[str, str] | None
    value_map: dict[str, str] | None
    pad_zeros: int | None
    datetime_format: str | None
    timezone: str | None
    cast_to: CastTarget | None
    ignore: bool
    max_levenshtein_distance: int | None
    min_jaro_winkler_similarity: float | None

ExcelLoader

Bases: BaseLoader

Loader for Excel workbooks, backed by the optional excel extra.

Like JSON, this is eager: a spreadsheet is a random-access container with no streaming reader.

Source code in src/veridelta/_reading.py
class ExcelLoader(BaseLoader):
    """Loader for Excel workbooks, backed by the optional `excel` extra.

    Like JSON, this is eager: a spreadsheet is a random-access container with
    no streaming reader.
    """

    def load(self, config: SourceConfig) -> pl.LazyFrame:
        """Read one worksheet.

        Args:
            config (SourceConfig): Source configuration, whose options go to
                `pl.read_excel`, such as `sheet_name`.

        Returns:
            pl.LazyFrame: A lazy wrapper over the rows read.

        Raises:
            ConfigError: If the `excel` extra is missing, or the options select more
                than one worksheet.
        """
        if fastexcel is None:
            raise ConfigError(missing_extra("excel", "Reading an Excel file"))
        loaded = pl.read_excel(  # pyright: ignore[reportUnknownVariableType] - untyped **options
            config.path, **config.options
        )
        if not isinstance(loaded, pl.DataFrame):
            raise ConfigError(
                f"Excel source '{shown_location(config.path)}' resolved to multiple worksheets. "
                "Name exactly one with the 'sheet_name' or 'sheet_id' option."
            )
        return loaded.lazy()

load(config)

Read one worksheet.

Parameters:

Name Type Description Default
config SourceConfig

Source configuration, whose options go to pl.read_excel, such as sheet_name.

required

Returns:

Type Description
LazyFrame

pl.LazyFrame: A lazy wrapper over the rows read.

Raises:

Type Description
ConfigError

If the excel extra is missing, or the options select more than one worksheet.

Source code in src/veridelta/_reading.py
def load(self, config: SourceConfig) -> pl.LazyFrame:
    """Read one worksheet.

    Args:
        config (SourceConfig): Source configuration, whose options go to
            `pl.read_excel`, such as `sheet_name`.

    Returns:
        pl.LazyFrame: A lazy wrapper over the rows read.

    Raises:
        ConfigError: If the `excel` extra is missing, or the options select more
            than one worksheet.
    """
    if fastexcel is None:
        raise ConfigError(missing_extra("excel", "Reading an Excel file"))
    loaded = pl.read_excel(  # pyright: ignore[reportUnknownVariableType] - untyped **options
        config.path, **config.options
    )
    if not isinstance(loaded, pl.DataFrame):
        raise ConfigError(
            f"Excel source '{shown_location(config.path)}' resolved to multiple worksheets. "
            "Name exactly one with the 'sheet_name' or 'sheet_id' option."
        )
    return loaded.lazy()

JSONLoader

Bases: BaseLoader

Loader for a JSON file holding one array of records.

Polars has no lazy JSON reader: a JSON array cannot be parsed incrementally the way newline-delimited records can. The file is read whole and wrapped. Prefer ndjson for large files.

Source code in src/veridelta/_reading.py
class JSONLoader(BaseLoader):
    """Loader for a JSON file holding one array of records.

    Polars has no lazy JSON reader: a JSON array cannot be parsed incrementally the
    way newline-delimited records can. The file is read whole and wrapped. Prefer
    `ndjson` for large files.
    """

    def load(self, config: SourceConfig) -> pl.LazyFrame:
        """Read a JSON file.

        Args:
            config (SourceConfig): Source configuration, whose options go to `pl.read_json`.

        Returns:
            pl.LazyFrame: A lazy wrapper over the rows read.
        """
        return pl.read_json(config.path, **config.options).lazy()

load(config)

Read a JSON file.

Parameters:

Name Type Description Default
config SourceConfig

Source configuration, whose options go to pl.read_json.

required

Returns:

Type Description
LazyFrame

pl.LazyFrame: A lazy wrapper over the rows read.

Source code in src/veridelta/_reading.py
def load(self, config: SourceConfig) -> pl.LazyFrame:
    """Read a JSON file.

    Args:
        config (SourceConfig): Source configuration, whose options go to `pl.read_json`.

    Returns:
        pl.LazyFrame: A lazy wrapper over the rows read.
    """
    return pl.read_json(config.path, **config.options).lazy()

LoaderFactory

Resolve a file, lakehouse, database, or DuckDB SourceRef to a LazyFrame.

A file source goes to the loader for its format, and its schema is read before it returns, so a file that cannot be read fails here, by name. A Delta Lake or Iceberg source returns its connector's lazy scan, and a database or DuckDB source is read once through its connector, which then closes. A warehouse source is refused: its comparison runs as SQL pushdown through DiffEngine.run_from_configs.

Attributes:

Name Type Description
_loaders ClassVar[dict[str, BaseLoader]]

Format name to loader. It is the one list of formats, and SourceType names the same set. Error messages list the supported formats from it.

Source code in src/veridelta/_reading.py
class LoaderFactory:
    """Resolve a file, lakehouse, database, or DuckDB `SourceRef` to a LazyFrame.

    A file source goes to the loader for its `format`, and its schema is read
    before it returns, so a file that cannot be read fails here, by name. A
    Delta Lake or Iceberg source returns its connector's lazy scan, and a database or DuckDB source is
    read once through its connector, which then closes. A warehouse source is
    refused: its comparison runs as SQL pushdown through
    `DiffEngine.run_from_configs`.

    Attributes:
        _loaders (ClassVar[dict[str, BaseLoader]]): Format name to loader. It is
            the one list of formats, and `SourceType` names the same set. Error
            messages list the supported formats from it.
    """

    _loaders: ClassVar[dict[str, BaseLoader]] = {
        "csv": CSVLoader(),
        "parquet": ParquetLoader(),
        "json": JSONLoader(),
        "ndjson": NDJSONLoader(),
        "arrow": ArrowLoader(),
        "avro": AvroLoader(),
        "excel": ExcelLoader(),
    }

    @classmethod
    def get_loader(cls, source_type: str) -> BaseLoader:
        """Return the loader for a file format.

        Args:
            source_type (str): Format name, such as `csv` or `parquet`.

        Returns:
            BaseLoader: The loader.

        Raises:
            ConfigError: If the format has no loader.
        """
        loader = cls._loaders.get(source_type)
        if loader is None:
            supported = ", ".join(sorted(cls._loaders))
            raise ConfigError(
                f"Source format '{source_type}' has no loader. Supported formats: {supported}."
            )
        return loader

    @classmethod
    def load(cls, config: SourceRef) -> pl.LazyFrame:
        """Load a file, lakehouse, database, or DuckDB source into a LazyFrame.

        Args:
            config (SourceRef): File, Delta, Iceberg, database, or DuckDB
                configuration.

        Returns:
            pl.LazyFrame: Unevaluated scan graph, or a lazy wrapper over the
                rows a database or DuckDB source read.

        Raises:
            ConnectorError: If `config` is a warehouse source, a file is missing
                or cannot be read, or a lakehouse scan, database read, or DuckDB
                read fails.
            ConfigError: If the file format has no loader, or a database
                `table` names a scheme Veridelta cannot quote for.
        """
        reader = _READERS.get(type(config))
        if reader is not None:
            with reader(config) as connector:
                connector.connect()
                return connector.lazyframe()
        if isinstance(config, SourceConfig):
            loader = cls.get_loader(config.format)
            try:
                frame = loader.load(config)
                # A scan reads nothing until the comparison runs, where a missing file
                # would fail unexplained. Reading the schema opens the file now.
                columns = len(frame.collect_schema())
            except FileNotFoundError as exc:
                raise ConnectorError(
                    f"The {config.format} file '{shown_location(config.path)}' does not exist."
                ) from exc
            except (OSError, pl.exceptions.PolarsError) as exc:
                raise ConnectorError(
                    f"Reading the {config.format} file '{shown_location(config.path)}' failed: "
                    f"{without_location(str(exc), config.path)}"
                ) from exc
            logger.info(
                "Opened the %s file '%s' (%d columns)",
                config.format,
                shown_location(config.path),
                columns,
            )
            return frame
        raise ConnectorError(
            "Warehouse sources cannot be loaded via LoaderFactory; "
            "use DiffEngine.run_from_configs for SQL pushdown."
        )

get_loader(source_type) classmethod

Return the loader for a file format.

Parameters:

Name Type Description Default
source_type str

Format name, such as csv or parquet.

required

Returns:

Name Type Description
BaseLoader BaseLoader

The loader.

Raises:

Type Description
ConfigError

If the format has no loader.

Source code in src/veridelta/_reading.py
@classmethod
def get_loader(cls, source_type: str) -> BaseLoader:
    """Return the loader for a file format.

    Args:
        source_type (str): Format name, such as `csv` or `parquet`.

    Returns:
        BaseLoader: The loader.

    Raises:
        ConfigError: If the format has no loader.
    """
    loader = cls._loaders.get(source_type)
    if loader is None:
        supported = ", ".join(sorted(cls._loaders))
        raise ConfigError(
            f"Source format '{source_type}' has no loader. Supported formats: {supported}."
        )
    return loader

load(config) classmethod

Load a file, lakehouse, database, or DuckDB source into a LazyFrame.

Parameters:

Name Type Description Default
config SourceRef

File, Delta, Iceberg, database, or DuckDB configuration.

required

Returns:

Type Description
LazyFrame

pl.LazyFrame: Unevaluated scan graph, or a lazy wrapper over the rows a database or DuckDB source read.

Raises:

Type Description
ConnectorError

If config is a warehouse source, a file is missing or cannot be read, or a lakehouse scan, database read, or DuckDB read fails.

ConfigError

If the file format has no loader, or a database table names a scheme Veridelta cannot quote for.

Source code in src/veridelta/_reading.py
@classmethod
def load(cls, config: SourceRef) -> pl.LazyFrame:
    """Load a file, lakehouse, database, or DuckDB source into a LazyFrame.

    Args:
        config (SourceRef): File, Delta, Iceberg, database, or DuckDB
            configuration.

    Returns:
        pl.LazyFrame: Unevaluated scan graph, or a lazy wrapper over the
            rows a database or DuckDB source read.

    Raises:
        ConnectorError: If `config` is a warehouse source, a file is missing
            or cannot be read, or a lakehouse scan, database read, or DuckDB
            read fails.
        ConfigError: If the file format has no loader, or a database
            `table` names a scheme Veridelta cannot quote for.
    """
    reader = _READERS.get(type(config))
    if reader is not None:
        with reader(config) as connector:
            connector.connect()
            return connector.lazyframe()
    if isinstance(config, SourceConfig):
        loader = cls.get_loader(config.format)
        try:
            frame = loader.load(config)
            # A scan reads nothing until the comparison runs, where a missing file
            # would fail unexplained. Reading the schema opens the file now.
            columns = len(frame.collect_schema())
        except FileNotFoundError as exc:
            raise ConnectorError(
                f"The {config.format} file '{shown_location(config.path)}' does not exist."
            ) from exc
        except (OSError, pl.exceptions.PolarsError) as exc:
            raise ConnectorError(
                f"Reading the {config.format} file '{shown_location(config.path)}' failed: "
                f"{without_location(str(exc), config.path)}"
            ) from exc
        logger.info(
            "Opened the %s file '%s' (%d columns)",
            config.format,
            shown_location(config.path),
            columns,
        )
        return frame
    raise ConnectorError(
        "Warehouse sources cannot be loaded via LoaderFactory; "
        "use DiffEngine.run_from_configs for SQL pushdown."
    )

NDJSONLoader

Bases: BaseLoader

Streaming loader for newline-delimited JSON over pl.scan_ndjson.

One record per line is the JSON shape Polars can read incrementally, so this is the format to prefer over json for large exports.

Source code in src/veridelta/_reading.py
class NDJSONLoader(BaseLoader):
    """Streaming loader for newline-delimited JSON over `pl.scan_ndjson`.

    One record per line is the JSON shape Polars can read incrementally, so
    this is the format to prefer over `json` for large exports.
    """

    def load(self, config: SourceConfig) -> pl.LazyFrame:
        """Scan a newline-delimited JSON file.

        Args:
            config (SourceConfig): Source configuration, whose options go to `pl.scan_ndjson`.

        Returns:
            pl.LazyFrame: The unevaluated rows.
        """
        return pl.scan_ndjson(config.path, **config.options)

load(config)

Scan a newline-delimited JSON file.

Parameters:

Name Type Description Default
config SourceConfig

Source configuration, whose options go to pl.scan_ndjson.

required

Returns:

Type Description
LazyFrame

pl.LazyFrame: The unevaluated rows.

Source code in src/veridelta/_reading.py
def load(self, config: SourceConfig) -> pl.LazyFrame:
    """Scan a newline-delimited JSON file.

    Args:
        config (SourceConfig): Source configuration, whose options go to `pl.scan_ndjson`.

    Returns:
        pl.LazyFrame: The unevaluated rows.
    """
    return pl.scan_ndjson(config.path, **config.options)

ParquetLoader

Bases: BaseLoader

Streaming Parquet loader over pl.scan_parquet.

Accepts whatever path or glob the Polars scanner accepts. Because the scan stays lazy, columns dropped by ignore rules are never read from disk.

Source code in src/veridelta/_reading.py
class ParquetLoader(BaseLoader):
    """Streaming Parquet loader over `pl.scan_parquet`.

    Accepts whatever path or glob the Polars scanner accepts. Because the scan
    stays lazy, columns dropped by `ignore` rules are never read from disk.
    """

    def load(self, config: SourceConfig) -> pl.LazyFrame:
        """Scan a Parquet file.

        Args:
            config (SourceConfig): Source configuration, whose options go to `pl.scan_parquet`.

        Returns:
            pl.LazyFrame: The unevaluated rows.
        """
        return pl.scan_parquet(config.path, **config.options)

load(config)

Scan a Parquet file.

Parameters:

Name Type Description Default
config SourceConfig

Source configuration, whose options go to pl.scan_parquet.

required

Returns:

Type Description
LazyFrame

pl.LazyFrame: The unevaluated rows.

Source code in src/veridelta/_reading.py
def load(self, config: SourceConfig) -> pl.LazyFrame:
    """Scan a Parquet file.

    Args:
        config (SourceConfig): Source configuration, whose options go to `pl.scan_parquet`.

    Returns:
        pl.LazyFrame: The unevaluated rows.
    """
    return pl.scan_parquet(config.path, **config.options)

Configuration loading

Functions that read and check a YAML file, or build a configuration from two files, and return configuration models. The models themselves live in veridelta.models, which Configuration models documents.

Configuration files: loading, checking, and their JSON Schema.

load_config reads a YAML file into the models the engine runs on, and config_json_schema describes the same files for editors and validators.

SCHEMA_URL = 'https://veridelta.github.io/veridelta/schema/veridelta.schema.json' module-attribute

Where the docs site publishes the configuration schema for editors to fetch.

config_json_schema()

Return a JSON Schema for configuration files, for editors and validators.

It is generated from the same models load_config validates with, then adjusted where the loader does something before validating:

  • type is required in every warehouse, lakehouse, and database block, because the loader reads a block without one as a file source.
  • Patterned and enumerated strings inside source and target, such as table, also accept a ${NAME} reference, which the loader expands.

The schema is stricter than the loader in one way: it does not model Pydantic's lax coercion, so a quoted number such as threshold: "0.1" is flagged even though it loads.

Returns:

Type Description
dict[str, Any]

dict[str, Any]: A Draft 2020-12 JSON Schema.

Source code in src/veridelta/config.py
def config_json_schema() -> dict[str, Any]:
    """Return a JSON Schema for configuration files, for editors and validators.

    It is generated from the same models `load_config` validates with, then
    adjusted where the loader does something before validating:

    - `type` is required in every warehouse, lakehouse, and database block,
      because the loader reads a block without one as a file source.
    - Patterned and enumerated strings inside `source` and `target`, such as
      `table`, also accept a `${NAME}` reference, which the loader expands.

    The schema is stricter than the loader in one way: it does not model
    Pydantic's lax coercion, so a quoted number such as `threshold: "0.1"` is
    flagged even though it loads.

    Returns:
        dict[str, Any]: A Draft 2020-12 JSON Schema.
    """
    schema = _RootConfig.model_json_schema()
    definitions: dict[str, dict[str, Any]] = schema["$defs"]
    branches: dict[str, str] = schema["properties"]["source"]["discriminator"]["mapping"]
    for tag, ref in branches.items():
        branch = definitions[ref.rsplit("/", 1)[1]]
        if tag != "file":
            branch["required"] = [*branch.get("required", []), "type"]
        for name, prop in branch["properties"].items():
            # Text restricted by `pattern` or `enum`, directly or in an `anyOf`.
            options: list[dict[str, Any]] = prop.get("anyOf", [prop])
            if name != "type" and any("pattern" in o or "enum" in o for o in options):
                branch["properties"][name] = _accept_env_reference(prop)
    return {
        "$schema": "https://json-schema.org/draft/2020-12/schema",
        "$id": SCHEMA_URL,
        **schema,
        "title": "Veridelta configuration",
        "description": "A Veridelta comparison: primary keys, rules, and the source and target.",
    }

files_config(source, target, primary_keys)

Build the configuration that compares two files on their primary keys.

It holds what a file with only primary_keys and a path for each side holds, and is checked the same way, so each file's format follows its suffix. The paths are read as given: a shell has already expanded its own variables, so a ${NAME} in one is left as it is.

Parameters:

Name Type Description Default
source str | Path

The source file.

required
target str | Path

The target file.

required
primary_keys Sequence[str]

The columns that identify a row on both sides.

required

Returns:

Type Description
tuple[DiffConfig, SourceConfig, SourceConfig]

tuple[DiffConfig, SourceConfig, SourceConfig]: The comparison settings, then the source and the target, as load_config returns them.

Raises:

Type Description
ConfigError

If primary_keys is empty.

Examples:

>>> diff, source, target = files_config("legacy.csv", "modern.parquet", ["id"])
>>> diff.primary_keys, source.path, target.format
(['id'], 'legacy.csv', 'parquet')
Source code in src/veridelta/config.py
def files_config(
    source: str | Path, target: str | Path, primary_keys: Sequence[str]
) -> tuple[DiffConfig, SourceConfig, SourceConfig]:
    """Build the configuration that compares two files on their primary keys.

    It holds what a file with only `primary_keys` and a `path` for each side holds,
    and is checked the same way, so each file's format follows its suffix. The
    paths are read as given: a shell has already expanded its own variables, so a
    `${NAME}` in one is left as it is.

    Args:
        source (str | Path): The source file.
        target (str | Path): The target file.
        primary_keys (Sequence[str]): The columns that identify a row on both sides.

    Returns:
        tuple[DiffConfig, SourceConfig, SourceConfig]: The comparison settings, then
            the source and the target, as `load_config` returns them.

    Raises:
        ConfigError: If `primary_keys` is empty.

    Examples:
        >>> diff, source, target = files_config("legacy.csv", "modern.parquet", ["id"])
        >>> diff.primary_keys, source.path, target.format
        (['id'], 'legacy.csv', 'parquet')
    """
    sides = [SourceConfig.model_validate({"path": str(path)}) for path in (source, target)]
    try:
        diff = DiffConfig.model_validate({"primary_keys": list(primary_keys)})
    except ValidationError as e:
        raise _validation_failure(e) from e
    return diff, sides[0], sides[1]

load_config(path, *, unset_env=None)

Load and validate a configuration file.

The source and target blocks become source configurations, and every other root key belongs to the DiffConfig. A file source may omit type, which defaults to file.

Strings inside source and target may reference environment variables as ${NAME}, or ${NAME:-default} to fall back when the variable is unset or empty, so credentials can stay out of the file. $${ writes a literal ${. Root settings and rules are read verbatim, which keeps a ${1} in a regex replacement intact.

Passing a list as unset_env checks a file without its secrets: an unset variable with no default then reads as its own name, so ${TABLE} becomes TABLE, and its name is appended to the list once. A block that fails validation after such a guess says which variables it guessed.

Parameters:

Name Type Description Default
path str | Path

Path to the YAML file.

required
unset_env list[str] | None

Collects unset variables instead of raising for them. None, the default, raises.

None

Returns:

Type Description
tuple[DiffConfig, SourceRef, SourceRef]

tuple[DiffConfig, SourceRef, SourceRef]: The comparison settings, then the source and the target.

Raises:

Type Description
ConfigError

If the file is missing or not valid YAML, lacks a source or target block, references an unset environment variable, holds a malformed reference, or fails validation.

Examples:

>>> from pathlib import Path
>>> from tempfile import TemporaryDirectory
>>> text = "{primary_keys: [id], source: {path: a.csv}, target: {path: b.csv}}"
>>> with TemporaryDirectory() as folder:
...     path = Path(folder, "veridelta.yaml")
...     _ = path.write_text(text)
...     diff, source, target = load_config(path)
>>> diff.primary_keys, source.path
(['id'], 'a.csv')
Source code in src/veridelta/config.py
def load_config(
    path: str | Path, *, unset_env: list[str] | None = None
) -> tuple[DiffConfig, SourceRef, SourceRef]:
    """Load and validate a configuration file.

    The `source` and `target` blocks become source configurations, and every other
    root key belongs to the `DiffConfig`. A file source may omit `type`, which
    defaults to `file`.

    Strings inside `source` and `target` may reference environment variables as
    `${NAME}`, or `${NAME:-default}` to fall back when the variable is unset or
    empty, so credentials can stay out of the file. `$${` writes a literal `${`.
    Root settings and rules are read verbatim, which keeps a `${1}` in a regex
    replacement intact.

    Passing a list as `unset_env` checks a file without its secrets: an unset
    variable with no default then reads as its own name, so `${TABLE}` becomes
    `TABLE`, and its name is appended to the list once. A block that fails
    validation after such a guess says which variables it guessed.

    Args:
        path (str | Path): Path to the YAML file.
        unset_env (list[str] | None): Collects unset variables instead of raising
            for them. None, the default, raises.

    Returns:
        tuple[DiffConfig, SourceRef, SourceRef]: The comparison settings, then the
            source and the target.

    Raises:
        ConfigError: If the file is missing or not valid YAML, lacks a `source` or
            `target` block, references an unset environment variable, holds a
            malformed reference, or fails validation.

    Examples:
        >>> from pathlib import Path
        >>> from tempfile import TemporaryDirectory
        >>> text = "{primary_keys: [id], source: {path: a.csv}, target: {path: b.csv}}"
        >>> with TemporaryDirectory() as folder:
        ...     path = Path(folder, "veridelta.yaml")
        ...     _ = path.write_text(text)
        ...     diff, source, target = load_config(path)
        >>> diff.primary_keys, source.path
        (['id'], 'a.csv')
    """
    file_path = Path(path)

    if not file_path.is_file():
        raise ConfigError(f"Configuration file not found or is not a file: {file_path.absolute()}")

    try:
        with file_path.open("r", encoding="utf-8") as f:
            parsed_yaml: Any = yaml.safe_load(f)
    except yaml.YAMLError as yaml_err:
        raise ConfigError(f"Failed to parse YAML file:\n{yaml_err}") from yaml_err

    if not isinstance(parsed_yaml, dict):
        raise ConfigError("Invalid YAML structure: Root element must be a dictionary.")

    raw_config = cast("dict[str, Any]", parsed_yaml)
    if "source" not in raw_config or "target" not in raw_config:
        raise ConfigError("Configuration must contain both 'source' and 'target' blocks.")

    source_cfg = _parse_source_ref(raw_config.pop("source"), label="source", unset=unset_env)
    target_cfg = _parse_source_ref(raw_config.pop("target"), label="target", unset=unset_env)
    try:
        diff_cfg = DiffConfig.model_validate(raw_config)
    except ValidationError as e:
        raise _validation_failure(e) from e
    return diff_cfg, source_cfg, target_cfg

referenced_variables(text)

Name each environment variable a configuration's text references, once, in order.

The text is read as written, so a file that does not load names its variables too, and a reference outside source and target, which the loader leaves as it is, counts as well.

Parameters:

Name Type Description Default
text str

The text of a configuration file.

required

Returns:

Type Description
list[str]

list[str]: The name in each ${NAME} or ${NAME:-default}. An escaped $${ names none.

Examples:

>>> referenced_variables("{password: ${PASSWORD}, role: '${ROLE:-ANALYST}$${X}'}")
['PASSWORD', 'ROLE']
Source code in src/veridelta/config.py
def referenced_variables(text: str) -> list[str]:
    """Name each environment variable a configuration's text references, once, in order.

    The text is read as written, so a file that does not load names its
    variables too, and a reference outside `source` and `target`, which the
    loader leaves as it is, counts as well.

    Args:
        text (str): The text of a configuration file.

    Returns:
        list[str]: The name in each `${NAME}` or `${NAME:-default}`. An escaped
            `$${` names none.

    Examples:
        >>> referenced_variables("{password: ${PASSWORD}, role: '${ROLE:-ANALYST}$${X}'}")
        ['PASSWORD', 'ROLE']
    """
    names: list[str] = []
    for match in _ENV_REFERENCE.finditer(text):
        name = match["name"]
        if name is not None and name not in names:
            names.append(name)
    return names

Exceptions

The errors Veridelta raises. Each derives from VerideltaError, so one except clause catches them all without hiding Python's own errors.

The errors Veridelta raises.

Each derives from VerideltaError, so one except clause catches them all.

ConfigError

Bases: VerideltaError

Raised when a configuration is invalid or the data breaks its rules.

Such failures include a primary key missing from a dataset, a rule the column types cannot satisfy, and a column that schema_mode forbids.

Source code in src/veridelta/exceptions.py
class ConfigError(VerideltaError):
    """Raised when a configuration is invalid or the data breaks its rules.

    Such failures include a primary key missing from a dataset, a rule the
    column types cannot satisfy, and a column that `schema_mode` forbids.
    """

ConnectorError

Bases: VerideltaError

Raised when a source connector cannot complete an operation.

Such failures include a missing optional extra, a failed query or scan, a pair of backends that cannot be compared, and a call before a session opens.

Source code in src/veridelta/exceptions.py
class ConnectorError(VerideltaError):
    """Raised when a source connector cannot complete an operation.

    Such failures include a missing optional extra, a failed query or scan, a
    pair of backends that cannot be compared, and a call before a session opens.
    """

DataIntegrityError

Bases: VerideltaError

Raised when the data breaks an assumption the comparison relies on.

Repeated primary keys in either dataset raise it before any join runs, since a repeated key multiplies the joined rows.

Source code in src/veridelta/exceptions.py
class DataIntegrityError(VerideltaError):
    """Raised when the data breaks an assumption the comparison relies on.

    Repeated primary keys in either dataset raise it before any join runs,
    since a repeated key multiplies the joined rows.
    """

DatasetError

Bases: VerideltaError

Raised when a sample dataset cannot be downloaded.

veridelta.datasets fetches tutorial data over the network, and a failed or interrupted download raises this error.

Source code in src/veridelta/exceptions.py
class DatasetError(VerideltaError):
    """Raised when a sample dataset cannot be downloaded.

    `veridelta.datasets` fetches tutorial data over the network, and a failed
    or interrupted download raises this error.
    """

VerideltaError

Bases: Exception

Base class for every error Veridelta raises.

Catch it to handle any Veridelta failure without also catching Python's own errors, such as MemoryError or ValueError.

Source code in src/veridelta/exceptions.py
class VerideltaError(Exception):
    """Base class for every error Veridelta raises.

    Catch it to handle any Veridelta failure without also catching Python's own
    errors, such as `MemoryError` or `ValueError`.
    """

missing_extra(extra, needed_for)

Return the one message for an optional extra that is not installed.

The error that carries it stays the one its caller raises: a reader's ConfigError or a connector's ConnectorError.

Parameters:

Name Type Description Default
extra str

The extra, such as snowflake.

required
needed_for str

What needs it, as the subject of a sentence, such as Connecting to Snowflake.

required

Returns:

Name Type Description
str str

What needs the extra, and the command that installs it.

Source code in src/veridelta/exceptions.py
def missing_extra(extra: str, needed_for: str) -> str:
    """Return the one message for an optional extra that is not installed.

    The error that carries it stays the one its caller raises: a reader's
    `ConfigError` or a connector's `ConnectorError`.

    Args:
        extra (str): The extra, such as `snowflake`.
        needed_for (str): What needs it, as the subject of a sentence, such as
            `Connecting to Snowflake`.

    Returns:
        str: What needs the extra, and the command that installs it.
    """
    return (
        f"{needed_for} needs the optional '{extra}' extra, which is not installed. "
        f"Install it with: uv add 'veridelta[{extra}]'"
    )

Connectors

Warehouse sessions, lakehouse scanners, and the database and DuckDB readers. VerideltaConnector is the lifecycle they share, ReaderConnector and PushdownSession are the two kinds the engine drives, and SQLPushdownCompiler writes each dialect's comparison SQL.

Warehouse pushdown, lakehouse-native, database, and DuckDB connector abstractions.

PushdownQueryType = Literal['mismatch', 'added', 'missing', 'count', 'duplicates', 'columns', 'schema', 'value_maps', 'settings', 'samples'] module-attribute

Warehouse pushdown round-trip: comparison rows, tallies, totals, key checks, probes, value map evidence, a check of the server's settings, or a sample of changed rows with their values.

BigQueryConnector

Bases: PushdownSession

BigQuery warehouse connector backed by the optional BigQuery extra.

connect() builds a bigquery.Client for the configured project, from a service account key file when credentials_path is set and from Application Default Credentials otherwise. Every statement runs as a GoogleSQL job under one QueryJobConfig, which names the default dataset and the maximum_bytes_billed cap. Install the client with uv add 'veridelta[bigquery]'; it is imported on first connect.

Attributes:

Name Type Description
compiler SQLPushdownCompiler

BigQuery-dialect compiler (backtick quoting, GoogleSQL type names) the engine uses for every statement.

Source code in src/veridelta/connectors/warehouse.py
class BigQueryConnector(PushdownSession):
    """BigQuery warehouse connector backed by the optional BigQuery extra.

    `connect()` builds a `bigquery.Client` for the configured project, from a
    service account key file when `credentials_path` is set and from
    Application Default Credentials otherwise. Every statement runs as a
    GoogleSQL job under one `QueryJobConfig`, which names the default dataset
    and the `maximum_bytes_billed` cap. Install the client with
    `uv add 'veridelta[bigquery]'`; it is imported on first connect.

    Attributes:
        compiler (SQLPushdownCompiler): BigQuery-dialect compiler (backtick
            quoting, GoogleSQL type names) the engine uses for every statement.
    """

    def __init__(self, config: BigQueryConfig) -> None:
        """Initialize the connector with validated BigQuery settings.

        Args:
            config (BigQueryConfig): Frozen project, table, and job settings.
        """
        self._config = config
        self.compiler = SQLPushdownCompiler(SQLDialect.BIGQUERY)
        self._client: Any = None
        self._job_config: Any = None

    def connect(self) -> None:
        """Create the BigQuery client and the job settings every statement uses.

        Raises:
            ConnectorError: If the BigQuery extra is missing or the client
                cannot be created, such as when no credentials are found.
        """
        driver = bigquery
        if driver is None:
            try:
                driver = importlib.import_module("google.cloud.bigquery")
            except ImportError:
                raise ConnectorError(_BIGQUERY_EXTRA) from None
        project, location = self._config.project, self._config.location
        # Legacy SQL rejects backtick quoting, so GoogleSQL is set explicitly
        # rather than trusted to stay the client's default.
        settings: dict[str, Any] = {"use_legacy_sql": False}
        if self._config.dataset is not None:
            settings["default_dataset"] = f"{project}.{self._config.dataset}"
        if self._config.maximum_bytes_billed is not None:
            settings["maximum_bytes_billed"] = self._config.maximum_bytes_billed
        try:
            if self._config.credentials_path is None:
                client = driver.Client(project=project, location=location)
            else:
                client = driver.Client.from_service_account_json(
                    self._config.credentials_path, project=project, location=location
                )
            job_config = driver.QueryJobConfig(**settings)
        except Exception as exc:
            logger.warning("BigQuery connection to project %s failed", project)
            # Not chained, as for the other warehouses: the message already holds the
            # driver's reason, and its traceback can quote the credentials it read.
            raise ConnectorError(f"Failed to connect to BigQuery: {exc}") from None
        self._client, self._job_config = client, job_config
        logger.info("Connected to BigQuery project %s", project)

    def execute_pushdown(
        self, statement: str, query_type: PushdownQueryType = "mismatch"
    ) -> pl.LazyFrame:
        """Execute compiler SQL as a BigQuery job and return a LazyFrame.

        Args:
            statement (str): SQL produced by `SQLPushdownCompiler`.
            query_type (PushdownQueryType): Which comparison round-trip this
                statement represents; recorded in the log line for the call.

        Returns:
            pl.LazyFrame: Unevaluated frame wrapped around the Arrow result.

        Raises:
            ConnectorError: If the connector is not connected, the job fails,
                or it returns no Arrow batches.
        """
        self._require_client()
        batches = _run_bigquery_query(
            self._client, statement, self._job_config, query_type=query_type
        )
        return _lazy_from_arrow(batches)

    def close(self) -> None:
        """Close the BigQuery client, if one is open.

        Idempotent. Afterwards `execute_pushdown` raises `ConnectorError` until
        `connect()` is called again.
        """
        if self._client is None:
            return
        client = self._client
        self._client = self._job_config = None
        _close_session(client, "BigQuery")

    def _require_client(self) -> None:
        """Ensure `connect()` has created a client."""
        if self._client is None:
            raise ConnectorError(_UNCONNECTED)

__init__(config)

Initialize the connector with validated BigQuery settings.

Parameters:

Name Type Description Default
config BigQueryConfig

Frozen project, table, and job settings.

required
Source code in src/veridelta/connectors/warehouse.py
def __init__(self, config: BigQueryConfig) -> None:
    """Initialize the connector with validated BigQuery settings.

    Args:
        config (BigQueryConfig): Frozen project, table, and job settings.
    """
    self._config = config
    self.compiler = SQLPushdownCompiler(SQLDialect.BIGQUERY)
    self._client: Any = None
    self._job_config: Any = None

close()

Close the BigQuery client, if one is open.

Idempotent. Afterwards execute_pushdown raises ConnectorError until connect() is called again.

Source code in src/veridelta/connectors/warehouse.py
def close(self) -> None:
    """Close the BigQuery client, if one is open.

    Idempotent. Afterwards `execute_pushdown` raises `ConnectorError` until
    `connect()` is called again.
    """
    if self._client is None:
        return
    client = self._client
    self._client = self._job_config = None
    _close_session(client, "BigQuery")

connect()

Create the BigQuery client and the job settings every statement uses.

Raises:

Type Description
ConnectorError

If the BigQuery extra is missing or the client cannot be created, such as when no credentials are found.

Source code in src/veridelta/connectors/warehouse.py
def connect(self) -> None:
    """Create the BigQuery client and the job settings every statement uses.

    Raises:
        ConnectorError: If the BigQuery extra is missing or the client
            cannot be created, such as when no credentials are found.
    """
    driver = bigquery
    if driver is None:
        try:
            driver = importlib.import_module("google.cloud.bigquery")
        except ImportError:
            raise ConnectorError(_BIGQUERY_EXTRA) from None
    project, location = self._config.project, self._config.location
    # Legacy SQL rejects backtick quoting, so GoogleSQL is set explicitly
    # rather than trusted to stay the client's default.
    settings: dict[str, Any] = {"use_legacy_sql": False}
    if self._config.dataset is not None:
        settings["default_dataset"] = f"{project}.{self._config.dataset}"
    if self._config.maximum_bytes_billed is not None:
        settings["maximum_bytes_billed"] = self._config.maximum_bytes_billed
    try:
        if self._config.credentials_path is None:
            client = driver.Client(project=project, location=location)
        else:
            client = driver.Client.from_service_account_json(
                self._config.credentials_path, project=project, location=location
            )
        job_config = driver.QueryJobConfig(**settings)
    except Exception as exc:
        logger.warning("BigQuery connection to project %s failed", project)
        # Not chained, as for the other warehouses: the message already holds the
        # driver's reason, and its traceback can quote the credentials it read.
        raise ConnectorError(f"Failed to connect to BigQuery: {exc}") from None
    self._client, self._job_config = client, job_config
    logger.info("Connected to BigQuery project %s", project)

execute_pushdown(statement, query_type='mismatch')

Execute compiler SQL as a BigQuery job and return a LazyFrame.

Parameters:

Name Type Description Default
statement str

SQL produced by SQLPushdownCompiler.

required
query_type PushdownQueryType

Which comparison round-trip this statement represents; recorded in the log line for the call.

'mismatch'

Returns:

Type Description
LazyFrame

pl.LazyFrame: Unevaluated frame wrapped around the Arrow result.

Raises:

Type Description
ConnectorError

If the connector is not connected, the job fails, or it returns no Arrow batches.

Source code in src/veridelta/connectors/warehouse.py
def execute_pushdown(
    self, statement: str, query_type: PushdownQueryType = "mismatch"
) -> pl.LazyFrame:
    """Execute compiler SQL as a BigQuery job and return a LazyFrame.

    Args:
        statement (str): SQL produced by `SQLPushdownCompiler`.
        query_type (PushdownQueryType): Which comparison round-trip this
            statement represents; recorded in the log line for the call.

    Returns:
        pl.LazyFrame: Unevaluated frame wrapped around the Arrow result.

    Raises:
        ConnectorError: If the connector is not connected, the job fails,
            or it returns no Arrow batches.
    """
    self._require_client()
    batches = _run_bigquery_query(
        self._client, statement, self._job_config, query_type=query_type
    )
    return _lazy_from_arrow(batches)

DatabaseConnector

Bases: ReaderConnector

Read one database table or query into Polars through ConnectorX.

connect() reads eagerly and keeps the frame, and lazyframe() hands it to the local engine as a LazyFrame over those rows. The read is the one place the rows are fetched, so it happens once per connect(). With partition_on set, ConnectorX splits that read into ranges over parallel connections, after Veridelta confirms the column holds no NULL and reads its lowest and highest values. The comparison runs in Polars, never in the database.

Source code in src/veridelta/connectors/database.py
class DatabaseConnector(ReaderConnector):
    """Read one database table or query into Polars through ConnectorX.

    `connect()` reads eagerly and keeps the frame, and `lazyframe()` hands it to
    the local engine as a LazyFrame over those rows. The read is the one place
    the rows are fetched, so it happens once per `connect()`. With
    `partition_on` set, ConnectorX splits that read into ranges over parallel
    connections, after Veridelta confirms the column holds no NULL and reads
    its lowest and highest values. The comparison runs in Polars, never in the
    database.
    """

    _unconnected = _UNCONNECTED

    def __init__(self, config: DatabaseConfig, *, probe: bool = False) -> None:
        """Initialize the connector with validated database settings.

        Args:
            config (DatabaseConfig): Frozen URI, credentials, and table or query.
            probe (bool): Whether to read the table's columns and no rows, for a
                schema check. Only a `table` can be probed.
        """
        self._config = config
        self._probe = probe

    def connect(self) -> None:
        """Read the configured table or query into memory.

        Raises:
            ConnectorError: If the `database` extra is missing, a SQLite file
                does not exist, the partition column holds a NULL or anything
                but integers, or the read fails.
            ConfigError: If `table` names a database Veridelta cannot quote for,
                or a probe was asked of a `query`.
        """
        if connectorx is None:
            raise ConnectorError(_DATABASE_EXTRA)
        scheme = urlsplit(self._config.uri).scheme.lower()
        statement = self._statement(scheme)
        uri = _connection_uri(self._config)
        if scheme == "sqlite":
            uri = _existing_sqlite_uri(uri)

        started = time.perf_counter()
        declared: dict[str, pl.Decimal] = {}
        try:
            # A Postgres `table` read keeps each `numeric` column's declared scale.
            if scheme in POSTGRES_SCHEMES and self._config.table is not None:
                statement, declared = self._declared_statement(self._config.table, statement, uri)
            elif scheme == _MSSQL_SCHEME and self._config.table is not None and not self._probe:
                statement = self._utc_statement(self._config.table, statement, uri)
            frame = self._read(scheme, statement, uri)
        except ConnectorError:
            raise
        except ImportError:
            # A missing pyarrow surfaces here, and it is the same missing extra.
            raise ConnectorError(_DATABASE_EXTRA) from None
        except Exception as exc:
            logger.warning(
                "Database read of %s from %s failed after %.3fs",
                self._subject,
                self._config.redacted_uri,
                time.perf_counter() - started,
            )
            raise ConnectorError(
                f"Database read of {self._subject} from '{self._config.redacted_uri}' "
                f"failed: {_scrub(self._config, str(exc))}"
            ) from None
        logger.info(
            "Read %d rows of %s from %s in %.3fs",
            frame.height,
            self._subject,
            self._config.redacted_uri,
            time.perf_counter() - started,
        )
        self._frame = _with_declared_scale(frame, declared, self._subject).lazy()

    def _statement(self, scheme: str) -> str:
        """Return the SQL to read: the compiled table select, or the query as written."""
        if self._config.table is not None and self._probe:
            return compile_database_probe(scheme, self._config.table)
        if self._config.table is not None:
            return compile_database_select(scheme, self._config.table)
        if self._probe:
            # Wrapping a statement is not portable: Oracle refuses `AS` on a
            # derived table, and SQL Server refuses `ORDER BY` inside one.
            raise ConfigError(PROBE_NEEDS_A_TABLE)
        # DatabaseConfig requires exactly one of `table` and `query`.
        return cast("str", self._config.query)

    def _read(self, scheme: str, statement: str, uri: str) -> pl.DataFrame:
        """Run the read, split into ranges over parallel connections when configured."""
        column = self._config.partition_on
        if column is None or self._probe:
            return pl.read_database_uri(statement, uri)
        # The model sets `table` and `partitions` whenever `partition_on` is set.
        table, partitions = cast("str", self._config.table), cast("int", self._config.partitions)
        counts = compile_database_partition_counts(scheme, table, column)
        nulls, valued = pl.read_database_uri(counts, uri).row(0)
        if nulls:
            # ConnectorX reads only rows inside its ranges, and a NULL is in none of them.
            raise ConnectorError(
                f"Column '{column}' of {self._subject} is NULL in {nulls:,} of its rows, "
                "which a partitioned read would leave out. Partition on a column without "
                "NULLs, or remove 'partition_on' and 'partitions'."
            )
        if not valued:
            # An empty table has no range to split, and one read keeps its columns.
            return pl.read_database_uri(statement, uri)
        bounds = compile_database_partition_range(scheme, table, column)
        low, high = pl.read_database_uri(bounds, uri).row(0)
        if not (isinstance(low, int) and isinstance(high, int)):
            raise ConnectorError(
                f"Column '{column}' of {self._subject} does not hold integers, so the read "
                "cannot be split on it. Partition on an integer column."
            )
        # Passing the range skips ConnectorX's own lookup, whose SQLite version opens the
        # path without percent-decoding it: every path on Windows, or one with a space.
        return pl.read_database_uri(
            statement,
            uri,
            partition_on=column,
            partition_num=partitions,
            partition_range=(low, high),
        )

    def _declared_statement(
        self, table: str, statement: str, uri: str
    ) -> tuple[str, dict[str, pl.Decimal]]:
        """Look up a Postgres table's `numeric` columns and read them as text, to keep scale."""
        catalog = pl.read_database_uri(compile_postgres_columns_query(table), uri)
        declared = _declared_decimals(catalog)
        if not declared:
            return statement, declared
        columns = catalog.get_column("attname").to_list()
        return compile_postgres_text_select(table, columns, declared, probe=self._probe), declared

    def _utc_statement(self, table: str, statement: str, uri: str) -> str:
        """Find a SQL Server table's `DATETIMEOFFSET` columns, and read them at offset zero.

        ConnectorX shifts a `DATETIMEOFFSET` value by its offset a second time.
        A read of the table's columns and no rows finds them, as the only
        columns ConnectorX reads as a time in UTC. The read itself would type
        them the same way, which a catalog query could only approximate.
        """
        columns = pl.read_database_uri(compile_database_probe(_MSSQL_SCHEME, table), uri).schema
        in_utc = [
            name
            for name, dtype in columns.items()
            if isinstance(dtype, pl.Datetime) and dtype.time_zone is not None
        ]
        if not in_utc:
            return statement
        bracketed = [name for name in columns if "]" in name]
        if bracketed and self._config.partition_on is not None:
            # ConnectorX parses a partitioned read to split it. Its parser reads a
            # doubled `]` as one and writes it back undoubled, which SQL Server refuses.
            raise ConnectorError(
                f"Column '{bracketed[0]}' of {self._subject} has ']' in its name, which a "
                "partitioned read cannot pass through ConnectorX. Remove 'partition_on' and "
                "'partitions' to read the table in one piece."
            )
        return compile_mssql_utc_select(table, list(columns), in_utc)

    @property
    def _subject(self) -> str:
        """Name what is read without repeating any SQL."""
        return read_subject(self._config.table)

__init__(config, *, probe=False)

Initialize the connector with validated database settings.

Parameters:

Name Type Description Default
config DatabaseConfig

Frozen URI, credentials, and table or query.

required
probe bool

Whether to read the table's columns and no rows, for a schema check. Only a table can be probed.

False
Source code in src/veridelta/connectors/database.py
def __init__(self, config: DatabaseConfig, *, probe: bool = False) -> None:
    """Initialize the connector with validated database settings.

    Args:
        config (DatabaseConfig): Frozen URI, credentials, and table or query.
        probe (bool): Whether to read the table's columns and no rows, for a
            schema check. Only a `table` can be probed.
    """
    self._config = config
    self._probe = probe

connect()

Read the configured table or query into memory.

Raises:

Type Description
ConnectorError

If the database extra is missing, a SQLite file does not exist, the partition column holds a NULL or anything but integers, or the read fails.

ConfigError

If table names a database Veridelta cannot quote for, or a probe was asked of a query.

Source code in src/veridelta/connectors/database.py
def connect(self) -> None:
    """Read the configured table or query into memory.

    Raises:
        ConnectorError: If the `database` extra is missing, a SQLite file
            does not exist, the partition column holds a NULL or anything
            but integers, or the read fails.
        ConfigError: If `table` names a database Veridelta cannot quote for,
            or a probe was asked of a `query`.
    """
    if connectorx is None:
        raise ConnectorError(_DATABASE_EXTRA)
    scheme = urlsplit(self._config.uri).scheme.lower()
    statement = self._statement(scheme)
    uri = _connection_uri(self._config)
    if scheme == "sqlite":
        uri = _existing_sqlite_uri(uri)

    started = time.perf_counter()
    declared: dict[str, pl.Decimal] = {}
    try:
        # A Postgres `table` read keeps each `numeric` column's declared scale.
        if scheme in POSTGRES_SCHEMES and self._config.table is not None:
            statement, declared = self._declared_statement(self._config.table, statement, uri)
        elif scheme == _MSSQL_SCHEME and self._config.table is not None and not self._probe:
            statement = self._utc_statement(self._config.table, statement, uri)
        frame = self._read(scheme, statement, uri)
    except ConnectorError:
        raise
    except ImportError:
        # A missing pyarrow surfaces here, and it is the same missing extra.
        raise ConnectorError(_DATABASE_EXTRA) from None
    except Exception as exc:
        logger.warning(
            "Database read of %s from %s failed after %.3fs",
            self._subject,
            self._config.redacted_uri,
            time.perf_counter() - started,
        )
        raise ConnectorError(
            f"Database read of {self._subject} from '{self._config.redacted_uri}' "
            f"failed: {_scrub(self._config, str(exc))}"
        ) from None
    logger.info(
        "Read %d rows of %s from %s in %.3fs",
        frame.height,
        self._subject,
        self._config.redacted_uri,
        time.perf_counter() - started,
    )
    self._frame = _with_declared_scale(frame, declared, self._subject).lazy()

DatabricksConnector

Bases: _CursorSession

Databricks SQL warehouse connector backed by the optional Databricks extra.

connect() opens a databricks.sql session against the configured SQL warehouse HTTP path; execute_pushdown runs compiler SQL on a fresh cursor and fetches the result as Arrow. Install the driver with uv add 'veridelta[databricks]'; without it, connect() raises ConnectorError with that hint instead of an ImportError.

Attributes:

Name Type Description
compiler SQLPushdownCompiler

Databricks-dialect compiler (backtick quoting, Spark type names) the engine uses for every statement.

Source code in src/veridelta/connectors/warehouse.py
class DatabricksConnector(_CursorSession):
    """Databricks SQL warehouse connector backed by the optional Databricks extra.

    `connect()` opens a `databricks.sql` session against the configured SQL
    warehouse HTTP path; `execute_pushdown` runs compiler SQL on a fresh cursor
    and fetches the result as Arrow. Install the driver with
    `uv add 'veridelta[databricks]'`; without it, `connect()` raises
    `ConnectorError` with that hint instead of an `ImportError`.

    Attributes:
        compiler (SQLPushdownCompiler): Databricks-dialect compiler (backtick
            quoting, Spark type names) the engine uses for every statement.
    """

    _backend = "Databricks"
    _fetch = "fetchall_arrow"

    def __init__(self, config: DatabricksConfig) -> None:
        """Initialize the connector with validated Databricks settings.

        Args:
            config (DatabricksConfig): Frozen workspace hostname and HTTP path.
        """
        self._config = config
        self.compiler = SQLPushdownCompiler(SQLDialect.DATABRICKS)

    def connect(self) -> None:
        """Open a Databricks SQL session for subsequent pushdown statements.

        Raises:
            ConnectorError: If the Databricks extra is missing or authentication
                fails.
        """
        missing = self._missing_extra()
        if missing is not None:
            raise ConnectorError(missing)
        try:
            self._session = databricks_sql.connect(
                server_hostname=self._config.server_hostname,
                http_path=self._config.http_path,
                access_token=self._config.access_token,
                catalog=self._config.catalog,
                schema=self._config.schema_name,
            )
        except ConnectorError:
            raise
        except Exception as exc:
            logger.warning("Databricks connection to %s failed", self._config.server_hostname)
            # Not chained: a traceback would print the driver's message unmasked.
            message = mask_secrets(str(exc), self._config.access_token)
            raise ConnectorError(f"Failed to connect to Databricks: {message}") from None
        logger.info(
            "Connected to Databricks host %s, path %s",
            self._config.server_hostname,
            self._config.http_path,
        )

    def _missing_extra(self) -> str | None:
        """Return the Databricks install hint when its driver is not installed."""
        return _DATABRICKS_EXTRA if databricks_sql is None else None

__init__(config)

Initialize the connector with validated Databricks settings.

Parameters:

Name Type Description Default
config DatabricksConfig

Frozen workspace hostname and HTTP path.

required
Source code in src/veridelta/connectors/warehouse.py
def __init__(self, config: DatabricksConfig) -> None:
    """Initialize the connector with validated Databricks settings.

    Args:
        config (DatabricksConfig): Frozen workspace hostname and HTTP path.
    """
    self._config = config
    self.compiler = SQLPushdownCompiler(SQLDialect.DATABRICKS)

connect()

Open a Databricks SQL session for subsequent pushdown statements.

Raises:

Type Description
ConnectorError

If the Databricks extra is missing or authentication fails.

Source code in src/veridelta/connectors/warehouse.py
def connect(self) -> None:
    """Open a Databricks SQL session for subsequent pushdown statements.

    Raises:
        ConnectorError: If the Databricks extra is missing or authentication
            fails.
    """
    missing = self._missing_extra()
    if missing is not None:
        raise ConnectorError(missing)
    try:
        self._session = databricks_sql.connect(
            server_hostname=self._config.server_hostname,
            http_path=self._config.http_path,
            access_token=self._config.access_token,
            catalog=self._config.catalog,
            schema=self._config.schema_name,
        )
    except ConnectorError:
        raise
    except Exception as exc:
        logger.warning("Databricks connection to %s failed", self._config.server_hostname)
        # Not chained: a traceback would print the driver's message unmasked.
        message = mask_secrets(str(exc), self._config.access_token)
        raise ConnectorError(f"Failed to connect to Databricks: {message}") from None
    logger.info(
        "Connected to Databricks host %s, path %s",
        self._config.server_hostname,
        self._config.http_path,
    )

DeltaLakeConnector

Bases: ReaderConnector

Delta Lake scanner backed by pl.scan_delta.

connect() opens a lazy scan of DeltaLakeConfig.table_uri, pinned to version when one is set, and lazyframe() hands that scan to the local engine. No SQL is involved: the comparison runs in Polars. Requires the delta extra (uv add 'veridelta[delta]').

Source code in src/veridelta/connectors/lakehouse.py
class DeltaLakeConnector(ReaderConnector):
    """Delta Lake scanner backed by `pl.scan_delta`.

    `connect()` opens a lazy scan of `DeltaLakeConfig.table_uri`, pinned to
    `version` when one is set, and `lazyframe()` hands that scan to the local
    engine. No SQL is involved: the comparison runs in Polars. Requires the
    `delta` extra (`uv add 'veridelta[delta]'`).
    """

    _unconnected = _UNCONNECTED

    def __init__(self, config: DeltaLakeConfig) -> None:
        """Initialize the connector with validated Delta Lake settings.

        Args:
            config (DeltaLakeConfig): Frozen table URI and optional version.
        """
        self._config = config

    def connect(self) -> None:
        """Open a lazy `pl.scan_delta` of the configured table and read its log.

        Raises:
            ConnectorError: If the `deltalake` extra is missing, or the table
                or version cannot be read.
        """
        storage_options = self._config.storage_options or None
        try:
            frame = pl.scan_delta(
                self._config.table_uri,
                version=self._config.version,
                storage_options=storage_options,
            )
            # Polars defers the scan, so read the log now: a missing table or
            # version then fails here, named, instead of partway through a run.
            frame.collect_schema()
        except ImportError as exc:
            raise ConnectorError(_DELTA_EXTRA) from exc
        except Exception as exc:
            where = shown_location(self._config.table_uri)
            logger.warning("Delta Lake scan of %s failed", where)
            raise ConnectorError(
                f"Delta Lake scan of '{where}' failed: "
                f"{without_location(str(exc), self._config.table_uri)}"
            ) from exc
        self._frame = frame
        logger.info(
            "Opened Delta Lake scan of %s (version=%s)",
            shown_location(self._config.table_uri),
            "latest" if self._config.version is None else self._config.version,
        )

__init__(config)

Initialize the connector with validated Delta Lake settings.

Parameters:

Name Type Description Default
config DeltaLakeConfig

Frozen table URI and optional version.

required
Source code in src/veridelta/connectors/lakehouse.py
def __init__(self, config: DeltaLakeConfig) -> None:
    """Initialize the connector with validated Delta Lake settings.

    Args:
        config (DeltaLakeConfig): Frozen table URI and optional version.
    """
    self._config = config

connect()

Open a lazy pl.scan_delta of the configured table and read its log.

Raises:

Type Description
ConnectorError

If the deltalake extra is missing, or the table or version cannot be read.

Source code in src/veridelta/connectors/lakehouse.py
def connect(self) -> None:
    """Open a lazy `pl.scan_delta` of the configured table and read its log.

    Raises:
        ConnectorError: If the `deltalake` extra is missing, or the table
            or version cannot be read.
    """
    storage_options = self._config.storage_options or None
    try:
        frame = pl.scan_delta(
            self._config.table_uri,
            version=self._config.version,
            storage_options=storage_options,
        )
        # Polars defers the scan, so read the log now: a missing table or
        # version then fails here, named, instead of partway through a run.
        frame.collect_schema()
    except ImportError as exc:
        raise ConnectorError(_DELTA_EXTRA) from exc
    except Exception as exc:
        where = shown_location(self._config.table_uri)
        logger.warning("Delta Lake scan of %s failed", where)
        raise ConnectorError(
            f"Delta Lake scan of '{where}' failed: "
            f"{without_location(str(exc), self._config.table_uri)}"
        ) from exc
    self._frame = frame
    logger.info(
        "Opened Delta Lake scan of %s (version=%s)",
        shown_location(self._config.table_uri),
        "latest" if self._config.version is None else self._config.version,
    )

DuckDBConnector

Bases: ReaderConnector

Read one DuckDB or MotherDuck table or query into Polars.

connect() reads eagerly and keeps the frame, and lazyframe() hands it to the local engine as a LazyFrame over those rows. The read is the one place the rows are fetched, so it happens once per connect(). The comparison runs in Polars, never in DuckDB.

Source code in src/veridelta/connectors/duckdb.py
class DuckDBConnector(ReaderConnector):
    """Read one DuckDB or MotherDuck table or query into Polars.

    `connect()` reads eagerly and keeps the frame, and `lazyframe()` hands it to
    the local engine as a LazyFrame over those rows. The read is the one place
    the rows are fetched, so it happens once per `connect()`. The comparison
    runs in Polars, never in DuckDB.
    """

    _unconnected = _UNCONNECTED

    def __init__(self, config: DuckDBConfig, *, probe: bool = False) -> None:
        """Initialize the connector with validated DuckDB settings.

        Args:
            config (DuckDBConfig): Frozen database, token, and table or query.
            probe (bool): Whether to read the table's columns and no rows, for a
                schema check. Only a `table` can be probed.
        """
        self._config = config
        self._probe = probe

    def connect(self) -> None:
        """Read the configured table or query into memory.

        Raises:
            ConnectorError: If the `duckdb` extra is missing, a MotherDuck
                database has no token, a column has no Polars type, or the
                read fails.
            ConfigError: If a probe was asked of a `query`.
        """
        if duckdb is None:
            raise ConnectorError(_DUCKDB_EXTRA)
        statement = self._statement()
        token = _token(self._config)
        started = time.perf_counter()
        try:
            connection = _open(self._config, token)
            try:
                frame = _read(connection, statement, self._subject)
            finally:
                connection.close()
        except ConnectorError:
            raise
        except ImportError:
            # A missing pyarrow surfaces here, and it is the same missing extra.
            raise ConnectorError(_DUCKDB_EXTRA) from None
        except Exception as exc:
            logger.warning(
                "DuckDB read of %s from %s failed after %.3fs",
                self._subject,
                self._config.database,
                time.perf_counter() - started,
            )
            raise ConnectorError(
                f"DuckDB read of {self._subject} from '{self._config.database}' failed: "
                f"{mask_secrets(str(exc), token)}"
            ) from None
        logger.info(
            "Read %d rows of %s from %s in %.3fs",
            frame.height,
            self._subject,
            self._config.database,
            time.perf_counter() - started,
        )
        self._frame = frame.lazy()

    def _statement(self) -> str:
        """Return the SQL to read: the compiled table select, or the query as written."""
        if self._config.table is not None:
            return compile_duckdb_select(self._config.table, probe=self._probe)
        if self._probe:
            raise ConfigError(PROBE_NEEDS_A_TABLE)
        # DuckDBConfig requires exactly one of `table` and `query`.
        return cast("str", self._config.query)

    @property
    def _subject(self) -> str:
        """Name what is read without repeating any SQL."""
        return read_subject(self._config.table)

__init__(config, *, probe=False)

Initialize the connector with validated DuckDB settings.

Parameters:

Name Type Description Default
config DuckDBConfig

Frozen database, token, and table or query.

required
probe bool

Whether to read the table's columns and no rows, for a schema check. Only a table can be probed.

False
Source code in src/veridelta/connectors/duckdb.py
def __init__(self, config: DuckDBConfig, *, probe: bool = False) -> None:
    """Initialize the connector with validated DuckDB settings.

    Args:
        config (DuckDBConfig): Frozen database, token, and table or query.
        probe (bool): Whether to read the table's columns and no rows, for a
            schema check. Only a `table` can be probed.
    """
    self._config = config
    self._probe = probe

connect()

Read the configured table or query into memory.

Raises:

Type Description
ConnectorError

If the duckdb extra is missing, a MotherDuck database has no token, a column has no Polars type, or the read fails.

ConfigError

If a probe was asked of a query.

Source code in src/veridelta/connectors/duckdb.py
def connect(self) -> None:
    """Read the configured table or query into memory.

    Raises:
        ConnectorError: If the `duckdb` extra is missing, a MotherDuck
            database has no token, a column has no Polars type, or the
            read fails.
        ConfigError: If a probe was asked of a `query`.
    """
    if duckdb is None:
        raise ConnectorError(_DUCKDB_EXTRA)
    statement = self._statement()
    token = _token(self._config)
    started = time.perf_counter()
    try:
        connection = _open(self._config, token)
        try:
            frame = _read(connection, statement, self._subject)
        finally:
            connection.close()
    except ConnectorError:
        raise
    except ImportError:
        # A missing pyarrow surfaces here, and it is the same missing extra.
        raise ConnectorError(_DUCKDB_EXTRA) from None
    except Exception as exc:
        logger.warning(
            "DuckDB read of %s from %s failed after %.3fs",
            self._subject,
            self._config.database,
            time.perf_counter() - started,
        )
        raise ConnectorError(
            f"DuckDB read of {self._subject} from '{self._config.database}' failed: "
            f"{mask_secrets(str(exc), token)}"
        ) from None
    logger.info(
        "Read %d rows of %s from %s in %.3fs",
        frame.height,
        self._subject,
        self._config.database,
        time.perf_counter() - started,
    )
    self._frame = frame.lazy()

DuckDBPushdownSession

Bases: PushdownSession

Run compiled comparison SQL inside DuckDB or MotherDuck.

Opened for two tables in one database that both set pushdown. It holds one connection from connect() to close(), opened as a read opens it: a file read-only, MotherDuck read-write with its token, both in UTC. Only counts and keys come back, and a row sample when one is asked for.

Source code in src/veridelta/connectors/duckdb.py
class DuckDBPushdownSession(PushdownSession):
    """Run compiled comparison SQL inside DuckDB or MotherDuck.

    Opened for two tables in one database that both set `pushdown`. It holds
    one connection from `connect()` to `close()`, opened as a read opens it:
    a file read-only, MotherDuck read-write with its token, both in UTC. Only
    counts and keys come back, and a row sample when one is asked for.
    """

    def __init__(self, config: DuckDBConfig) -> None:
        """Initialize the session for one side's connection settings.

        Args:
            config (DuckDBConfig): A DuckDB table that sets `pushdown`.
        """
        self._config = config
        self.compiler = SQLPushdownCompiler(SQLDialect.DUCKDB)
        self._connection: Any = None
        self._token: str | None = None

    def connect(self) -> None:
        """Open the database for the statements to come.

        Raises:
            ConnectorError: If the `duckdb` extra is missing, a MotherDuck
                database has no token, or the database cannot be opened.
        """
        if duckdb is None:
            raise ConnectorError(_DUCKDB_EXTRA)
        token = _token(self._config)
        try:
            connection = _open(self._config, token)
        except Exception as exc:
            logger.warning("DuckDB connection to %s failed", self._config.database)
            raise ConnectorError(
                f"DuckDB connection to '{self._config.database}' failed: "
                f"{mask_secrets(str(exc), token)}"
            ) from None
        logger.info("Connected to DuckDB database %s", self._config.database)
        self._connection, self._token = connection, token

    def execute_pushdown(
        self, statement: str, query_type: PushdownQueryType = "mismatch"
    ) -> pl.LazyFrame:
        """Run one compiled statement and return its rows lazily.

        Args:
            statement (str): SQL from the DuckDB compiler.
            query_type (PushdownQueryType): Which round-trip this is, for logs
                and errors.

        Returns:
            pl.LazyFrame: The statement's result.

        Raises:
            ConnectorError: If the session is not connected, a result column
                has no Polars type, or the statement fails.
        """
        frame = self._run(statement, query_type)
        return frame.lazy()

    def close(self) -> None:
        """Close the connection. Idempotent; `connect()` opens it again."""
        if self._connection is not None:
            self._connection.close()
            logger.info("Closed DuckDB database %s", self._config.database)
        self._connection = None

    def _run(self, statement: str, query_type: str) -> pl.DataFrame:
        """Run a statement on the open connection, logging and reporting without secrets."""
        if self._connection is None:
            raise ConnectorError(_SESSION_UNCONNECTED)
        started = time.perf_counter()
        try:
            frame = _read(self._connection, statement, f"the {query_type} statement")
        except ConnectorError:
            raise
        except ImportError:
            raise ConnectorError(_DUCKDB_EXTRA) from None
        except Exception as exc:
            logger.warning(
                "DuckDB %s statement on %s failed after %.3fs",
                query_type,
                self._config.database,
                time.perf_counter() - started,
            )
            raise ConnectorError(
                f"DuckDB {query_type} statement on '{self._config.database}' failed: "
                f"{mask_secrets(str(exc), self._token)}"
            ) from None
        logger.info(
            "Ran DuckDB %s statement on %s in %.3fs, %d rows",
            query_type,
            self._config.database,
            time.perf_counter() - started,
            frame.height,
        )
        return frame

__init__(config)

Initialize the session for one side's connection settings.

Parameters:

Name Type Description Default
config DuckDBConfig

A DuckDB table that sets pushdown.

required
Source code in src/veridelta/connectors/duckdb.py
def __init__(self, config: DuckDBConfig) -> None:
    """Initialize the session for one side's connection settings.

    Args:
        config (DuckDBConfig): A DuckDB table that sets `pushdown`.
    """
    self._config = config
    self.compiler = SQLPushdownCompiler(SQLDialect.DUCKDB)
    self._connection: Any = None
    self._token: str | None = None

close()

Close the connection. Idempotent; connect() opens it again.

Source code in src/veridelta/connectors/duckdb.py
def close(self) -> None:
    """Close the connection. Idempotent; `connect()` opens it again."""
    if self._connection is not None:
        self._connection.close()
        logger.info("Closed DuckDB database %s", self._config.database)
    self._connection = None

connect()

Open the database for the statements to come.

Raises:

Type Description
ConnectorError

If the duckdb extra is missing, a MotherDuck database has no token, or the database cannot be opened.

Source code in src/veridelta/connectors/duckdb.py
def connect(self) -> None:
    """Open the database for the statements to come.

    Raises:
        ConnectorError: If the `duckdb` extra is missing, a MotherDuck
            database has no token, or the database cannot be opened.
    """
    if duckdb is None:
        raise ConnectorError(_DUCKDB_EXTRA)
    token = _token(self._config)
    try:
        connection = _open(self._config, token)
    except Exception as exc:
        logger.warning("DuckDB connection to %s failed", self._config.database)
        raise ConnectorError(
            f"DuckDB connection to '{self._config.database}' failed: "
            f"{mask_secrets(str(exc), token)}"
        ) from None
    logger.info("Connected to DuckDB database %s", self._config.database)
    self._connection, self._token = connection, token

execute_pushdown(statement, query_type='mismatch')

Run one compiled statement and return its rows lazily.

Parameters:

Name Type Description Default
statement str

SQL from the DuckDB compiler.

required
query_type PushdownQueryType

Which round-trip this is, for logs and errors.

'mismatch'

Returns:

Type Description
LazyFrame

pl.LazyFrame: The statement's result.

Raises:

Type Description
ConnectorError

If the session is not connected, a result column has no Polars type, or the statement fails.

Source code in src/veridelta/connectors/duckdb.py
def execute_pushdown(
    self, statement: str, query_type: PushdownQueryType = "mismatch"
) -> pl.LazyFrame:
    """Run one compiled statement and return its rows lazily.

    Args:
        statement (str): SQL from the DuckDB compiler.
        query_type (PushdownQueryType): Which round-trip this is, for logs
            and errors.

    Returns:
        pl.LazyFrame: The statement's result.

    Raises:
        ConnectorError: If the session is not connected, a result column
            has no Polars type, or the statement fails.
    """
    frame = self._run(statement, query_type)
    return frame.lazy()

IcebergConnector

Bases: ReaderConnector

Apache Iceberg scanner backed by pl.scan_iceberg.

connect() opens a lazy scan of IcebergConfig.table_uri, pinned to snapshot_id when one is set, and lazyframe() hands that scan to the local engine. As with Delta Lake, the comparison runs in Polars. Requires the iceberg extra (uv add 'veridelta[iceberg]').

Source code in src/veridelta/connectors/lakehouse.py
class IcebergConnector(ReaderConnector):
    """Apache Iceberg scanner backed by `pl.scan_iceberg`.

    `connect()` opens a lazy scan of `IcebergConfig.table_uri`, pinned to
    `snapshot_id` when one is set, and `lazyframe()` hands that scan to the
    local engine. As with Delta Lake, the comparison runs in Polars. Requires
    the `iceberg` extra (`uv add 'veridelta[iceberg]'`).
    """

    _unconnected = _UNCONNECTED

    def __init__(self, config: IcebergConfig) -> None:
        """Initialize the connector with validated Iceberg settings.

        Args:
            config (IcebergConfig): Frozen table URI and storage options.
        """
        self._config = config

    def connect(self) -> None:
        """Open a lazy `pl.scan_iceberg` of the configured table and read its metadata.

        Raises:
            ConnectorError: If the `pyiceberg` extra is missing, or the table
                or snapshot cannot be read.
        """
        try:
            frame = pl.scan_iceberg(
                self._config.table_uri,
                snapshot_id=self._config.snapshot_id,
                storage_options=self._config.storage_options or None,
            )
            # Polars defers the scan, so read the metadata now: a missing table
            # then fails here, named, instead of partway through a run. Polars
            # looks a snapshot up only to read rows, so time travel reads one.
            if self._config.snapshot_id is None:
                frame.collect_schema()
            else:
                frame.head(1).collect()
        except ImportError as exc:
            raise ConnectorError(_ICEBERG_EXTRA) from exc
        except Exception as exc:
            where = shown_location(self._config.table_uri)
            logger.warning("Iceberg scan of %s failed", where)
            raise ConnectorError(
                f"Iceberg scan of '{where}' failed: "
                f"{without_location(str(exc), self._config.table_uri)}"
            ) from exc
        self._frame = frame
        logger.info(
            "Opened Iceberg scan of %s (snapshot_id=%s)",
            shown_location(self._config.table_uri),
            "latest" if self._config.snapshot_id is None else self._config.snapshot_id,
        )

__init__(config)

Initialize the connector with validated Iceberg settings.

Parameters:

Name Type Description Default
config IcebergConfig

Frozen table URI and storage options.

required
Source code in src/veridelta/connectors/lakehouse.py
def __init__(self, config: IcebergConfig) -> None:
    """Initialize the connector with validated Iceberg settings.

    Args:
        config (IcebergConfig): Frozen table URI and storage options.
    """
    self._config = config

connect()

Open a lazy pl.scan_iceberg of the configured table and read its metadata.

Raises:

Type Description
ConnectorError

If the pyiceberg extra is missing, or the table or snapshot cannot be read.

Source code in src/veridelta/connectors/lakehouse.py
def connect(self) -> None:
    """Open a lazy `pl.scan_iceberg` of the configured table and read its metadata.

    Raises:
        ConnectorError: If the `pyiceberg` extra is missing, or the table
            or snapshot cannot be read.
    """
    try:
        frame = pl.scan_iceberg(
            self._config.table_uri,
            snapshot_id=self._config.snapshot_id,
            storage_options=self._config.storage_options or None,
        )
        # Polars defers the scan, so read the metadata now: a missing table
        # then fails here, named, instead of partway through a run. Polars
        # looks a snapshot up only to read rows, so time travel reads one.
        if self._config.snapshot_id is None:
            frame.collect_schema()
        else:
            frame.head(1).collect()
    except ImportError as exc:
        raise ConnectorError(_ICEBERG_EXTRA) from exc
    except Exception as exc:
        where = shown_location(self._config.table_uri)
        logger.warning("Iceberg scan of %s failed", where)
        raise ConnectorError(
            f"Iceberg scan of '{where}' failed: "
            f"{without_location(str(exc), self._config.table_uri)}"
        ) from exc
    self._frame = frame
    logger.info(
        "Opened Iceberg scan of %s (snapshot_id=%s)",
        shown_location(self._config.table_uri),
        "latest" if self._config.snapshot_id is None else self._config.snapshot_id,
    )

PostgresPushdownSession

Bases: PushdownSession

Run compiled comparison SQL inside Postgres, through ConnectorX.

Opened for two database sources on one Postgres connection that both set pushdown. Each statement is one ConnectorX read, which opens its own connection and returns Arrow, so nothing is held between statements and close() only forgets the connection. Only counts and keys come back.

connect() checks that the server reads string literals by the SQL standard, as the Postgres dialect writes them: with standard_conforming_strings off, a backslash would escape the closing quote of a value that ends in one.

Source code in src/veridelta/connectors/database.py
class PostgresPushdownSession(PushdownSession):
    """Run compiled comparison SQL inside Postgres, through ConnectorX.

    Opened for two database sources on one Postgres connection that both set
    `pushdown`. Each statement is one ConnectorX read, which opens its own
    connection and returns Arrow, so nothing is held between statements and
    `close()` only forgets the connection. Only counts and keys come back.

    `connect()` checks that the server reads string literals by the SQL
    standard, as the Postgres dialect writes them: with
    `standard_conforming_strings` off, a backslash would escape the closing
    quote of a value that ends in one.
    """

    def __init__(self, config: DatabaseConfig) -> None:
        """Initialize the session for one side's connection settings.

        Args:
            config (DatabaseConfig): A Postgres table that sets `pushdown`.
        """
        self._config = config
        self.compiler = SQLPushdownCompiler(SQLDialect.POSTGRES)
        self._uri: str | None = None

    def connect(self) -> None:
        """Check the server and keep the connection URI for the statements to come.

        Raises:
            ConnectorError: If the `database` extra is missing, the server cannot
                be reached, or it reads backslashes in string literals as escapes.
        """
        if connectorx is None:
            raise ConnectorError(_DATABASE_EXTRA)
        uri = _connection_uri(self._config)
        setting = self._read(_LITERAL_RULES, uri, query_type="settings")
        if setting.item(0, 0) != "on":
            raise ConnectorError(
                f"standard_conforming_strings is off on Postgres at "
                f"'{self._config.redacted_uri}', so it would read a backslash in a string "
                "literal as an escape. Turn it on for the role or database, or compare "
                "locally (pushdown: false)."
            )
        self._uri = uri

    def execute_pushdown(
        self, statement: str, query_type: PushdownQueryType = "mismatch"
    ) -> pl.LazyFrame:
        """Run one compiled statement and return its rows lazily.

        Args:
            statement (str): SQL from the Postgres compiler.
            query_type (PushdownQueryType): Which round-trip this is, for logs
                and errors.

        Returns:
            pl.LazyFrame: The statement's result.

        Raises:
            ConnectorError: If the session is not connected or the statement fails.
        """
        frame = self._read(statement, self._connected_uri(), query_type=query_type)
        return frame.lazy()

    def declared_types(self, table: str) -> dict[str, pl.Decimal]:
        """Return the declared precision and scale of a table's `numeric` columns.

        ConnectorX describes every `numeric` as `Decimal(38, 10)`, so a schema
        probe alone cannot tell `numeric(10, 2)` from `numeric(12, 4)`.

        Args:
            table (str): The table, as configured.

        Returns:
            dict[str, pl.Decimal]: Each `numeric` column Polars holds exactly,
                by name. A column without a declared precision, wider than 38
                digits, or with a negative scale is left out.

        Raises:
            ConnectorError: If the session is not connected or the query fails.
        """
        statement = compile_postgres_columns_query(table)
        return _declared_decimals(self._read(statement, self._connected_uri(), query_type="schema"))

    def close(self) -> None:
        """Forget the connection. Idempotent; `connect()` checks the server again."""
        self._uri = None

    def _connected_uri(self) -> str:
        """Return the URI `connect()` kept."""
        if self._uri is None:
            raise ConnectorError(_POSTGRES_UNCONNECTED)
        return self._uri

    def _read(self, statement: str, uri: str, *, query_type: str) -> pl.DataFrame:
        """Run a statement through ConnectorX, logging and reporting without secrets."""
        started = time.perf_counter()
        try:
            frame = pl.read_database_uri(statement, uri)
        except ImportError:
            raise ConnectorError(_DATABASE_EXTRA) from None
        except Exception as exc:
            logger.warning(
                "Postgres %s statement on %s failed after %.3fs",
                query_type,
                self._config.redacted_uri,
                time.perf_counter() - started,
            )
            raise ConnectorError(
                f"Postgres {query_type} statement on '{self._config.redacted_uri}' failed: "
                f"{_scrub(self._config, str(exc))}"
            ) from None
        logger.info(
            "Ran Postgres %s statement on %s in %.3fs, %d rows",
            query_type,
            self._config.redacted_uri,
            time.perf_counter() - started,
            frame.height,
        )
        return frame

__init__(config)

Initialize the session for one side's connection settings.

Parameters:

Name Type Description Default
config DatabaseConfig

A Postgres table that sets pushdown.

required
Source code in src/veridelta/connectors/database.py
def __init__(self, config: DatabaseConfig) -> None:
    """Initialize the session for one side's connection settings.

    Args:
        config (DatabaseConfig): A Postgres table that sets `pushdown`.
    """
    self._config = config
    self.compiler = SQLPushdownCompiler(SQLDialect.POSTGRES)
    self._uri: str | None = None

close()

Forget the connection. Idempotent; connect() checks the server again.

Source code in src/veridelta/connectors/database.py
def close(self) -> None:
    """Forget the connection. Idempotent; `connect()` checks the server again."""
    self._uri = None

connect()

Check the server and keep the connection URI for the statements to come.

Raises:

Type Description
ConnectorError

If the database extra is missing, the server cannot be reached, or it reads backslashes in string literals as escapes.

Source code in src/veridelta/connectors/database.py
def connect(self) -> None:
    """Check the server and keep the connection URI for the statements to come.

    Raises:
        ConnectorError: If the `database` extra is missing, the server cannot
            be reached, or it reads backslashes in string literals as escapes.
    """
    if connectorx is None:
        raise ConnectorError(_DATABASE_EXTRA)
    uri = _connection_uri(self._config)
    setting = self._read(_LITERAL_RULES, uri, query_type="settings")
    if setting.item(0, 0) != "on":
        raise ConnectorError(
            f"standard_conforming_strings is off on Postgres at "
            f"'{self._config.redacted_uri}', so it would read a backslash in a string "
            "literal as an escape. Turn it on for the role or database, or compare "
            "locally (pushdown: false)."
        )
    self._uri = uri

declared_types(table)

Return the declared precision and scale of a table's numeric columns.

ConnectorX describes every numeric as Decimal(38, 10), so a schema probe alone cannot tell numeric(10, 2) from numeric(12, 4).

Parameters:

Name Type Description Default
table str

The table, as configured.

required

Returns:

Type Description
dict[str, Decimal]

dict[str, pl.Decimal]: Each numeric column Polars holds exactly, by name. A column without a declared precision, wider than 38 digits, or with a negative scale is left out.

Raises:

Type Description
ConnectorError

If the session is not connected or the query fails.

Source code in src/veridelta/connectors/database.py
def declared_types(self, table: str) -> dict[str, pl.Decimal]:
    """Return the declared precision and scale of a table's `numeric` columns.

    ConnectorX describes every `numeric` as `Decimal(38, 10)`, so a schema
    probe alone cannot tell `numeric(10, 2)` from `numeric(12, 4)`.

    Args:
        table (str): The table, as configured.

    Returns:
        dict[str, pl.Decimal]: Each `numeric` column Polars holds exactly,
            by name. A column without a declared precision, wider than 38
            digits, or with a negative scale is left out.

    Raises:
        ConnectorError: If the session is not connected or the query fails.
    """
    statement = compile_postgres_columns_query(table)
    return _declared_decimals(self._read(statement, self._connected_uri(), query_type="schema"))

execute_pushdown(statement, query_type='mismatch')

Run one compiled statement and return its rows lazily.

Parameters:

Name Type Description Default
statement str

SQL from the Postgres compiler.

required
query_type PushdownQueryType

Which round-trip this is, for logs and errors.

'mismatch'

Returns:

Type Description
LazyFrame

pl.LazyFrame: The statement's result.

Raises:

Type Description
ConnectorError

If the session is not connected or the statement fails.

Source code in src/veridelta/connectors/database.py
def execute_pushdown(
    self, statement: str, query_type: PushdownQueryType = "mismatch"
) -> pl.LazyFrame:
    """Run one compiled statement and return its rows lazily.

    Args:
        statement (str): SQL from the Postgres compiler.
        query_type (PushdownQueryType): Which round-trip this is, for logs
            and errors.

    Returns:
        pl.LazyFrame: The statement's result.

    Raises:
        ConnectorError: If the session is not connected or the statement fails.
    """
    frame = self._read(statement, self._connected_uri(), query_type=query_type)
    return frame.lazy()

PushdownSession

Bases: VerideltaConnector

A connector that runs compiled comparison SQL where the data lives.

The engine compiles every statement with compiler and runs it through execute_pushdown, one round-trip per PushdownQueryType. Only what a statement returns comes back: counts, keys, and the samples asked for, never the tables themselves.

Attributes:

Name Type Description
compiler SQLPushdownCompiler

The dialect's compiler, which the subclass sets before connect().

Source code in src/veridelta/connectors/base.py
class PushdownSession(VerideltaConnector):
    """A connector that runs compiled comparison SQL where the data lives.

    The engine compiles every statement with `compiler` and runs it through
    `execute_pushdown`, one round-trip per `PushdownQueryType`. Only what a
    statement returns comes back: counts, keys, and the samples asked for,
    never the tables themselves.

    Attributes:
        compiler (SQLPushdownCompiler): The dialect's compiler, which the
            subclass sets before `connect()`.
    """

    compiler: SQLPushdownCompiler

    @abstractmethod
    def execute_pushdown(
        self, statement: str, query_type: PushdownQueryType = "mismatch"
    ) -> pl.LazyFrame:
        """Execute compiled SQL and return an unevaluated result graph.

        Args:
            statement (str): Compiler-produced SQL. Never user text.
            query_type (PushdownQueryType): Which round-trip the SQL represents
                (`mismatch`, `added`, `missing`, `count`, `duplicates`,
                `columns`, `schema`, `value_maps`, `settings`, or `samples`).

        Returns:
            pl.LazyFrame: Unevaluated result graph. Must not be collected here.

        Raises:
            ConnectorError: If the session is not connected or the statement
                fails.
        """

execute_pushdown(statement, query_type='mismatch') abstractmethod

Execute compiled SQL and return an unevaluated result graph.

Parameters:

Name Type Description Default
statement str

Compiler-produced SQL. Never user text.

required
query_type PushdownQueryType

Which round-trip the SQL represents (mismatch, added, missing, count, duplicates, columns, schema, value_maps, settings, or samples).

'mismatch'

Returns:

Type Description
LazyFrame

pl.LazyFrame: Unevaluated result graph. Must not be collected here.

Raises:

Type Description
ConnectorError

If the session is not connected or the statement fails.

Source code in src/veridelta/connectors/base.py
@abstractmethod
def execute_pushdown(
    self, statement: str, query_type: PushdownQueryType = "mismatch"
) -> pl.LazyFrame:
    """Execute compiled SQL and return an unevaluated result graph.

    Args:
        statement (str): Compiler-produced SQL. Never user text.
        query_type (PushdownQueryType): Which round-trip the SQL represents
            (`mismatch`, `added`, `missing`, `count`, `duplicates`,
            `columns`, `schema`, `value_maps`, `settings`, or `samples`).

    Returns:
        pl.LazyFrame: Unevaluated result graph. Must not be collected here.

    Raises:
        ConnectorError: If the session is not connected or the statement
            fails.
    """

ReaderConnector

Bases: VerideltaConnector

A connector the local engine reads through lazyframe().

connect() opens a scan or reads the rows and keeps them as _frame, lazyframe() hands them to the engine, and the comparison runs in Polars. close() drops them. A reader has no execute_pushdown: nothing is pushed into a source that is compared locally.

Source code in src/veridelta/connectors/base.py
class ReaderConnector(VerideltaConnector):
    """A connector the local engine reads through `lazyframe()`.

    `connect()` opens a scan or reads the rows and keeps them as `_frame`,
    `lazyframe()` hands them to the engine, and the comparison runs in Polars.
    `close()` drops them. A reader has no `execute_pushdown`: nothing is pushed
    into a source that is compared locally.
    """

    _unconnected: ClassVar[str] = "Connector is not connected. Call connect() first."
    """What `lazyframe()` raises before `connect()`, naming the kind of source."""

    _frame: pl.LazyFrame | None = None
    """What `connect()` opened, and None before it and after `close()`."""

    def lazyframe(self) -> pl.LazyFrame:
        """Return what `connect()` opened, as an unevaluated LazyFrame.

        Returns:
            pl.LazyFrame: A lazy scan, or a lazy wrapper over the rows read.

        Raises:
            ConnectorError: If `connect()` has not been called.
        """
        if self._frame is None:
            raise ConnectorError(self._unconnected)
        return self._frame

    def close(self) -> None:
        """Drop what `connect()` opened. Idempotent; `connect()` opens it again."""
        self._frame = None

close()

Drop what connect() opened. Idempotent; connect() opens it again.

Source code in src/veridelta/connectors/base.py
def close(self) -> None:
    """Drop what `connect()` opened. Idempotent; `connect()` opens it again."""
    self._frame = None

lazyframe()

Return what connect() opened, as an unevaluated LazyFrame.

Returns:

Type Description
LazyFrame

pl.LazyFrame: A lazy scan, or a lazy wrapper over the rows read.

Raises:

Type Description
ConnectorError

If connect() has not been called.

Source code in src/veridelta/connectors/base.py
def lazyframe(self) -> pl.LazyFrame:
    """Return what `connect()` opened, as an unevaluated LazyFrame.

    Returns:
        pl.LazyFrame: A lazy scan, or a lazy wrapper over the rows read.

    Raises:
        ConnectorError: If `connect()` has not been called.
    """
    if self._frame is None:
        raise ConnectorError(self._unconnected)
    return self._frame

SQLDialect

Bases: StrEnum

Warehouse SQL dialects supported by the pushdown compiler.

POSTGRES runs where two database sources on one Postgres connection opt into pushdown, and DUCKDB where two DuckDB tables in one database do. The differential test harness also runs DUCKDB output, so the compiled SQL itself, not a rewrite of it, is checked against the local engine.

Source code in src/veridelta/connectors/sql.py
class SQLDialect(StrEnum):
    """Warehouse SQL dialects supported by the pushdown compiler.

    `POSTGRES` runs where two database sources on one Postgres connection opt
    into pushdown, and `DUCKDB` where two DuckDB tables in one database do.
    The differential test harness also runs `DUCKDB` output, so the compiled
    SQL itself, not a rewrite of it, is checked against the local engine.
    """

    SNOWFLAKE = "snowflake"
    DATABRICKS = "databricks"
    DUCKDB = "duckdb"
    BIGQUERY = "bigquery"
    POSTGRES = "postgres"

SQLPushdownCompiler

Compile DiffRule semantics into dialect-specific SQL strings.

One instance targets one SQLDialect. The engine drives a warehouse run with these statements, in this order:

  1. compile_schema_probe_query per side, to learn column names and types.
  2. compile_duplicate_key_query per side, to assert normalized keys are unique before anything joins on them.
  3. compile_count_query per side, for the threshold denominator.
  4. compile_query for inner-join rows where a compared column differs.
  5. compile_added_query and compile_missing_query for the anti-joins.
  6. compile_column_mismatch_query for the per-column tally.

A value map proposal runs steps 1 and 2, then compile_value_map_query.

Every join reads keys through the same stages 1-7 as compared columns, driven by key_rules, so rows match on the keys the local engine sees.

Rules reach the compiler already folded over the configuration's default_* settings, so every compared column arrives as one fully specified DiffRule.

The compiler is also the security boundary for warehouse SQL. Identifiers are allowlisted segment by segment and then dialect-quoted, data literals are escaped through _literal, and every dialect keyword comes from a module-level table keyed by SQLDialect, so an unsupported combination raises rather than borrowing another dialect's spelling.

Attributes:

Name Type Description
dialect SQLDialect

Target dialect for quoting, casts, and functions.

Source code in src/veridelta/connectors/sql.py
 699
 700
 701
 702
 703
 704
 705
 706
 707
 708
 709
 710
 711
 712
 713
 714
 715
 716
 717
 718
 719
 720
 721
 722
 723
 724
 725
 726
 727
 728
 729
 730
 731
 732
 733
 734
 735
 736
 737
 738
 739
 740
 741
 742
 743
 744
 745
 746
 747
 748
 749
 750
 751
 752
 753
 754
 755
 756
 757
 758
 759
 760
 761
 762
 763
 764
 765
 766
 767
 768
 769
 770
 771
 772
 773
 774
 775
 776
 777
 778
 779
 780
 781
 782
 783
 784
 785
 786
 787
 788
 789
 790
 791
 792
 793
 794
 795
 796
 797
 798
 799
 800
 801
 802
 803
 804
 805
 806
 807
 808
 809
 810
 811
 812
 813
 814
 815
 816
 817
 818
 819
 820
 821
 822
 823
 824
 825
 826
 827
 828
 829
 830
 831
 832
 833
 834
 835
 836
 837
 838
 839
 840
 841
 842
 843
 844
 845
 846
 847
 848
 849
 850
 851
 852
 853
 854
 855
 856
 857
 858
 859
 860
 861
 862
 863
 864
 865
 866
 867
 868
 869
 870
 871
 872
 873
 874
 875
 876
 877
 878
 879
 880
 881
 882
 883
 884
 885
 886
 887
 888
 889
 890
 891
 892
 893
 894
 895
 896
 897
 898
 899
 900
 901
 902
 903
 904
 905
 906
 907
 908
 909
 910
 911
 912
 913
 914
 915
 916
 917
 918
 919
 920
 921
 922
 923
 924
 925
 926
 927
 928
 929
 930
 931
 932
 933
 934
 935
 936
 937
 938
 939
 940
 941
 942
 943
 944
 945
 946
 947
 948
 949
 950
 951
 952
 953
 954
 955
 956
 957
 958
 959
 960
 961
 962
 963
 964
 965
 966
 967
 968
 969
 970
 971
 972
 973
 974
 975
 976
 977
 978
 979
 980
 981
 982
 983
 984
 985
 986
 987
 988
 989
 990
 991
 992
 993
 994
 995
 996
 997
 998
 999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
1030
1031
1032
1033
1034
1035
1036
1037
1038
1039
1040
1041
1042
1043
1044
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
1075
1076
1077
1078
1079
1080
1081
1082
1083
1084
1085
1086
1087
1088
1089
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
1100
1101
1102
1103
1104
1105
1106
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
1118
1119
1120
1121
1122
1123
1124
1125
1126
1127
1128
1129
1130
1131
1132
1133
1134
1135
1136
1137
1138
1139
1140
1141
1142
1143
1144
1145
1146
1147
1148
1149
1150
1151
1152
1153
1154
1155
1156
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
1169
1170
1171
1172
1173
1174
1175
1176
1177
1178
1179
1180
1181
1182
1183
1184
1185
1186
1187
1188
1189
1190
1191
1192
1193
1194
1195
1196
1197
1198
1199
1200
1201
1202
1203
1204
1205
1206
1207
1208
1209
1210
1211
1212
1213
1214
1215
1216
1217
1218
1219
1220
1221
1222
1223
1224
1225
1226
1227
1228
1229
1230
1231
1232
1233
1234
1235
1236
1237
1238
1239
1240
1241
1242
1243
1244
1245
1246
1247
1248
1249
1250
1251
1252
1253
1254
1255
1256
1257
1258
1259
1260
1261
1262
1263
1264
1265
1266
1267
1268
1269
1270
1271
1272
1273
1274
1275
1276
1277
1278
1279
1280
1281
1282
1283
1284
1285
1286
1287
1288
1289
1290
1291
1292
1293
1294
1295
1296
1297
1298
1299
1300
1301
1302
1303
1304
1305
1306
1307
1308
1309
1310
1311
1312
1313
1314
1315
1316
1317
1318
1319
1320
1321
1322
1323
1324
1325
1326
1327
1328
1329
1330
1331
1332
1333
1334
1335
1336
1337
1338
1339
1340
1341
1342
1343
1344
1345
1346
1347
1348
1349
1350
1351
1352
1353
1354
1355
1356
1357
1358
1359
1360
1361
1362
1363
1364
1365
1366
1367
1368
1369
1370
1371
1372
1373
1374
1375
1376
1377
1378
1379
1380
1381
1382
1383
1384
1385
1386
1387
1388
1389
1390
1391
1392
1393
1394
1395
1396
1397
1398
1399
1400
1401
1402
1403
1404
1405
1406
1407
1408
1409
1410
1411
1412
1413
1414
1415
1416
1417
1418
1419
1420
1421
1422
1423
1424
1425
1426
1427
1428
1429
1430
1431
1432
1433
1434
1435
1436
1437
1438
1439
1440
1441
1442
1443
1444
1445
1446
1447
1448
1449
1450
1451
1452
1453
1454
1455
1456
1457
1458
1459
1460
1461
1462
1463
1464
1465
1466
1467
1468
1469
1470
1471
1472
1473
1474
1475
1476
1477
1478
1479
1480
1481
1482
1483
1484
1485
1486
1487
1488
1489
1490
1491
1492
1493
1494
1495
1496
1497
1498
1499
1500
1501
1502
1503
1504
1505
1506
1507
1508
1509
1510
1511
1512
1513
1514
1515
1516
1517
1518
1519
1520
1521
1522
1523
1524
1525
1526
1527
1528
1529
1530
1531
1532
1533
1534
1535
1536
1537
1538
1539
1540
1541
1542
1543
1544
1545
1546
1547
1548
1549
1550
1551
1552
1553
1554
1555
1556
1557
1558
1559
1560
1561
1562
1563
1564
1565
1566
1567
1568
1569
1570
1571
1572
1573
1574
1575
1576
1577
1578
1579
1580
1581
1582
1583
1584
1585
1586
1587
1588
1589
1590
1591
1592
1593
1594
1595
1596
1597
1598
1599
1600
1601
1602
1603
1604
1605
1606
1607
1608
1609
1610
1611
1612
1613
1614
1615
1616
1617
1618
1619
1620
1621
1622
1623
1624
1625
1626
1627
1628
1629
1630
1631
1632
1633
1634
1635
1636
1637
1638
1639
1640
1641
1642
1643
1644
1645
1646
1647
1648
1649
1650
1651
1652
1653
1654
1655
1656
1657
1658
1659
1660
1661
1662
1663
1664
1665
1666
1667
1668
1669
1670
1671
1672
1673
1674
1675
1676
1677
1678
1679
1680
1681
1682
1683
1684
1685
1686
1687
1688
1689
1690
1691
1692
1693
1694
1695
1696
1697
1698
1699
1700
1701
1702
1703
1704
1705
1706
1707
1708
1709
1710
1711
1712
1713
1714
1715
1716
1717
1718
1719
1720
1721
1722
1723
1724
1725
1726
1727
1728
1729
1730
1731
1732
1733
1734
1735
1736
1737
1738
1739
1740
1741
1742
1743
1744
1745
1746
1747
1748
1749
1750
1751
1752
1753
1754
1755
1756
1757
1758
1759
1760
1761
1762
1763
1764
1765
1766
1767
1768
1769
1770
1771
1772
1773
1774
1775
1776
1777
1778
1779
1780
1781
1782
1783
1784
1785
1786
1787
1788
1789
1790
1791
1792
1793
1794
1795
1796
1797
1798
1799
1800
1801
1802
1803
1804
1805
1806
1807
1808
1809
1810
1811
1812
1813
1814
1815
1816
1817
1818
1819
1820
1821
1822
1823
1824
1825
1826
1827
1828
1829
1830
1831
1832
1833
1834
1835
1836
1837
1838
1839
1840
1841
1842
1843
1844
1845
1846
1847
1848
1849
1850
1851
1852
1853
1854
1855
1856
1857
1858
1859
1860
1861
1862
1863
1864
1865
1866
1867
1868
1869
1870
1871
1872
1873
1874
1875
1876
1877
1878
1879
1880
1881
1882
1883
1884
1885
1886
1887
1888
class SQLPushdownCompiler:
    """Compile `DiffRule` semantics into dialect-specific SQL strings.

    One instance targets one `SQLDialect`. The engine drives a warehouse run
    with these statements, in this order:

    1. `compile_schema_probe_query` per side, to learn column names and types.
    2. `compile_duplicate_key_query` per side, to assert normalized keys are
       unique before anything joins on them.
    3. `compile_count_query` per side, for the `threshold` denominator.
    4. `compile_query` for inner-join rows where a compared column differs.
    5. `compile_added_query` and `compile_missing_query` for the anti-joins.
    6. `compile_column_mismatch_query` for the per-column tally.

    A value map proposal runs steps 1 and 2, then `compile_value_map_query`.

    Every join reads keys through the same stages 1-7 as compared columns,
    driven by `key_rules`, so rows match on the keys the local engine sees.

    Rules reach the compiler already folded over the configuration's
    `default_*` settings, so every compared column arrives as one fully
    specified `DiffRule`.

    The compiler is also the security boundary for warehouse SQL. Identifiers
    are allowlisted segment by segment and then dialect-quoted, data literals
    are escaped through `_literal`, and every dialect keyword comes from a
    module-level table keyed by `SQLDialect`, so an unsupported combination
    raises rather than borrowing another dialect's spelling.

    Attributes:
        dialect (SQLDialect): Target dialect for quoting, casts, and functions.
    """

    def __init__(self, dialect: SQLDialect) -> None:
        """Initialize a compiler for a single warehouse dialect.

        Args:
            dialect (SQLDialect): Target SQL dialect.
        """
        self.dialect = dialect

    def compile_column_predicate(
        self,
        rule: DiffRule,
        source_column: str,
        target_column: str | None = None,
        *,
        source_dtype: pl.DataType | None = None,
        target_dtype: pl.DataType | None = None,
    ) -> str:
        """Compile a boolean match predicate for one source/target column pair.

        Follows the canonical transform order documented on `DiffRule`, which is
        the single source of truth shared with the local engine. So a pushdown
        run and a local run evaluate the same pipeline. A setting a dialect
        cannot reproduce raises `ConfigError` instead of compiling to something
        that differs:

        - `min_jaro_winkler_similarity`, on every dialect;
        - `max_levenshtein_distance`, on Postgres and DuckDB;
        - `datetime_format` on Postgres, or with a directive the dialect lacks;
        - a `regex_replace` replacement that refers to a group by name.

        Args:
            rule (DiffRule): Semantic comparison overrides for the column.
            source_column (str): Column name on the source relation.
            target_column (str | None): Column name on the target relation. Defaults
                to `source_column` when omitted.
            source_dtype (pl.DataType | None): Probed source dtype, used to drop
                sentinels the column cannot hold. Each side is filtered
                separately because the two relations can disagree on a type.
                When None, sentinels are emitted unfiltered.
            target_dtype (pl.DataType | None): Probed target dtype.

        Returns:
            str: Boolean SQL expression that is true when the column values match.

        Raises:
            ConfigError: If the rule sets one of the settings listed above that
                this dialect refuses.
            ConnectorError: If identifiers are empty or not allowlisted.
        """
        tgt_name = target_column if target_column is not None else source_column
        src_expr = self._normalize_expr(
            self._qualify(_SOURCE_ALIAS, source_column),
            rule,
            source_dtype,
            is_source=True,
        )
        tgt_expr = self._normalize_expr(
            self._qualify(_TARGET_ALIAS, tgt_name),
            rule,
            target_dtype,
            is_source=False,
        )
        return self._compare(src_expr, tgt_expr, rule)

    def compile_query(
        self,
        source_table: str,
        target_table: str,
        primary_keys: list[str],
        rules: list[DiffRule],
        *,
        source_types: ColumnTypes | None = None,
        target_types: ColumnTypes | None = None,
        key_rules: Sequence[DiffRule] | None = None,
        wide_integers: frozenset[str] = frozenset(),
        type_drift: frozenset[str] = frozenset(),
    ) -> str:
        """Assemble a changed-row inner-join query from tables, keys, and rules.

        Stages 1 through 7 run once per column in a pair of CTEs, keys
        included. The join and the match predicates then read those projected
        values, so each stage appears once in the statement however many
        predicates read it.

        Args:
            source_table (str): Source relation (optionally dotted catalog path).
            target_table (str): Target relation (optionally dotted catalog path).
            primary_keys (list[str]): Join keys, spelled as the target stores them.
            rules (list[DiffRule]): Per-column semantic overrides.
            source_types (ColumnTypes | None): Probed source dtypes, used to drop
                null sentinels the column cannot hold. When None, sentinels are
                emitted unfiltered.
            target_types (ColumnTypes | None): Probed target dtypes.
            key_rules (Sequence[DiffRule] | None): One rule per key that needs
                normalizing, naming the stored source column and, for a renamed
                key, its `rename_to`. Keys without one are joined as stored.
            wide_integers (frozenset[str]): Compared columns, by target name,
                that hold integers on both sides after normalization. Their
                tolerance is measured in `_WIDE_INTEGER_TYPES`, so a difference
                wider than the stored type neither wraps nor overflows.
            type_drift (frozenset[str]): Compared columns, by target name, that
                `strict_types` fails because the two sides hold different
                types after normalization. No value of theirs ever matches, and
                two NULLs meet only under `treat_null_as_equal`.

        Returns:
            str: `SELECT ... FROM src INNER JOIN tgt ON ... WHERE NOT (...)` statement
            keeping the joined rows where at least one compared column differs.
            Each predicate is wrapped in `COALESCE(pred, FALSE)` so a one-sided
            NULL reads as a mismatch rather than as an unknown that `WHERE`
            drops, matching the local engine's `fill_null(False)`. When no
            column is compared the statement selects no rows, since without
            match expressions the local engine reports nothing as changed.

        Raises:
            ConfigError: If a rule sets `min_jaro_winkler_similarity`, or a
                `datetime_format` uses a directive this dialect cannot express.
            ConnectorError: If tables or keys are empty, a rule is pattern-only,
                `rename_to` is used with multiple `column_names`, or a key rule
                does not name exactly one primary key.
        """
        keys = self._key_columns(primary_keys, key_rules)
        compared = self._compared_columns(rules)
        with_clause = self._normalized_with_clause(
            source_table,
            target_table,
            [*keys, *compared],
            source_types=source_types,
            target_types=target_types,
        )
        select_list = ", ".join(self._qualify(_SOURCE_ALIAS, pk) for pk in primary_keys)
        join = self._normalized_join("INNER", primary_keys)
        statement = f"{with_clause} SELECT {select_list} {join}"
        predicates = self._column_predicates(
            compared,
            wide_integers=wide_integers,
            type_drift=type_drift,
        )
        if not predicates:
            # Nothing to compare means nothing can have changed. Without this
            # guard the bare join would report every shared key as drift.
            return f"{statement} WHERE 1 = 0"
        return f"{statement} WHERE {self._changed_condition(predicates)}"

    def compile_changed_sample_query(
        self,
        source_table: str,
        target_table: str,
        primary_keys: list[str],
        rules: list[DiffRule],
        *,
        limit: int,
        source_types: ColumnTypes | None = None,
        target_types: ColumnTypes | None = None,
        key_rules: Sequence[DiffRule] | None = None,
        wide_integers: frozenset[str] = frozenset(),
        type_drift: frozenset[str] = frozenset(),
    ) -> SampleQuery | None:
        """Assemble a query for the first changed rows, with both sides' values.

        This is `compile_query` with values: the same normalized CTEs, join,
        and WHERE clause, so every sampled row is one `compile_query` reports
        as changed, and each match flag is the predicate that decided it. Each
        output column takes a positional alias, so a long column name cannot
        pass an identifier limit and a key cannot clash with a suffixed column;
        `SampleQuery.renames` gives the local engine's names back. Rows come in
        key order, so the same tables give the same sample.

        Args:
            source_table (str): Source relation (optionally dotted catalog path).
            target_table (str): Target relation (optionally dotted catalog path).
            primary_keys (list[str]): Join keys, spelled as the target stores them.
            rules (list[DiffRule]): Per-column semantic overrides.
            limit (int): Most rows to return, at least 1.
            source_types (ColumnTypes | None): Probed source dtypes, as for
                `compile_query`.
            target_types (ColumnTypes | None): Probed target dtypes.
            key_rules (Sequence[DiffRule] | None): Key normalization, as for
                `compile_query`.
            wide_integers (frozenset[str]): Integer columns to measure in a wide
                type, as for `compile_query`.
            type_drift (frozenset[str]): Columns `strict_types` fails, as for
                `compile_query`.

        Returns:
            SampleQuery | None: The statement and its alias names, or None when
            no column is compared, since then no row can have changed.

        Raises:
            ConfigError: If a rule sets `min_jaro_winkler_similarity`, or a
                `datetime_format` uses a directive this dialect cannot express.
            ConnectorError: If `limit` is not a positive `int`, tables or keys
                are empty, a rule is pattern-only, `rename_to` is used with
                multiple `column_names`, or a key rule does not name exactly one
                primary key.
        """
        rendered_limit = self._integer(limit)
        if limit < 1:
            raise ConnectorError(f"A row sample needs a LIMIT of at least 1, got {limit}.")
        keys = self._key_columns(primary_keys, key_rules)
        compared = self._compared_columns(rules)
        if not compared:
            return None

        with_clause = self._normalized_with_clause(
            source_table,
            target_table,
            [*keys, *compared],
            source_types=source_types,
            target_types=target_types,
        )
        predicates = self._column_predicates(
            compared,
            wide_integers=wide_integers,
            type_drift=type_drift,
        )
        renames: dict[str, str] = {}
        projections: list[str] = []
        for index, key in enumerate(primary_keys):
            alias = f"{SAMPLE_KEY_PREFIX}{index}"
            projections.append(f"{self._qualify(_SOURCE_ALIAS, key)} AS {self._quote_ident(alias)}")
            renames[alias] = key
        for index, ((_source_column, target_column, _rule), predicate) in enumerate(
            zip(compared, predicates, strict=True)
        ):
            outputs = (
                (SAMPLE_SOURCE_PREFIX, self._qualify(_SOURCE_ALIAS, target_column), "source"),
                (SAMPLE_TARGET_PREFIX, self._qualify(_TARGET_ALIAS, target_column), "target"),
                (SAMPLE_MATCH_PREFIX, f"COALESCE({predicate}, FALSE)", "is_match"),
            )
            for prefix, expression, suffix in outputs:
                alias = f"{prefix}{index}"
                projections.append(f"{expression} AS {self._quote_ident(alias)}")
                renames[alias] = f"{target_column}_{suffix}"
        join = self._normalized_join("INNER", primary_keys)
        order = ", ".join(
            self._quote_ident(f"{SAMPLE_KEY_PREFIX}{index}") for index in range(len(primary_keys))
        )
        statement = (
            f"{with_clause} SELECT {', '.join(projections)} {join} "
            f"WHERE {self._changed_condition(predicates)} ORDER BY {order} LIMIT {rendered_limit}"
        )
        return SampleQuery(statement, renames)

    def compile_missing_query(
        self,
        source_table: str,
        target_table: str,
        primary_keys: list[str],
        *,
        source_types: ColumnTypes | None = None,
        target_types: ColumnTypes | None = None,
        key_rules: Sequence[DiffRule] | None = None,
    ) -> str:
        """Assemble a LEFT JOIN anti-join for rows present only in the source.

        Args:
            source_table (str): Source relation (optionally dotted catalog path).
            target_table (str): Target relation (optionally dotted catalog path).
            primary_keys (list[str]): Join keys, spelled as the target stores them.
            source_types (ColumnTypes | None): Probed source dtypes.
            target_types (ColumnTypes | None): Probed target dtypes.
            key_rules (Sequence[DiffRule] | None): Key normalization, as for
                `compile_query`.

        Returns:
            str: Source keys whose normalized value has no target counterpart.

        Raises:
            ConnectorError: If tables or keys are empty, or a key rule does not
                name exactly one primary key.
        """
        return self._compile_anti_join(
            source_table,
            target_table,
            primary_keys,
            join_kind="LEFT",
            source_types=source_types,
            target_types=target_types,
            key_rules=key_rules,
        )

    def compile_added_query(
        self,
        source_table: str,
        target_table: str,
        primary_keys: list[str],
        *,
        source_types: ColumnTypes | None = None,
        target_types: ColumnTypes | None = None,
        key_rules: Sequence[DiffRule] | None = None,
    ) -> str:
        """Assemble a RIGHT JOIN anti-join for rows present only in the target.

        Args:
            source_table (str): Source relation (optionally dotted catalog path).
            target_table (str): Target relation (optionally dotted catalog path).
            primary_keys (list[str]): Join keys, spelled as the target stores them.
            source_types (ColumnTypes | None): Probed source dtypes.
            target_types (ColumnTypes | None): Probed target dtypes.
            key_rules (Sequence[DiffRule] | None): Key normalization, as for
                `compile_query`.

        Returns:
            str: Target keys whose normalized value has no source counterpart.

        Raises:
            ConnectorError: If tables or keys are empty, or a key rule does not
                name exactly one primary key.
        """
        return self._compile_anti_join(
            source_table,
            target_table,
            primary_keys,
            join_kind="RIGHT",
            source_types=source_types,
            target_types=target_types,
            key_rules=key_rules,
        )

    def compile_column_mismatch_query(
        self,
        source_table: str,
        target_table: str,
        primary_keys: list[str],
        rules: list[DiffRule],
        *,
        source_types: ColumnTypes | None = None,
        target_types: ColumnTypes | None = None,
        key_rules: Sequence[DiffRule] | None = None,
        wide_integers: frozenset[str] = frozenset(),
        type_drift: frozenset[str] = frozenset(),
    ) -> str | None:
        """Assemble a per-column mismatch tally over the joined rows.

        Each column contributes one `SUM(CASE ...)` term, so a single round trip
        fills `DiffSummary.column_mismatches` the way the local engine does.
        `COALESCE(pred, FALSE)` is load-bearing: under three-valued logic a NULL
        predicate is neither true nor false, and without the coalesce those rows
        would silently count as matches instead of mismatches. The local engine
        resolves the same case with `val_match.fill_null(False)`.

        Args:
            source_table (str): Source relation (optionally dotted catalog path).
            target_table (str): Target relation (optionally dotted catalog path).
            primary_keys (list[str]): Join keys present on both relations.
            rules (list[DiffRule]): Per-column semantic overrides.
            source_types (ColumnTypes | None): Probed source dtypes, used to drop
                null sentinels the column cannot hold. When None, sentinels are
                emitted unfiltered.
            target_types (ColumnTypes | None): Probed target dtypes.
            key_rules (Sequence[DiffRule] | None): Key normalization, as for
                `compile_query`.
            wide_integers (frozenset[str]): Integer columns to measure in a wide
                type, as for `compile_query`.
            type_drift (frozenset[str]): Columns `strict_types` fails, as for
                `compile_query`.

        Returns:
            str | None: Single-row aggregate statement, or None when no rule
            yields a comparable column, mirroring the local engine's decision to
            skip the tally when there are no match expressions.

        Raises:
            ConfigError: If a rule sets `min_jaro_winkler_similarity`, or a
                `datetime_format` uses a directive this dialect cannot express.
            ConnectorError: If tables or keys are empty, a rule is pattern-only,
                `rename_to` is used with multiple `column_names`, or a key rule
                does not name exactly one primary key.
        """
        keys = self._key_columns(primary_keys, key_rules)
        compared = self._compared_columns(rules)
        if not compared:
            return None

        with_clause = self._normalized_with_clause(
            source_table,
            target_table,
            [*keys, *compared],
            source_types=source_types,
            target_types=target_types,
        )
        predicates = self._column_predicates(
            compared,
            wide_integers=wide_integers,
            type_drift=type_drift,
        )
        terms = [
            f"SUM(CASE WHEN COALESCE({predicate}, FALSE) THEN 0 ELSE 1 END) "
            f"AS {self._quote_ident(target_column)}"
            for (_source_column, target_column, _rule), predicate in zip(
                compared, predicates, strict=True
            )
        ]
        join = self._normalized_join("INNER", primary_keys)
        return f"{with_clause} SELECT {', '.join(terms)} {join}"

    def compile_value_map_query(
        self,
        source_table: str,
        target_table: str,
        primary_keys: list[str],
        rules: list[DiffRule],
        *,
        min_support: int,
        sample_fraction: float = 1.0,
        source_types: ColumnTypes | None = None,
        target_types: ColumnTypes | None = None,
        key_rules: Sequence[DiffRule] | None = None,
    ) -> str | None:
        """Assemble one statement that counts how source and target values line up.

        Keys and candidate columns are normalized in the usual CTE pair and
        joined once. Each column then contributes one `UNION ALL` branch,
        labeled by its position in `rules` rather than by name, so no column
        name becomes a string literal. A branch leaves out NULL sources and
        values the column's existing map produced, as the local engine does.

        Only exact predicates run here: the target differs from the source, at
        least `min_support` rows agree, and the agreeing rows are more than
        half of the source value's rows. Every confidence floor is above one
        half, so this keeps a superset of what qualifies, and the engine
        applies the floor itself, in floating point exactly as locally.

        Args:
            source_table (str): Source relation (optionally dotted catalog path).
            target_table (str): Target relation (optionally dotted catalog path).
            primary_keys (list[str]): Join keys, spelled as the target stores them.
            rules (list[DiffRule]): One rule per candidate column, naming its
                stored source column and, when renamed, the target's name.
            min_support (int): Agreeing rows a pair needs.
            sample_fraction (float): Share of source keys to read, chosen by a
                hash of the normalized keys. 1 reads every row.
            source_types (ColumnTypes | None): Probed source dtypes.
            target_types (ColumnTypes | None): Probed target dtypes.
            key_rules (Sequence[DiffRule] | None): Key normalization, as for
                `compile_query`.

        Returns:
            str | None: A statement returning `VALUE_MAP_COLUMN_ALIAS`,
            `VALUE_MAP_SOURCE_ALIAS`, `VALUE_MAP_TARGET_ALIAS`,
            `VALUE_MAP_ROWS_ALIAS`, and `VALUE_MAP_AGREEING_ALIAS` per pair, or
            None when there is no candidate column.

        Raises:
            ConnectorError: If tables or keys are empty, a rule is pattern-only,
                a key rule does not name exactly one primary key, or
                `min_support` is not an integer.
        """
        keys = self._key_columns(primary_keys, key_rules)
        compared = self._compared_columns(rules)
        if not compared:
            return None
        with_clause = self._normalized_with_clause(
            source_table,
            target_table,
            [*keys, *compared],
            source_types=source_types,
            target_types=target_types,
        )
        joined = self._value_map_joined(
            [target for _source, target, _rule in compared],
            primary_keys,
            sample_fraction,
        )
        branches = " UNION ALL ".join(
            self._value_map_branch(label, rule)
            for label, (_source, _target, rule) in enumerate(compared)
        )
        return f"{with_clause}, {joined} {self._value_map_tally(branches, min_support)}"

    def _value_map_joined(
        self,
        columns: list[str],
        primary_keys: list[str],
        sample_fraction: float,
    ) -> str:
        """Build the CTE pairing each candidate's normalized values on the keys."""
        projections = ", ".join(
            f"{self._qualify(_SOURCE_ALIAS, column)} AS {self._quote_ident(f'{SAMPLE_SOURCE_PREFIX}{label}')}, "
            f"{self._qualify(_TARGET_ALIAS, column)} AS {self._quote_ident(f'{SAMPLE_TARGET_PREFIX}{label}')}"
            for label, column in enumerate(columns)
        )
        join = self._normalized_join("INNER", primary_keys)
        sample = ""
        if sample_fraction < 1:
            keys = [self._qualify(_SOURCE_ALIAS, key) for key in primary_keys]
            cutoff = self._integer(round(sample_fraction * SAMPLE_BUCKETS))
            sample = f" WHERE {self._sample_bucket(keys)} < {cutoff}"
        return f"{self._quote_ident(_JOINED_CTE)} AS (SELECT {projections} {join}{sample})"

    def _sample_bucket(self, keys: list[str]) -> str:
        """Hash qualified keys into one of `SAMPLE_BUCKETS` buckets."""
        joined = ", ".join(keys)
        function = _SAMPLE_HASH_FUNCTIONS[self.dialect]
        buckets = self._integer(SAMPLE_BUCKETS)
        if self.dialect is SQLDialect.DATABRICKS:
            # pmod is always non-negative.
            return f"pmod({function}({joined}), {buckets})"
        if self.dialect is SQLDialect.DUCKDB:
            # DuckDB's hash is unsigned.
            return f"{function}({joined}) % {buckets}"
        # These hashes are signed, and MOD keeps the dividend's sign. FARM_FINGERPRINT
        # takes one string, and hashtextextended one text and a seed.
        hashed = {
            SQLDialect.SNOWFLAKE: f"{function}({joined})",
            SQLDialect.BIGQUERY: f"{function}(TO_JSON_STRING(STRUCT({joined})))",
            SQLDialect.POSTGRES: f"{function}(ROW({joined})::text, 0)",
        }[self.dialect]
        return f"MOD(MOD({hashed}, {buckets}) + {buckets}, {buckets})"

    def _value_map_branch(self, label: int, rule: DiffRule) -> str:
        """Select one candidate's pairs, labeled, without NULL or already mapped sources."""
        source = self._quote_ident(f"{SAMPLE_SOURCE_PREFIX}{label}")
        target = self._quote_ident(f"{SAMPLE_TARGET_PREFIX}{label}")
        conditions = [f"{source} IS NOT NULL"]
        if rule.value_map:
            outputs = ", ".join(
                self._literal(value) for value in dict.fromkeys(rule.value_map.values())
            )
            conditions.append(f"{source} NOT IN ({outputs})")
        return (
            f"SELECT {self._integer(label)} AS {self._quote_ident(VALUE_MAP_COLUMN_ALIAS)}, "
            f"{source} AS {self._quote_ident(VALUE_MAP_SOURCE_ALIAS)}, "
            f"{target} AS {self._quote_ident(VALUE_MAP_TARGET_ALIAS)} "
            f"FROM {self._quote_ident(_JOINED_CTE)} WHERE {' AND '.join(conditions)}"
        )

    def _value_map_tally(self, branches: str, min_support: int) -> str:
        """Count the labeled pairs and keep the ones that can qualify."""
        column = self._quote_ident(VALUE_MAP_COLUMN_ALIAS)
        source = self._quote_ident(VALUE_MAP_SOURCE_ALIAS)
        target = self._quote_ident(VALUE_MAP_TARGET_ALIAS)
        rows = self._quote_ident(VALUE_MAP_ROWS_ALIAS)
        agreeing = self._quote_ident(VALUE_MAP_AGREEING_ALIAS)
        count_type = self._cast_keyword("Int64")
        grouped = (
            f"SELECT {column}, {source}, {target}, COUNT(*) AS {agreeing} "
            f"FROM ({branches}) AS {self._quote_ident('_veridelta_pairs')} "
            f"GROUP BY {column}, {source}, {target}"
        )
        # The per-value total wraps the pair count, a form every supported dialect accepts.
        counted = (
            f"SELECT {column}, {source}, {target}, {agreeing}, "
            f"SUM({agreeing}) OVER (PARTITION BY {column}, {source}) AS {rows} "
            f"FROM ({grouped}) AS {self._quote_ident('_veridelta_groups')}"
        )
        # A NULL target forms its own group: it counts toward the total, then `<>` drops it.
        return (
            f"SELECT {column}, {source}, {target}, "
            f"CAST({rows} AS {count_type}) AS {rows}, "
            f"CAST({agreeing} AS {count_type}) AS {agreeing} "
            f"FROM ({counted}) AS {self._quote_ident('_veridelta_counts')} "
            f"WHERE {target} <> {source} AND {agreeing} >= {self._integer(min_support)} "
            f"AND 2 * {agreeing} > {rows}"
        )

    def compile_count_query(self, table: str) -> str:
        """Assemble a total row count query for one relation.

        The count supplies the denominator for `DiffSummary.mismatch_ratio`, so
        warehouse runs honor `threshold` the same way local comparisons do.

        Args:
            table (str): Relation to count (optionally dotted catalog path).

        Returns:
            str: `SELECT COUNT(*) AS alias FROM relation` with the alias quoted
            for the active dialect.

        Raises:
            ConnectorError: If the relation name is empty or not allowlisted.
        """
        return (
            f"SELECT COUNT(*) AS {self._quote_ident(COUNT_ALIAS)} "
            f"FROM {self._quote_relation(table)}"
        )

    def compile_duplicate_key_query(
        self,
        table: str,
        primary_keys: list[str],
        *,
        is_source: bool,
        key_rules: Sequence[DiffRule] | None = None,
        types: ColumnTypes | None = None,
    ) -> str:
        """Assemble a count of the rows whose normalized key is not unique.

        The local engine asserts uniqueness after normalization and reports
        every row that shares its key with another. Summing the size of each
        key group larger than one is that same number, and `GROUP BY` puts NULL
        keys in one group, as Polars counts them as duplicates of each other.

        Args:
            table (str): Relation to check (optionally dotted catalog path).
            primary_keys (list[str]): Keys, spelled as the target stores them.
            is_source (bool): Whether the relation is the source, which reads a
                renamed key under its stored name and applies `value_map`.
            key_rules (Sequence[DiffRule] | None): Key normalization, as for
                `compile_query`.
            types (ColumnTypes | None): Probed dtypes for this relation.

        Returns:
            str: A single-row statement whose `COUNT_ALIAS` column is zero when
            every normalized key is unique.

        Raises:
            ConnectorError: If the table or keys are empty, or a key rule does
                not name exactly one primary key.
        """
        keys = self._key_columns(primary_keys, key_rules)
        alias = "src" if is_source else "tgt"
        normalized = self._normalized_select(table, alias, keys, types=types, is_source=is_source)
        key_list = ", ".join(self._quote_ident(pk) for pk in primary_keys)
        rows = self._quote_ident(DUPLICATE_ROWS_ALIAS)
        return (
            f"SELECT COALESCE(SUM({rows}), 0) AS {self._quote_ident(COUNT_ALIAS)} "
            f"FROM (SELECT COUNT(*) AS {rows} "
            f"FROM ({normalized}) AS {self._quote_ident(KEYS_ALIAS)} "
            f"GROUP BY {key_list} HAVING COUNT(*) > 1) AS {self._quote_ident(DUPLICATES_ALIAS)}"
        )

    def compile_schema_probe_query(self, table: str) -> str:
        """Assemble a zero-row projection used to read a relation's columns.

        Args:
            table (str): Relation to probe (optionally dotted catalog path).

        Returns:
            str: `SELECT * FROM relation WHERE 1 = 0`, which returns column
            metadata without scanning rows.

        Raises:
            ConnectorError: If the relation name is empty or not allowlisted.
        """
        return f"SELECT * FROM {self._quote_relation(table)} WHERE 1 = 0"

    def _compile_anti_join(
        self,
        source_table: str,
        target_table: str,
        primary_keys: list[str],
        *,
        join_kind: str,
        source_types: ColumnTypes | None,
        target_types: ColumnTypes | None,
        key_rules: Sequence[DiffRule] | None,
    ) -> str:
        """Assemble a LEFT or RIGHT JOIN anti-join selecting keys from one side."""
        keys = self._key_columns(primary_keys, key_rules)
        with_clause = self._normalized_with_clause(
            source_table,
            target_table,
            keys,
            source_types=source_types,
            target_types=target_types,
        )
        select_alias, null_alias = (
            (_SOURCE_ALIAS, _TARGET_ALIAS)
            if join_kind == "LEFT"
            else (_TARGET_ALIAS, _SOURCE_ALIAS)
        )
        select_list = ", ".join(self._qualify(select_alias, pk) for pk in primary_keys)
        join = self._normalized_join(join_kind, primary_keys)
        where_clause = " AND ".join(
            f"{self._qualify(null_alias, pk)} IS NULL" for pk in primary_keys
        )
        return f"{with_clause} SELECT {select_list} {join} WHERE {where_clause}"

    def _normalized_join(self, join_kind: str, primary_keys: list[str]) -> str:
        """Join the two normalized CTEs on their keys."""
        on_clause = " AND ".join(
            f"{self._qualify(_SOURCE_ALIAS, pk)} = {self._qualify(_TARGET_ALIAS, pk)}"
            for pk in primary_keys
        )
        return (
            f"FROM {self._quote_ident(_SOURCE_CTE)} AS {self._quote_ident(_SOURCE_ALIAS)} "
            f"{join_kind} JOIN {self._quote_ident(_TARGET_CTE)} "
            f"AS {self._quote_ident(_TARGET_ALIAS)} ON {on_clause}"
        )

    def _normalize_expr(
        self,
        expr: str,
        rule: DiffRule,
        dtype: pl.DataType | None,
        *,
        is_source: bool,
    ) -> str:
        """Apply stages 1 through 7 to one side of a comparison."""
        # With no probed dtype, every sentinel is emitted as configured.
        sentinels = (
            list(rule.null_values or [])
            if dtype is None
            else usable_sentinels(rule.null_values, dtype)
        )
        expr = self._apply_null_values(expr, sentinels)
        if self._is_text_side(dtype):
            expr = self._apply_regex_replace(expr, rule)
            expr = self._apply_whitespace(expr, rule)
            expr = self._apply_case(expr, rule)
            if is_source:
                expr = self._apply_value_map(expr, rule)
        expr = self._apply_pad_zeros(expr, rule)
        expr = self._apply_datetime_format(expr, rule, dtype)
        # Stage 6b, `timezone`, emits no SQL; see `reject_unzoned_timezone` in `_resolution.py`.
        return self._apply_cast(expr, rule, dtype)

    def _key_columns(
        self, primary_keys: list[str], key_rules: Sequence[DiffRule] | None
    ) -> list[_Projection]:
        """Pair each primary key with its stored source name and normalizing rule."""
        if not primary_keys:
            raise ConnectorError("At least one primary key is required for pushdown joins.")
        by_key: dict[str, DiffRule] = {}
        for rule in key_rules or ():
            if len(rule.column_names) != 1:
                raise ConnectorError("A key rule must name exactly one stored column.")
            key = rule.rename_to or rule.column_names[0]
            if key not in primary_keys:
                raise ConnectorError(f"Key rule target '{key}' is not a primary key.")
            by_key[key] = rule
        return [
            (by_key[key].column_names[0] if key in by_key else key, key, by_key.get(key))
            for key in primary_keys
        ]

    def _column_predicates(
        self,
        compared: list[tuple[str, str, DiffRule]],
        *,
        wide_integers: frozenset[str],
        type_drift: frozenset[str],
    ) -> list[str]:
        """Build each compared column's match predicate over the normalized CTEs."""
        return [
            self._compare(
                self._qualify(_SOURCE_ALIAS, target_column),
                self._qualify(_TARGET_ALIAS, target_column),
                rule,
                wide=target_column in wide_integers,
                drift=target_column in type_drift,
            )
            for _source_column, target_column, rule in compared
        ]

    @staticmethod
    def _changed_condition(predicates: list[str]) -> str:
        """Join match predicates into the condition a changed row meets."""
        joined = " AND ".join(f"COALESCE({predicate}, FALSE)" for predicate in predicates)
        return f"NOT ({joined})"

    def _compared_columns(self, rules: list[DiffRule]) -> list[tuple[str, str, DiffRule]]:
        """Resolve the columns that a join query compares."""
        collected: dict[str, tuple[str, str, DiffRule]] = {}
        for rule in rules:
            if rule.pattern is not None and not rule.column_names:
                raise ConnectorError(
                    "Pattern-only DiffRule cannot be compiled without a resolved column name."
                )
            if rule.ignore or not rule.column_names:
                continue
            if rule.rename_to is not None and len(rule.column_names) != 1:
                raise ConnectorError(
                    "rename_to is only valid when column_names has exactly one entry."
                )
            for column in rule.column_names:
                target_column = rule.rename_to if rule.rename_to is not None else column
                # The first rule to name a target column wins, as in the local engine.
                collected.setdefault(target_column, (column, target_column, rule))
        return list(collected.values())

    def _normalized_with_clause(
        self,
        source_table: str,
        target_table: str,
        columns: Sequence[_Projection],
        *,
        source_types: ColumnTypes | None,
        target_types: ColumnTypes | None,
    ) -> str:
        """Build the CTE pair that applies stages 1-7 once per column."""
        # Projecting each normalized value once keeps later SQL linear in the number of columns.
        src_select = self._normalized_select(
            source_table, _SOURCE_ALIAS, columns, types=source_types, is_source=True
        )
        tgt_select = self._normalized_select(
            target_table, _TARGET_ALIAS, columns, types=target_types, is_source=False
        )
        return (
            f"WITH {self._quote_ident(_SOURCE_CTE)} AS ({src_select}), "
            f"{self._quote_ident(_TARGET_CTE)} AS ({tgt_select})"
        )

    def _normalized_select(
        self,
        table: str,
        alias: str,
        columns: Sequence[_Projection],
        *,
        types: ColumnTypes | None,
        is_source: bool,
    ) -> str:
        """Project one normalized expression per key and compared column."""
        projections: list[str] = []
        for source_column, target_column, rule in columns:
            raw_name = source_column if is_source else target_column
            expr = self._qualify(alias, raw_name)
            if rule is not None:
                dtype = None if types is None else types.get(raw_name)
                expr = self._normalize_expr(expr, rule, dtype, is_source=is_source)
            projections.append(f"{expr} AS {self._quote_ident(target_column)}")
        return (
            f"SELECT {', '.join(projections)} "
            f"FROM {self._quote_relation(table)} AS {self._quote_ident(alias)}"
        )

    def _quote_ident(self, name: str) -> str:
        """Quote a single SQL identifier for the active dialect."""
        # The allowlist admits no quote character, so the name needs no escaping.
        quote = _IDENTIFIER_QUOTES[self.dialect]
        return f"{quote}{_allowlisted(name)}{quote}"

    def _quote_relation(self, name: str) -> str:
        """Quote a possibly dotted table, schema, or catalog path."""
        return ".".join(self._quote_ident(part) for part in _relation_segments(name))

    def _qualify(self, alias: str, column: str) -> str:
        """Return `alias.column` with both parts quoted."""
        return f"{self._quote_ident(alias)}.{self._quote_ident(column)}"

    def _literal(self, value: str) -> str:
        """Render a single-quoted SQL string literal for the active dialect."""
        return _string_literal(self.dialect, value)

    def _integer(self, value: object) -> str:
        """Render an integer SQL operand, refusing anything that is not an `int`."""
        # Other numbers render through `repr`, so `True` or a NumPy scalar would reach SQL.
        if isinstance(value, bool) or not isinstance(value, int):
            raise ConnectorError(f"SQL integer operands must be int, got {value!r}.")
        return str(value)

    def _apply_null_values(self, expr: str, sentinels: Sequence[SentinelValue]) -> str:
        """Coerce sentinel values to NULL with a single `CASE` expression."""
        if not sentinels:
            return expr
        rendered = ", ".join(self._sentinel_literal(value) for value in sentinels)
        # Sentinels are stage 1, so `expr` is a bare column and repeating it costs nothing.
        return f"CASE WHEN {expr} IN ({rendered}) THEN NULL ELSE {expr} END"

    def _sentinel_literal(self, value: SentinelValue) -> str:
        """Render one sentinel as a SQL literal of its own type."""
        # bool first, since isinstance(True, int) is True in Python.
        if isinstance(value, bool):
            return "TRUE" if value else "FALSE"
        # A quoted number such as `'-999'` brings back the cast error that type filtering prevents.
        if isinstance(value, (int, float)):
            return repr(value)
        return self._literal(value)

    def _apply_regex_replace(self, expr: str, rule: DiffRule) -> str:
        """Apply `REGEXP_REPLACE` for each pattern/replacement pair, to every match."""
        if not rule.regex_replace:
            return expr
        wrapped = expr
        flags = _REGEX_REPLACE_FLAGS[self.dialect]
        for pattern, replacement in rule.regex_replace.items():
            written = self._regex_replacement(pattern, replacement)
            wrapped = (
                f"REGEXP_REPLACE({wrapped}, {self._literal(pattern)}, "
                f"{self._literal(written)}{flags})"
            )
        return wrapped

    def _regex_replacement(self, pattern: str, replacement: str) -> str:
        """Rewrite a Polars replacement in the dialect's replacement syntax."""
        written: list[str] = []
        follows_group = False
        for token in _replacement_tokens(pattern, replacement):
            if isinstance(token, int):
                written.append(self._group_reference(token))
                follows_group = True
                continue
            # Polars reads a backslash as plain text; each dialect reads it as an escape.
            if self.dialect is SQLDialect.DATABRICKS:
                # Databricks follows Java: `$` starts a group, and a digit right after a group
                # extends its number.
                token = token.replace("\\", "\\\\").replace("$", "\\$")
                if follows_group and token[:1].isdigit():
                    token = f"\\{token}"
            else:
                token = token.replace("\\", "\\\\")
            written.append(token)
            follows_group = False
        return "".join(written)

    def _group_reference(self, group: int) -> str:
        """Write a reference to a numbered group in the dialect's replacement syntax."""
        if self.dialect is SQLDialect.DATABRICKS:
            return f"${group}"
        # Postgres reads `\0` as plain text and spells the whole match `\&`.
        if group == 0 and self.dialect is SQLDialect.POSTGRES:
            return "\\&"
        return f"\\{group}"

    def _apply_whitespace(self, expr: str, rule: DiffRule) -> str:
        """Trim the characters Polars strips, from the side `whitespace_mode` names."""
        mode = rule.whitespace_mode
        if mode not in _TRIM_FUNCTIONS:
            return expr
        # A bare `TRIM` strips only spaces on Snowflake, Databricks, and DuckDB.
        characters = self._literal(_WHITESPACE_CHARACTERS)
        # Databricks' two-argument `ltrim` and `rtrim` take characters first and are deprecated.
        if self.dialect is SQLDialect.DATABRICKS:
            return f"TRIM({_TRIM_SIDES[mode]} {characters} FROM {expr})"
        return f"{_TRIM_FUNCTIONS[mode]}({expr}, {characters})"

    def _apply_case(self, expr: str, rule: DiffRule) -> str:
        """Lowercase an expression when `case_insensitive` is enabled."""
        if rule.case_insensitive:
            return f"LOWER({expr})"
        return expr

    def _apply_value_map(self, expr: str, rule: DiffRule) -> str:
        """Map source values with nested `IFF` on Snowflake, `CASE` elsewhere."""
        if not rule.value_map:
            return expr
        if self.dialect is SQLDialect.SNOWFLAKE:
            wrapped = expr
            for key, value in reversed(list(rule.value_map.items())):
                wrapped = f"IFF({expr} = {self._literal(key)}, {self._literal(value)}, {wrapped})"
            return wrapped

        branches = " ".join(
            f"WHEN {expr} = {self._literal(key)} THEN {self._literal(value)}"
            for key, value in rule.value_map.items()
        )
        return f"CASE {branches} ELSE {expr} END"

    def _is_text_side(self, dtype: pl.DataType | None) -> bool:
        """Return whether stages 2 through 4 apply to one side of a comparison."""
        # Polars' `.str` rejects `Categorical` and `Enum`, so pushdown skips them too, for parity.
        return dtype is None or isinstance(dtype, pl.String)

    def _apply_pad_zeros(self, expr: str, rule: DiffRule) -> str:
        """Left-pad with zeros the way Python's `str.zfill` does."""
        if rule.pad_zeros is None:
            return expr
        text = f"CAST({expr} AS {self._cast_keyword('String')})"
        width = rule.pad_zeros
        # Width zero still casts, as the local engine does, because later stages branch on text.
        if width == 0:
            return text
        # `LPAD` alone pads before a sign, so `-12` becomes `0-12`, not `-012`, and it truncates
        # longer input, which Polars leaves untouched.
        sign = f"SUBSTR({text}, 1, 1)"
        return (
            f"CASE WHEN LENGTH({text}) >= {width} THEN {text} "
            f"WHEN {sign} IN ('-', '+') "
            f"THEN {sign} || LPAD(SUBSTR({text}, 2), {width - 1}, '0') "
            f"ELSE LPAD({text}, {width}, '0') END"
        )

    def _apply_datetime_format(self, expr: str, rule: DiffRule, dtype: pl.DataType | None) -> str:
        """Parse text timestamps with the dialect's non-throwing parser."""
        if not rule.datetime_format:
            return expr
        if rule.pad_zeros is None and not self._is_text_side(dtype):
            return expr
        parse_function = _PARSE_FUNCTIONS.get(self.dialect)
        if parse_function is None:
            raise ConfigError(
                f"datetime_format cannot be pushed down to {self.dialect.value}: it has no "
                "parse that returns NULL for a value it cannot read, as a local run does, so "
                "one bad row would fail the whole statement. Compare locally instead "
                "(pushdown: false)."
            )
        pattern = self._translate_datetime_format(rule.datetime_format)
        if self.dialect is SQLDialect.BIGQUERY:
            # BigQuery takes the format first, and only a TIMESTAMP holds an offset.
            parse = (
                _BIGQUERY_OFFSET_PARSE if _reads_offset(rule.datetime_format) else parse_function
            )
            return f"{parse}({self._literal(pattern)}, {expr})"
        return f"{parse_function}({expr}, {self._literal(pattern)})"

    def _translate_datetime_format(self, fmt: str) -> str:
        """Rewrite a Python `strptime` format in the dialect's format language."""
        directives = _STRPTIME_DIRECTIVES[self.dialect]
        quote = _FORMAT_LITERAL_QUOTES[self.dialect]
        out: list[str] = []
        literal: list[str] = []

        # Quoting each literal run keeps a separator from reading as a format element.
        def flush() -> None:
            if literal:
                out.append(f"{quote}{''.join(literal)}{quote}")
                literal.clear()

        # A substitution pass would mistake literal letters for directives and keep unknown ones.
        index = 0
        while index < len(fmt):
            char = fmt[index]
            if char != "%":
                if char not in _FORMAT_LITERALS:
                    allowed = "".join(sorted(_FORMAT_LITERALS))
                    raise ConfigError(
                        f"datetime_format '{fmt}' contains the literal character "
                        f"'{char}', which SQL pushdown cannot quote. Allowed "
                        f"separators: {allowed!r}."
                    )
                literal.append(char)
                index += 1
                continue

            if index + 1 >= len(fmt):
                raise ConfigError(f"datetime_format '{fmt}' ends with a dangling '%'.")
            code = fmt[index + 1]
            mapped = directives.get(code)
            # An unknown directive parses to NULL, which reads as a clean match, not an error.
            if mapped is None:
                supported = ", ".join(f"%{key}" for key in directives)
                raise ConfigError(
                    f"datetime_format '{fmt}' uses '%{code}', which SQL pushdown "
                    f"cannot translate for {self.dialect.value}. Supported "
                    f"directives: {supported}."
                )
            flush()
            out.append(mapped)
            index += 2

        flush()
        return "".join(out)

    def _apply_cast(self, expr: str, rule: DiffRule, dtype: pl.DataType | None) -> str:
        """Cast to the configured target type using the dialect's keyword."""
        if rule.cast_to is None:
            return expr
        # Every earlier stage that fires on a number leaves text or a timestamp
        # behind, so the probed dtype reaches the cast only when none of them does.
        earlier = rule.pad_zeros is not None or rule.datetime_format or rule.timezone
        precast = None if earlier else dtype
        if (
            rule.cast_to == "Boolean"
            and self.dialect in _ZERO_TEST_BOOLEANS
            and precast is not None
            and precast.is_numeric()
        ):
            # Polars reads every nonzero number as true, NaN included, as `<> 0` does.
            return f"({expr} <> 0)"
        if rule.cast_to == "Int64" and precast is not None and precast.is_float():
            # Polars truncates a float toward zero on the way to an integer.
            # Snowflake and DuckDB round instead, so 10.7 would compare as 11
            # under pushdown and 10 locally. Truncate explicitly rather than
            # inherit whichever behavior the warehouse happens to have. Only
            # floats need this: `Decimal` rounds to integers in Polars exactly
            # as SQL does.
            expr = f"CASE WHEN {expr} < 0 THEN CEIL({expr}) ELSE FLOOR({expr}) END"
        return f"CAST({expr} AS {self._cast_keyword(rule.cast_to)})"

    def _cast_keyword(self, target: CastTarget) -> str:
        """Look up the dialect keyword for a cast target."""
        keyword = _CAST_KEYWORDS[self.dialect].get(target)
        if keyword is None:
            raise ConnectorError(f"SQL pushdown has no {self.dialect.value} type for '{target}'.")
        return keyword

    def _numeric_predicate(
        self, src_expr: str, tgt_expr: str, rule: DiffRule, *, wide: bool = False
    ) -> str:
        """Build the engine-equivalent absolute/relative tolerance predicate."""
        abs_tol = repr(rule.absolute_tolerance or 0.0)
        rel_tol = repr(rule.relative_tolerance or 0.0)
        if wide:
            wide_type = _WIDE_INTEGER_TYPES[self.dialect]
            src = f"CAST({src_expr} AS {wide_type})"
            tgt = f"CAST({tgt_expr} AS {wide_type})"
            return (
                f"({self._value_equality(src, tgt)} OR "
                f"ABS({tgt} - {src}) <= {abs_tol} + ({rel_tol} * ABS({src})))"
            )
        # The allowance needs a finite source, since `0 * ABS(inf)` is NaN. Every supported engine
        # sorts NaN above all numbers, so `ABS(diff) <= NaN` would accept any target.
        infinity = _INFINITY_LITERALS[self.dialect]
        return (
            f"({self._value_equality(src_expr, tgt_expr)} OR (ABS({src_expr}) < {infinity} AND "
            f"ABS({tgt_expr} - {src_expr}) <= {abs_tol} + ({rel_tol} * ABS({src_expr}))))"
        )

    def _edit_distance_predicate(self, src_expr: str, tgt_expr: str, limit: int) -> str:
        """Build the predicate matching text within a Levenshtein distance."""
        distance = _EDIT_DISTANCE_FUNCTIONS.get(self.dialect)
        if distance is None:
            raise ConfigError(
                f"max_levenshtein_distance cannot be pushed down to {self.dialect.value}: "
                f"{_EDIT_DISTANCE_REFUSALS[self.dialect]}. Compare locally instead "
                "(pushdown: false)."
            )
        return f"({src_expr} = {tgt_expr} OR {distance}({src_expr}, {tgt_expr}) <= {limit!r})"

    def _loosened_predicate(
        self, src_expr: str, tgt_expr: str, rule: DiffRule, *, wide: bool = False
    ) -> str | None:
        """Build the stage 8 predicate for a rule that loosens equality."""
        if rule.min_jaro_winkler_similarity is not None:
            raise ConfigError(
                "min_jaro_winkler_similarity has no SQL translation: Snowflake's "
                "JAROWINKLER_SIMILARITY ignores case and returns a whole number from 0 "
                "to 100, and Databricks has no Jaro-Winkler function. Use "
                "max_levenshtein_distance, or compare the tables locally."
            )
        if rule.absolute_tolerance or rule.relative_tolerance:
            return self._numeric_predicate(src_expr, tgt_expr, rule, wide=wide)
        if rule.max_levenshtein_distance is not None:
            return self._edit_distance_predicate(src_expr, tgt_expr, rule.max_levenshtein_distance)
        return None

    def _compare(
        self,
        src_expr: str,
        tgt_expr: str,
        rule: DiffRule,
        *,
        wide: bool = False,
        drift: bool = False,
    ) -> str:
        """Build the final match predicate, including null-safe equality."""
        if drift:
            if rule.treat_null_as_equal:
                return f"({src_expr} IS NULL AND {tgt_expr} IS NULL)"
            return "FALSE"
        loosened = self._loosened_predicate(src_expr, tgt_expr, rule, wide=wide)
        if loosened is not None:
            if rule.treat_null_as_equal:
                return f"({src_expr} IS NULL AND {tgt_expr} IS NULL) OR ({loosened})"
            return loosened

        if rule.treat_null_as_equal:
            if self.dialect is SQLDialect.SNOWFLAKE:
                return f"EQUAL_NULL({src_expr}, {tgt_expr})"
            if self.dialect is SQLDialect.DATABRICKS:
                return f"{src_expr} <=> {tgt_expr}"
            # DuckDB, BigQuery, and Postgres; BigQuery's also treats two NaNs as equal.
            return f"{src_expr} IS NOT DISTINCT FROM {tgt_expr}"

        return self._value_equality(src_expr, tgt_expr)

    def _value_equality(self, src_expr: str, tgt_expr: str) -> str:
        """Build equality that, like Polars, treats two NaNs as equal."""
        # BigQuery's `=` calls two NaNs different, per IEEE 754. A NULL still compares as unknown.
        if self.dialect is SQLDialect.BIGQUERY:
            return (
                f"({src_expr} = {tgt_expr} OR "
                f"({src_expr} IS NOT DISTINCT FROM {tgt_expr} AND {src_expr} IS NOT NULL))"
            )
        return f"{src_expr} = {tgt_expr}"

__init__(dialect)

Initialize a compiler for a single warehouse dialect.

Parameters:

Name Type Description Default
dialect SQLDialect

Target SQL dialect.

required
Source code in src/veridelta/connectors/sql.py
def __init__(self, dialect: SQLDialect) -> None:
    """Initialize a compiler for a single warehouse dialect.

    Args:
        dialect (SQLDialect): Target SQL dialect.
    """
    self.dialect = dialect

compile_added_query(source_table, target_table, primary_keys, *, source_types=None, target_types=None, key_rules=None)

Assemble a RIGHT JOIN anti-join for rows present only in the target.

Parameters:

Name Type Description Default
source_table str

Source relation (optionally dotted catalog path).

required
target_table str

Target relation (optionally dotted catalog path).

required
primary_keys list[str]

Join keys, spelled as the target stores them.

required
source_types ColumnTypes | None

Probed source dtypes.

None
target_types ColumnTypes | None

Probed target dtypes.

None
key_rules Sequence[DiffRule] | None

Key normalization, as for compile_query.

None

Returns:

Name Type Description
str str

Target keys whose normalized value has no source counterpart.

Raises:

Type Description
ConnectorError

If tables or keys are empty, or a key rule does not name exactly one primary key.

Source code in src/veridelta/connectors/sql.py
def compile_added_query(
    self,
    source_table: str,
    target_table: str,
    primary_keys: list[str],
    *,
    source_types: ColumnTypes | None = None,
    target_types: ColumnTypes | None = None,
    key_rules: Sequence[DiffRule] | None = None,
) -> str:
    """Assemble a RIGHT JOIN anti-join for rows present only in the target.

    Args:
        source_table (str): Source relation (optionally dotted catalog path).
        target_table (str): Target relation (optionally dotted catalog path).
        primary_keys (list[str]): Join keys, spelled as the target stores them.
        source_types (ColumnTypes | None): Probed source dtypes.
        target_types (ColumnTypes | None): Probed target dtypes.
        key_rules (Sequence[DiffRule] | None): Key normalization, as for
            `compile_query`.

    Returns:
        str: Target keys whose normalized value has no source counterpart.

    Raises:
        ConnectorError: If tables or keys are empty, or a key rule does not
            name exactly one primary key.
    """
    return self._compile_anti_join(
        source_table,
        target_table,
        primary_keys,
        join_kind="RIGHT",
        source_types=source_types,
        target_types=target_types,
        key_rules=key_rules,
    )

compile_changed_sample_query(source_table, target_table, primary_keys, rules, *, limit, source_types=None, target_types=None, key_rules=None, wide_integers=frozenset(), type_drift=frozenset())

Assemble a query for the first changed rows, with both sides' values.

This is compile_query with values: the same normalized CTEs, join, and WHERE clause, so every sampled row is one compile_query reports as changed, and each match flag is the predicate that decided it. Each output column takes a positional alias, so a long column name cannot pass an identifier limit and a key cannot clash with a suffixed column; SampleQuery.renames gives the local engine's names back. Rows come in key order, so the same tables give the same sample.

Parameters:

Name Type Description Default
source_table str

Source relation (optionally dotted catalog path).

required
target_table str

Target relation (optionally dotted catalog path).

required
primary_keys list[str]

Join keys, spelled as the target stores them.

required
rules list[DiffRule]

Per-column semantic overrides.

required
limit int

Most rows to return, at least 1.

required
source_types ColumnTypes | None

Probed source dtypes, as for compile_query.

None
target_types ColumnTypes | None

Probed target dtypes.

None
key_rules Sequence[DiffRule] | None

Key normalization, as for compile_query.

None
wide_integers frozenset[str]

Integer columns to measure in a wide type, as for compile_query.

frozenset()
type_drift frozenset[str]

Columns strict_types fails, as for compile_query.

frozenset()

Returns:

Type Description
SampleQuery | None

SampleQuery | None: The statement and its alias names, or None when

SampleQuery | None

no column is compared, since then no row can have changed.

Raises:

Type Description
ConfigError

If a rule sets min_jaro_winkler_similarity, or a datetime_format uses a directive this dialect cannot express.

ConnectorError

If limit is not a positive int, tables or keys are empty, a rule is pattern-only, rename_to is used with multiple column_names, or a key rule does not name exactly one primary key.

Source code in src/veridelta/connectors/sql.py
def compile_changed_sample_query(
    self,
    source_table: str,
    target_table: str,
    primary_keys: list[str],
    rules: list[DiffRule],
    *,
    limit: int,
    source_types: ColumnTypes | None = None,
    target_types: ColumnTypes | None = None,
    key_rules: Sequence[DiffRule] | None = None,
    wide_integers: frozenset[str] = frozenset(),
    type_drift: frozenset[str] = frozenset(),
) -> SampleQuery | None:
    """Assemble a query for the first changed rows, with both sides' values.

    This is `compile_query` with values: the same normalized CTEs, join,
    and WHERE clause, so every sampled row is one `compile_query` reports
    as changed, and each match flag is the predicate that decided it. Each
    output column takes a positional alias, so a long column name cannot
    pass an identifier limit and a key cannot clash with a suffixed column;
    `SampleQuery.renames` gives the local engine's names back. Rows come in
    key order, so the same tables give the same sample.

    Args:
        source_table (str): Source relation (optionally dotted catalog path).
        target_table (str): Target relation (optionally dotted catalog path).
        primary_keys (list[str]): Join keys, spelled as the target stores them.
        rules (list[DiffRule]): Per-column semantic overrides.
        limit (int): Most rows to return, at least 1.
        source_types (ColumnTypes | None): Probed source dtypes, as for
            `compile_query`.
        target_types (ColumnTypes | None): Probed target dtypes.
        key_rules (Sequence[DiffRule] | None): Key normalization, as for
            `compile_query`.
        wide_integers (frozenset[str]): Integer columns to measure in a wide
            type, as for `compile_query`.
        type_drift (frozenset[str]): Columns `strict_types` fails, as for
            `compile_query`.

    Returns:
        SampleQuery | None: The statement and its alias names, or None when
        no column is compared, since then no row can have changed.

    Raises:
        ConfigError: If a rule sets `min_jaro_winkler_similarity`, or a
            `datetime_format` uses a directive this dialect cannot express.
        ConnectorError: If `limit` is not a positive `int`, tables or keys
            are empty, a rule is pattern-only, `rename_to` is used with
            multiple `column_names`, or a key rule does not name exactly one
            primary key.
    """
    rendered_limit = self._integer(limit)
    if limit < 1:
        raise ConnectorError(f"A row sample needs a LIMIT of at least 1, got {limit}.")
    keys = self._key_columns(primary_keys, key_rules)
    compared = self._compared_columns(rules)
    if not compared:
        return None

    with_clause = self._normalized_with_clause(
        source_table,
        target_table,
        [*keys, *compared],
        source_types=source_types,
        target_types=target_types,
    )
    predicates = self._column_predicates(
        compared,
        wide_integers=wide_integers,
        type_drift=type_drift,
    )
    renames: dict[str, str] = {}
    projections: list[str] = []
    for index, key in enumerate(primary_keys):
        alias = f"{SAMPLE_KEY_PREFIX}{index}"
        projections.append(f"{self._qualify(_SOURCE_ALIAS, key)} AS {self._quote_ident(alias)}")
        renames[alias] = key
    for index, ((_source_column, target_column, _rule), predicate) in enumerate(
        zip(compared, predicates, strict=True)
    ):
        outputs = (
            (SAMPLE_SOURCE_PREFIX, self._qualify(_SOURCE_ALIAS, target_column), "source"),
            (SAMPLE_TARGET_PREFIX, self._qualify(_TARGET_ALIAS, target_column), "target"),
            (SAMPLE_MATCH_PREFIX, f"COALESCE({predicate}, FALSE)", "is_match"),
        )
        for prefix, expression, suffix in outputs:
            alias = f"{prefix}{index}"
            projections.append(f"{expression} AS {self._quote_ident(alias)}")
            renames[alias] = f"{target_column}_{suffix}"
    join = self._normalized_join("INNER", primary_keys)
    order = ", ".join(
        self._quote_ident(f"{SAMPLE_KEY_PREFIX}{index}") for index in range(len(primary_keys))
    )
    statement = (
        f"{with_clause} SELECT {', '.join(projections)} {join} "
        f"WHERE {self._changed_condition(predicates)} ORDER BY {order} LIMIT {rendered_limit}"
    )
    return SampleQuery(statement, renames)

compile_column_mismatch_query(source_table, target_table, primary_keys, rules, *, source_types=None, target_types=None, key_rules=None, wide_integers=frozenset(), type_drift=frozenset())

Assemble a per-column mismatch tally over the joined rows.

Each column contributes one SUM(CASE ...) term, so a single round trip fills DiffSummary.column_mismatches the way the local engine does. COALESCE(pred, FALSE) is load-bearing: under three-valued logic a NULL predicate is neither true nor false, and without the coalesce those rows would silently count as matches instead of mismatches. The local engine resolves the same case with val_match.fill_null(False).

Parameters:

Name Type Description Default
source_table str

Source relation (optionally dotted catalog path).

required
target_table str

Target relation (optionally dotted catalog path).

required
primary_keys list[str]

Join keys present on both relations.

required
rules list[DiffRule]

Per-column semantic overrides.

required
source_types ColumnTypes | None

Probed source dtypes, used to drop null sentinels the column cannot hold. When None, sentinels are emitted unfiltered.

None
target_types ColumnTypes | None

Probed target dtypes.

None
key_rules Sequence[DiffRule] | None

Key normalization, as for compile_query.

None
wide_integers frozenset[str]

Integer columns to measure in a wide type, as for compile_query.

frozenset()
type_drift frozenset[str]

Columns strict_types fails, as for compile_query.

frozenset()

Returns:

Type Description
str | None

str | None: Single-row aggregate statement, or None when no rule

str | None

yields a comparable column, mirroring the local engine's decision to

str | None

skip the tally when there are no match expressions.

Raises:

Type Description
ConfigError

If a rule sets min_jaro_winkler_similarity, or a datetime_format uses a directive this dialect cannot express.

ConnectorError

If tables or keys are empty, a rule is pattern-only, rename_to is used with multiple column_names, or a key rule does not name exactly one primary key.

Source code in src/veridelta/connectors/sql.py
def compile_column_mismatch_query(
    self,
    source_table: str,
    target_table: str,
    primary_keys: list[str],
    rules: list[DiffRule],
    *,
    source_types: ColumnTypes | None = None,
    target_types: ColumnTypes | None = None,
    key_rules: Sequence[DiffRule] | None = None,
    wide_integers: frozenset[str] = frozenset(),
    type_drift: frozenset[str] = frozenset(),
) -> str | None:
    """Assemble a per-column mismatch tally over the joined rows.

    Each column contributes one `SUM(CASE ...)` term, so a single round trip
    fills `DiffSummary.column_mismatches` the way the local engine does.
    `COALESCE(pred, FALSE)` is load-bearing: under three-valued logic a NULL
    predicate is neither true nor false, and without the coalesce those rows
    would silently count as matches instead of mismatches. The local engine
    resolves the same case with `val_match.fill_null(False)`.

    Args:
        source_table (str): Source relation (optionally dotted catalog path).
        target_table (str): Target relation (optionally dotted catalog path).
        primary_keys (list[str]): Join keys present on both relations.
        rules (list[DiffRule]): Per-column semantic overrides.
        source_types (ColumnTypes | None): Probed source dtypes, used to drop
            null sentinels the column cannot hold. When None, sentinels are
            emitted unfiltered.
        target_types (ColumnTypes | None): Probed target dtypes.
        key_rules (Sequence[DiffRule] | None): Key normalization, as for
            `compile_query`.
        wide_integers (frozenset[str]): Integer columns to measure in a wide
            type, as for `compile_query`.
        type_drift (frozenset[str]): Columns `strict_types` fails, as for
            `compile_query`.

    Returns:
        str | None: Single-row aggregate statement, or None when no rule
        yields a comparable column, mirroring the local engine's decision to
        skip the tally when there are no match expressions.

    Raises:
        ConfigError: If a rule sets `min_jaro_winkler_similarity`, or a
            `datetime_format` uses a directive this dialect cannot express.
        ConnectorError: If tables or keys are empty, a rule is pattern-only,
            `rename_to` is used with multiple `column_names`, or a key rule
            does not name exactly one primary key.
    """
    keys = self._key_columns(primary_keys, key_rules)
    compared = self._compared_columns(rules)
    if not compared:
        return None

    with_clause = self._normalized_with_clause(
        source_table,
        target_table,
        [*keys, *compared],
        source_types=source_types,
        target_types=target_types,
    )
    predicates = self._column_predicates(
        compared,
        wide_integers=wide_integers,
        type_drift=type_drift,
    )
    terms = [
        f"SUM(CASE WHEN COALESCE({predicate}, FALSE) THEN 0 ELSE 1 END) "
        f"AS {self._quote_ident(target_column)}"
        for (_source_column, target_column, _rule), predicate in zip(
            compared, predicates, strict=True
        )
    ]
    join = self._normalized_join("INNER", primary_keys)
    return f"{with_clause} SELECT {', '.join(terms)} {join}"

compile_column_predicate(rule, source_column, target_column=None, *, source_dtype=None, target_dtype=None)

Compile a boolean match predicate for one source/target column pair.

Follows the canonical transform order documented on DiffRule, which is the single source of truth shared with the local engine. So a pushdown run and a local run evaluate the same pipeline. A setting a dialect cannot reproduce raises ConfigError instead of compiling to something that differs:

  • min_jaro_winkler_similarity, on every dialect;
  • max_levenshtein_distance, on Postgres and DuckDB;
  • datetime_format on Postgres, or with a directive the dialect lacks;
  • a regex_replace replacement that refers to a group by name.

Parameters:

Name Type Description Default
rule DiffRule

Semantic comparison overrides for the column.

required
source_column str

Column name on the source relation.

required
target_column str | None

Column name on the target relation. Defaults to source_column when omitted.

None
source_dtype DataType | None

Probed source dtype, used to drop sentinels the column cannot hold. Each side is filtered separately because the two relations can disagree on a type. When None, sentinels are emitted unfiltered.

None
target_dtype DataType | None

Probed target dtype.

None

Returns:

Name Type Description
str str

Boolean SQL expression that is true when the column values match.

Raises:

Type Description
ConfigError

If the rule sets one of the settings listed above that this dialect refuses.

ConnectorError

If identifiers are empty or not allowlisted.

Source code in src/veridelta/connectors/sql.py
def compile_column_predicate(
    self,
    rule: DiffRule,
    source_column: str,
    target_column: str | None = None,
    *,
    source_dtype: pl.DataType | None = None,
    target_dtype: pl.DataType | None = None,
) -> str:
    """Compile a boolean match predicate for one source/target column pair.

    Follows the canonical transform order documented on `DiffRule`, which is
    the single source of truth shared with the local engine. So a pushdown
    run and a local run evaluate the same pipeline. A setting a dialect
    cannot reproduce raises `ConfigError` instead of compiling to something
    that differs:

    - `min_jaro_winkler_similarity`, on every dialect;
    - `max_levenshtein_distance`, on Postgres and DuckDB;
    - `datetime_format` on Postgres, or with a directive the dialect lacks;
    - a `regex_replace` replacement that refers to a group by name.

    Args:
        rule (DiffRule): Semantic comparison overrides for the column.
        source_column (str): Column name on the source relation.
        target_column (str | None): Column name on the target relation. Defaults
            to `source_column` when omitted.
        source_dtype (pl.DataType | None): Probed source dtype, used to drop
            sentinels the column cannot hold. Each side is filtered
            separately because the two relations can disagree on a type.
            When None, sentinels are emitted unfiltered.
        target_dtype (pl.DataType | None): Probed target dtype.

    Returns:
        str: Boolean SQL expression that is true when the column values match.

    Raises:
        ConfigError: If the rule sets one of the settings listed above that
            this dialect refuses.
        ConnectorError: If identifiers are empty or not allowlisted.
    """
    tgt_name = target_column if target_column is not None else source_column
    src_expr = self._normalize_expr(
        self._qualify(_SOURCE_ALIAS, source_column),
        rule,
        source_dtype,
        is_source=True,
    )
    tgt_expr = self._normalize_expr(
        self._qualify(_TARGET_ALIAS, tgt_name),
        rule,
        target_dtype,
        is_source=False,
    )
    return self._compare(src_expr, tgt_expr, rule)

compile_count_query(table)

Assemble a total row count query for one relation.

The count supplies the denominator for DiffSummary.mismatch_ratio, so warehouse runs honor threshold the same way local comparisons do.

Parameters:

Name Type Description Default
table str

Relation to count (optionally dotted catalog path).

required

Returns:

Name Type Description
str str

SELECT COUNT(*) AS alias FROM relation with the alias quoted

str

for the active dialect.

Raises:

Type Description
ConnectorError

If the relation name is empty or not allowlisted.

Source code in src/veridelta/connectors/sql.py
def compile_count_query(self, table: str) -> str:
    """Assemble a total row count query for one relation.

    The count supplies the denominator for `DiffSummary.mismatch_ratio`, so
    warehouse runs honor `threshold` the same way local comparisons do.

    Args:
        table (str): Relation to count (optionally dotted catalog path).

    Returns:
        str: `SELECT COUNT(*) AS alias FROM relation` with the alias quoted
        for the active dialect.

    Raises:
        ConnectorError: If the relation name is empty or not allowlisted.
    """
    return (
        f"SELECT COUNT(*) AS {self._quote_ident(COUNT_ALIAS)} "
        f"FROM {self._quote_relation(table)}"
    )

compile_duplicate_key_query(table, primary_keys, *, is_source, key_rules=None, types=None)

Assemble a count of the rows whose normalized key is not unique.

The local engine asserts uniqueness after normalization and reports every row that shares its key with another. Summing the size of each key group larger than one is that same number, and GROUP BY puts NULL keys in one group, as Polars counts them as duplicates of each other.

Parameters:

Name Type Description Default
table str

Relation to check (optionally dotted catalog path).

required
primary_keys list[str]

Keys, spelled as the target stores them.

required
is_source bool

Whether the relation is the source, which reads a renamed key under its stored name and applies value_map.

required
key_rules Sequence[DiffRule] | None

Key normalization, as for compile_query.

None
types ColumnTypes | None

Probed dtypes for this relation.

None

Returns:

Name Type Description
str str

A single-row statement whose COUNT_ALIAS column is zero when

str

every normalized key is unique.

Raises:

Type Description
ConnectorError

If the table or keys are empty, or a key rule does not name exactly one primary key.

Source code in src/veridelta/connectors/sql.py
def compile_duplicate_key_query(
    self,
    table: str,
    primary_keys: list[str],
    *,
    is_source: bool,
    key_rules: Sequence[DiffRule] | None = None,
    types: ColumnTypes | None = None,
) -> str:
    """Assemble a count of the rows whose normalized key is not unique.

    The local engine asserts uniqueness after normalization and reports
    every row that shares its key with another. Summing the size of each
    key group larger than one is that same number, and `GROUP BY` puts NULL
    keys in one group, as Polars counts them as duplicates of each other.

    Args:
        table (str): Relation to check (optionally dotted catalog path).
        primary_keys (list[str]): Keys, spelled as the target stores them.
        is_source (bool): Whether the relation is the source, which reads a
            renamed key under its stored name and applies `value_map`.
        key_rules (Sequence[DiffRule] | None): Key normalization, as for
            `compile_query`.
        types (ColumnTypes | None): Probed dtypes for this relation.

    Returns:
        str: A single-row statement whose `COUNT_ALIAS` column is zero when
        every normalized key is unique.

    Raises:
        ConnectorError: If the table or keys are empty, or a key rule does
            not name exactly one primary key.
    """
    keys = self._key_columns(primary_keys, key_rules)
    alias = "src" if is_source else "tgt"
    normalized = self._normalized_select(table, alias, keys, types=types, is_source=is_source)
    key_list = ", ".join(self._quote_ident(pk) for pk in primary_keys)
    rows = self._quote_ident(DUPLICATE_ROWS_ALIAS)
    return (
        f"SELECT COALESCE(SUM({rows}), 0) AS {self._quote_ident(COUNT_ALIAS)} "
        f"FROM (SELECT COUNT(*) AS {rows} "
        f"FROM ({normalized}) AS {self._quote_ident(KEYS_ALIAS)} "
        f"GROUP BY {key_list} HAVING COUNT(*) > 1) AS {self._quote_ident(DUPLICATES_ALIAS)}"
    )

compile_missing_query(source_table, target_table, primary_keys, *, source_types=None, target_types=None, key_rules=None)

Assemble a LEFT JOIN anti-join for rows present only in the source.

Parameters:

Name Type Description Default
source_table str

Source relation (optionally dotted catalog path).

required
target_table str

Target relation (optionally dotted catalog path).

required
primary_keys list[str]

Join keys, spelled as the target stores them.

required
source_types ColumnTypes | None

Probed source dtypes.

None
target_types ColumnTypes | None

Probed target dtypes.

None
key_rules Sequence[DiffRule] | None

Key normalization, as for compile_query.

None

Returns:

Name Type Description
str str

Source keys whose normalized value has no target counterpart.

Raises:

Type Description
ConnectorError

If tables or keys are empty, or a key rule does not name exactly one primary key.

Source code in src/veridelta/connectors/sql.py
def compile_missing_query(
    self,
    source_table: str,
    target_table: str,
    primary_keys: list[str],
    *,
    source_types: ColumnTypes | None = None,
    target_types: ColumnTypes | None = None,
    key_rules: Sequence[DiffRule] | None = None,
) -> str:
    """Assemble a LEFT JOIN anti-join for rows present only in the source.

    Args:
        source_table (str): Source relation (optionally dotted catalog path).
        target_table (str): Target relation (optionally dotted catalog path).
        primary_keys (list[str]): Join keys, spelled as the target stores them.
        source_types (ColumnTypes | None): Probed source dtypes.
        target_types (ColumnTypes | None): Probed target dtypes.
        key_rules (Sequence[DiffRule] | None): Key normalization, as for
            `compile_query`.

    Returns:
        str: Source keys whose normalized value has no target counterpart.

    Raises:
        ConnectorError: If tables or keys are empty, or a key rule does not
            name exactly one primary key.
    """
    return self._compile_anti_join(
        source_table,
        target_table,
        primary_keys,
        join_kind="LEFT",
        source_types=source_types,
        target_types=target_types,
        key_rules=key_rules,
    )

compile_query(source_table, target_table, primary_keys, rules, *, source_types=None, target_types=None, key_rules=None, wide_integers=frozenset(), type_drift=frozenset())

Assemble a changed-row inner-join query from tables, keys, and rules.

Stages 1 through 7 run once per column in a pair of CTEs, keys included. The join and the match predicates then read those projected values, so each stage appears once in the statement however many predicates read it.

Parameters:

Name Type Description Default
source_table str

Source relation (optionally dotted catalog path).

required
target_table str

Target relation (optionally dotted catalog path).

required
primary_keys list[str]

Join keys, spelled as the target stores them.

required
rules list[DiffRule]

Per-column semantic overrides.

required
source_types ColumnTypes | None

Probed source dtypes, used to drop null sentinels the column cannot hold. When None, sentinels are emitted unfiltered.

None
target_types ColumnTypes | None

Probed target dtypes.

None
key_rules Sequence[DiffRule] | None

One rule per key that needs normalizing, naming the stored source column and, for a renamed key, its rename_to. Keys without one are joined as stored.

None
wide_integers frozenset[str]

Compared columns, by target name, that hold integers on both sides after normalization. Their tolerance is measured in _WIDE_INTEGER_TYPES, so a difference wider than the stored type neither wraps nor overflows.

frozenset()
type_drift frozenset[str]

Compared columns, by target name, that strict_types fails because the two sides hold different types after normalization. No value of theirs ever matches, and two NULLs meet only under treat_null_as_equal.

frozenset()

Returns:

Name Type Description
str str

SELECT ... FROM src INNER JOIN tgt ON ... WHERE NOT (...) statement

str

keeping the joined rows where at least one compared column differs.

str

Each predicate is wrapped in COALESCE(pred, FALSE) so a one-sided

str

NULL reads as a mismatch rather than as an unknown that WHERE

str

drops, matching the local engine's fill_null(False). When no

str

column is compared the statement selects no rows, since without

str

match expressions the local engine reports nothing as changed.

Raises:

Type Description
ConfigError

If a rule sets min_jaro_winkler_similarity, or a datetime_format uses a directive this dialect cannot express.

ConnectorError

If tables or keys are empty, a rule is pattern-only, rename_to is used with multiple column_names, or a key rule does not name exactly one primary key.

Source code in src/veridelta/connectors/sql.py
def compile_query(
    self,
    source_table: str,
    target_table: str,
    primary_keys: list[str],
    rules: list[DiffRule],
    *,
    source_types: ColumnTypes | None = None,
    target_types: ColumnTypes | None = None,
    key_rules: Sequence[DiffRule] | None = None,
    wide_integers: frozenset[str] = frozenset(),
    type_drift: frozenset[str] = frozenset(),
) -> str:
    """Assemble a changed-row inner-join query from tables, keys, and rules.

    Stages 1 through 7 run once per column in a pair of CTEs, keys
    included. The join and the match predicates then read those projected
    values, so each stage appears once in the statement however many
    predicates read it.

    Args:
        source_table (str): Source relation (optionally dotted catalog path).
        target_table (str): Target relation (optionally dotted catalog path).
        primary_keys (list[str]): Join keys, spelled as the target stores them.
        rules (list[DiffRule]): Per-column semantic overrides.
        source_types (ColumnTypes | None): Probed source dtypes, used to drop
            null sentinels the column cannot hold. When None, sentinels are
            emitted unfiltered.
        target_types (ColumnTypes | None): Probed target dtypes.
        key_rules (Sequence[DiffRule] | None): One rule per key that needs
            normalizing, naming the stored source column and, for a renamed
            key, its `rename_to`. Keys without one are joined as stored.
        wide_integers (frozenset[str]): Compared columns, by target name,
            that hold integers on both sides after normalization. Their
            tolerance is measured in `_WIDE_INTEGER_TYPES`, so a difference
            wider than the stored type neither wraps nor overflows.
        type_drift (frozenset[str]): Compared columns, by target name, that
            `strict_types` fails because the two sides hold different
            types after normalization. No value of theirs ever matches, and
            two NULLs meet only under `treat_null_as_equal`.

    Returns:
        str: `SELECT ... FROM src INNER JOIN tgt ON ... WHERE NOT (...)` statement
        keeping the joined rows where at least one compared column differs.
        Each predicate is wrapped in `COALESCE(pred, FALSE)` so a one-sided
        NULL reads as a mismatch rather than as an unknown that `WHERE`
        drops, matching the local engine's `fill_null(False)`. When no
        column is compared the statement selects no rows, since without
        match expressions the local engine reports nothing as changed.

    Raises:
        ConfigError: If a rule sets `min_jaro_winkler_similarity`, or a
            `datetime_format` uses a directive this dialect cannot express.
        ConnectorError: If tables or keys are empty, a rule is pattern-only,
            `rename_to` is used with multiple `column_names`, or a key rule
            does not name exactly one primary key.
    """
    keys = self._key_columns(primary_keys, key_rules)
    compared = self._compared_columns(rules)
    with_clause = self._normalized_with_clause(
        source_table,
        target_table,
        [*keys, *compared],
        source_types=source_types,
        target_types=target_types,
    )
    select_list = ", ".join(self._qualify(_SOURCE_ALIAS, pk) for pk in primary_keys)
    join = self._normalized_join("INNER", primary_keys)
    statement = f"{with_clause} SELECT {select_list} {join}"
    predicates = self._column_predicates(
        compared,
        wide_integers=wide_integers,
        type_drift=type_drift,
    )
    if not predicates:
        # Nothing to compare means nothing can have changed. Without this
        # guard the bare join would report every shared key as drift.
        return f"{statement} WHERE 1 = 0"
    return f"{statement} WHERE {self._changed_condition(predicates)}"

compile_schema_probe_query(table)

Assemble a zero-row projection used to read a relation's columns.

Parameters:

Name Type Description Default
table str

Relation to probe (optionally dotted catalog path).

required

Returns:

Name Type Description
str str

SELECT * FROM relation WHERE 1 = 0, which returns column

str

metadata without scanning rows.

Raises:

Type Description
ConnectorError

If the relation name is empty or not allowlisted.

Source code in src/veridelta/connectors/sql.py
def compile_schema_probe_query(self, table: str) -> str:
    """Assemble a zero-row projection used to read a relation's columns.

    Args:
        table (str): Relation to probe (optionally dotted catalog path).

    Returns:
        str: `SELECT * FROM relation WHERE 1 = 0`, which returns column
        metadata without scanning rows.

    Raises:
        ConnectorError: If the relation name is empty or not allowlisted.
    """
    return f"SELECT * FROM {self._quote_relation(table)} WHERE 1 = 0"

compile_value_map_query(source_table, target_table, primary_keys, rules, *, min_support, sample_fraction=1.0, source_types=None, target_types=None, key_rules=None)

Assemble one statement that counts how source and target values line up.

Keys and candidate columns are normalized in the usual CTE pair and joined once. Each column then contributes one UNION ALL branch, labeled by its position in rules rather than by name, so no column name becomes a string literal. A branch leaves out NULL sources and values the column's existing map produced, as the local engine does.

Only exact predicates run here: the target differs from the source, at least min_support rows agree, and the agreeing rows are more than half of the source value's rows. Every confidence floor is above one half, so this keeps a superset of what qualifies, and the engine applies the floor itself, in floating point exactly as locally.

Parameters:

Name Type Description Default
source_table str

Source relation (optionally dotted catalog path).

required
target_table str

Target relation (optionally dotted catalog path).

required
primary_keys list[str]

Join keys, spelled as the target stores them.

required
rules list[DiffRule]

One rule per candidate column, naming its stored source column and, when renamed, the target's name.

required
min_support int

Agreeing rows a pair needs.

required
sample_fraction float

Share of source keys to read, chosen by a hash of the normalized keys. 1 reads every row.

1.0
source_types ColumnTypes | None

Probed source dtypes.

None
target_types ColumnTypes | None

Probed target dtypes.

None
key_rules Sequence[DiffRule] | None

Key normalization, as for compile_query.

None

Returns:

Type Description
str | None

str | None: A statement returning VALUE_MAP_COLUMN_ALIAS,

str | None

VALUE_MAP_SOURCE_ALIAS, VALUE_MAP_TARGET_ALIAS,

str | None

VALUE_MAP_ROWS_ALIAS, and VALUE_MAP_AGREEING_ALIAS per pair, or

str | None

None when there is no candidate column.

Raises:

Type Description
ConnectorError

If tables or keys are empty, a rule is pattern-only, a key rule does not name exactly one primary key, or min_support is not an integer.

Source code in src/veridelta/connectors/sql.py
def compile_value_map_query(
    self,
    source_table: str,
    target_table: str,
    primary_keys: list[str],
    rules: list[DiffRule],
    *,
    min_support: int,
    sample_fraction: float = 1.0,
    source_types: ColumnTypes | None = None,
    target_types: ColumnTypes | None = None,
    key_rules: Sequence[DiffRule] | None = None,
) -> str | None:
    """Assemble one statement that counts how source and target values line up.

    Keys and candidate columns are normalized in the usual CTE pair and
    joined once. Each column then contributes one `UNION ALL` branch,
    labeled by its position in `rules` rather than by name, so no column
    name becomes a string literal. A branch leaves out NULL sources and
    values the column's existing map produced, as the local engine does.

    Only exact predicates run here: the target differs from the source, at
    least `min_support` rows agree, and the agreeing rows are more than
    half of the source value's rows. Every confidence floor is above one
    half, so this keeps a superset of what qualifies, and the engine
    applies the floor itself, in floating point exactly as locally.

    Args:
        source_table (str): Source relation (optionally dotted catalog path).
        target_table (str): Target relation (optionally dotted catalog path).
        primary_keys (list[str]): Join keys, spelled as the target stores them.
        rules (list[DiffRule]): One rule per candidate column, naming its
            stored source column and, when renamed, the target's name.
        min_support (int): Agreeing rows a pair needs.
        sample_fraction (float): Share of source keys to read, chosen by a
            hash of the normalized keys. 1 reads every row.
        source_types (ColumnTypes | None): Probed source dtypes.
        target_types (ColumnTypes | None): Probed target dtypes.
        key_rules (Sequence[DiffRule] | None): Key normalization, as for
            `compile_query`.

    Returns:
        str | None: A statement returning `VALUE_MAP_COLUMN_ALIAS`,
        `VALUE_MAP_SOURCE_ALIAS`, `VALUE_MAP_TARGET_ALIAS`,
        `VALUE_MAP_ROWS_ALIAS`, and `VALUE_MAP_AGREEING_ALIAS` per pair, or
        None when there is no candidate column.

    Raises:
        ConnectorError: If tables or keys are empty, a rule is pattern-only,
            a key rule does not name exactly one primary key, or
            `min_support` is not an integer.
    """
    keys = self._key_columns(primary_keys, key_rules)
    compared = self._compared_columns(rules)
    if not compared:
        return None
    with_clause = self._normalized_with_clause(
        source_table,
        target_table,
        [*keys, *compared],
        source_types=source_types,
        target_types=target_types,
    )
    joined = self._value_map_joined(
        [target for _source, target, _rule in compared],
        primary_keys,
        sample_fraction,
    )
    branches = " UNION ALL ".join(
        self._value_map_branch(label, rule)
        for label, (_source, _target, rule) in enumerate(compared)
    )
    return f"{with_clause}, {joined} {self._value_map_tally(branches, min_support)}"

SnowflakeConnector

Bases: _CursorSession

Snowflake SQL warehouse connector backed by the optional Snowflake extra.

connect() opens a snowflake.connector session from the frozen SnowflakeConfig; execute_pushdown runs compiler SQL on a fresh cursor and fetches the result as Arrow. Install the driver with uv add 'veridelta[snowflake]'; without it, connect() raises ConnectorError with that hint instead of an ImportError.

Attributes:

Name Type Description
compiler SQLPushdownCompiler

Snowflake-dialect compiler the engine uses to build every statement this connector executes.

Source code in src/veridelta/connectors/warehouse.py
class SnowflakeConnector(_CursorSession):
    """Snowflake SQL warehouse connector backed by the optional Snowflake extra.

    `connect()` opens a `snowflake.connector` session from the frozen
    `SnowflakeConfig`; `execute_pushdown` runs compiler SQL on a fresh cursor
    and fetches the result as Arrow. Install the driver with
    `uv add 'veridelta[snowflake]'`; without it, `connect()` raises
    `ConnectorError` with that hint instead of an `ImportError`.

    Attributes:
        compiler (SQLPushdownCompiler): Snowflake-dialect compiler the engine
            uses to build every statement this connector executes.
    """

    _backend = "Snowflake"
    _fetch = "fetch_arrow_all"
    _fetch_kwargs = _SNOWFLAKE_FETCH_KWARGS

    def __init__(self, config: SnowflakeConfig) -> None:
        """Initialize the connector with validated Snowflake settings.

        Args:
            config (SnowflakeConfig): Frozen account, warehouse, and database
                settings.
        """
        self._config = config
        self.compiler = SQLPushdownCompiler(SQLDialect.SNOWFLAKE)

    def connect(self) -> None:
        """Open a Snowflake session for subsequent pushdown statements.

        Raises:
            ConnectorError: If the Snowflake extra is missing or authentication
                fails.
        """
        missing = self._missing_extra()
        if missing is not None:
            raise ConnectorError(missing)
        config = self._config
        try:
            self._session = snowflake_connector.connect(
                account=config.account,
                user=config.user,
                warehouse=config.warehouse,
                database=config.database,
                schema=config.schema_name,
                role=config.role,
                **_snowflake_credentials(config),
            )
        except ConnectorError:
            raise
        except Exception as exc:
            logger.warning("Snowflake connection to account %s failed", config.account)
            # Not chained: a traceback would print the driver's message unmasked.
            message = mask_secrets(str(exc), config.password, config.private_key_passphrase)
            raise ConnectorError(f"Failed to connect to Snowflake: {message}") from None
        logger.info(
            "Connected to Snowflake account %s, warehouse %s",
            self._config.account,
            self._config.warehouse,
        )

    def _missing_extra(self) -> str | None:
        """Return the Snowflake install hint when its driver is not installed."""
        return _SNOWFLAKE_EXTRA if snowflake_connector is None else None

__init__(config)

Initialize the connector with validated Snowflake settings.

Parameters:

Name Type Description Default
config SnowflakeConfig

Frozen account, warehouse, and database settings.

required
Source code in src/veridelta/connectors/warehouse.py
def __init__(self, config: SnowflakeConfig) -> None:
    """Initialize the connector with validated Snowflake settings.

    Args:
        config (SnowflakeConfig): Frozen account, warehouse, and database
            settings.
    """
    self._config = config
    self.compiler = SQLPushdownCompiler(SQLDialect.SNOWFLAKE)

connect()

Open a Snowflake session for subsequent pushdown statements.

Raises:

Type Description
ConnectorError

If the Snowflake extra is missing or authentication fails.

Source code in src/veridelta/connectors/warehouse.py
def connect(self) -> None:
    """Open a Snowflake session for subsequent pushdown statements.

    Raises:
        ConnectorError: If the Snowflake extra is missing or authentication
            fails.
    """
    missing = self._missing_extra()
    if missing is not None:
        raise ConnectorError(missing)
    config = self._config
    try:
        self._session = snowflake_connector.connect(
            account=config.account,
            user=config.user,
            warehouse=config.warehouse,
            database=config.database,
            schema=config.schema_name,
            role=config.role,
            **_snowflake_credentials(config),
        )
    except ConnectorError:
        raise
    except Exception as exc:
        logger.warning("Snowflake connection to account %s failed", config.account)
        # Not chained: a traceback would print the driver's message unmasked.
        message = mask_secrets(str(exc), config.password, config.private_key_passphrase)
        raise ConnectorError(f"Failed to connect to Snowflake: {message}") from None
    logger.info(
        "Connected to Snowflake account %s, warehouse %s",
        self._config.account,
        self._config.warehouse,
    )

VerideltaConnector

Bases: ABC

The lifecycle every connector shares: connect(), close(), and the context manager.

A connector is one of two kinds, and the engine routes each source to one by its configuration:

  • A ReaderConnector reads a source for the local engine. The lakehouse connectors (DeltaLakeConnector, IcebergConnector) open a Polars scan_* handle, and the database and DuckDB connectors (DatabaseConnector, DuckDBConnector) read one table or query when they connect. Each hands the rows to the engine through lazyframe(), and the diff runs in Polars.
  • A PushdownSession runs compiled SQL where the data lives. The warehouse connectors (SnowflakeConnector, DatabricksConnector, BigQueryConnector) hold a driver session, and the pushdown sessions (PostgresPushdownSession, DuckDBPushdownSession) serve two tables that both set pushdown. The engine compiles comparison SQL with the session's compiler and calls execute_pushdown for each round-trip; results come back as Arrow wrapped in a LazyFrame.

Call connect() before anything else and close() when finished; the connector is also a context manager whose exit calls close(). After close() the connector is back in its unconnected state, so any further call raises ConnectorError until connect() runs again.

Source code in src/veridelta/connectors/base.py
class VerideltaConnector(ABC):
    """The lifecycle every connector shares: `connect()`, `close()`, and the context manager.

    A connector is one of two kinds, and the engine routes each source to one
    by its configuration:

    - A `ReaderConnector` reads a source for the local engine. The lakehouse
      connectors (`DeltaLakeConnector`, `IcebergConnector`) open a Polars
      `scan_*` handle, and the database and DuckDB connectors
      (`DatabaseConnector`, `DuckDBConnector`) read one table or query when
      they connect. Each hands the rows to the engine through `lazyframe()`,
      and the diff runs in Polars.
    - A `PushdownSession` runs compiled SQL where the data lives. The
      warehouse connectors (`SnowflakeConnector`, `DatabricksConnector`,
      `BigQueryConnector`) hold a driver session, and the pushdown sessions
      (`PostgresPushdownSession`, `DuckDBPushdownSession`) serve two tables
      that both set `pushdown`. The engine compiles comparison SQL with the
      session's `compiler` and calls `execute_pushdown` for each round-trip;
      results come back as Arrow wrapped in a LazyFrame.

    Call `connect()` before anything else and `close()` when finished; the
    connector is also a context manager whose exit calls `close()`. After
    `close()` the connector is back in its unconnected state, so any further
    call raises `ConnectorError` until `connect()` runs again.
    """

    @abstractmethod
    def connect(self) -> None:
        """Open the driver session, the scan, or the read that later calls use.

        Raises:
            ConnectorError: If the backend cannot be reached or its extra is missing.
        """

    def close(self) -> None:  # noqa: B027 - deliberate no-op default, see below
        """Release what `connect()` opened.

        Safe to call before `connect()` and safe to call twice. The default
        holds no resources; connectors that open a driver session, a scan, or
        a read override it. It is not abstract, so a subclass that holds
        nothing need not define it.
        """

    def __enter__(self) -> Self:
        """Return the connector unchanged; `connect()` stays an explicit call.

        Returns:
            The connector itself, so `with SnowflakeConnector(cfg) as c:` binds it.
        """
        return self

    def __exit__(
        self,
        exc_type: type[BaseException] | None,
        exc: BaseException | None,
        traceback: TracebackType | None,
    ) -> None:
        """Close the connector when leaving the context, error or not.

        Args:
            exc_type (type[BaseException] | None): Pending exception type.
            exc (BaseException | None): Pending exception.
            traceback (TracebackType | None): Pending traceback.
        """
        self.close()

__enter__()

Return the connector unchanged; connect() stays an explicit call.

Returns:

Type Description
Self

The connector itself, so with SnowflakeConnector(cfg) as c: binds it.

Source code in src/veridelta/connectors/base.py
def __enter__(self) -> Self:
    """Return the connector unchanged; `connect()` stays an explicit call.

    Returns:
        The connector itself, so `with SnowflakeConnector(cfg) as c:` binds it.
    """
    return self

__exit__(exc_type, exc, traceback)

Close the connector when leaving the context, error or not.

Parameters:

Name Type Description Default
exc_type type[BaseException] | None

Pending exception type.

required
exc BaseException | None

Pending exception.

required
traceback TracebackType | None

Pending traceback.

required
Source code in src/veridelta/connectors/base.py
def __exit__(
    self,
    exc_type: type[BaseException] | None,
    exc: BaseException | None,
    traceback: TracebackType | None,
) -> None:
    """Close the connector when leaving the context, error or not.

    Args:
        exc_type (type[BaseException] | None): Pending exception type.
        exc (BaseException | None): Pending exception.
        traceback (TracebackType | None): Pending traceback.
    """
    self.close()

close()

Release what connect() opened.

Safe to call before connect() and safe to call twice. The default holds no resources; connectors that open a driver session, a scan, or a read override it. It is not abstract, so a subclass that holds nothing need not define it.

Source code in src/veridelta/connectors/base.py
def close(self) -> None:  # noqa: B027 - deliberate no-op default, see below
    """Release what `connect()` opened.

    Safe to call before `connect()` and safe to call twice. The default
    holds no resources; connectors that open a driver session, a scan, or
    a read override it. It is not abstract, so a subclass that holds
    nothing need not define it.
    """

connect() abstractmethod

Open the driver session, the scan, or the read that later calls use.

Raises:

Type Description
ConnectorError

If the backend cannot be reached or its extra is missing.

Source code in src/veridelta/connectors/base.py
@abstractmethod
def connect(self) -> None:
    """Open the driver session, the scan, or the read that later calls use.

    Raises:
        ConnectorError: If the backend cannot be reached or its extra is missing.
    """

MCP server

The tools veridelta mcp serves, and the folder guard each one goes through. The SDK they run on comes with the mcp extra, and this module imports without it.

Serve Veridelta to an AI agent as Model Context Protocol tools.

veridelta mcp calls serve, which answers an agent's host over stdio through the official MCP SDK, from the mcp extra. Each tool calls a function here. A tool with a command returns the object that command prints with --json, so an agent that knows the command line knows the tools.

The person who starts the server names the folders it may read configuration files and data on this machine from, and a tool refuses a path outside them. A tool returns findings, counts, and column names, never a value from the file. Two tools return values from the data, and only when the person who starts the server allows it: then at most a set number of rows. A side's query runs only when that person allows queries too. The SDK is imported by the first build_server(), not with this module, so veridelta imports without the extra.

DEFAULT_ROW_CAP = 50 module-attribute

The most rows, or value map entries, one call returns unless the server is started with another --max-rows.

DEFAULT_ROW_LIMIT = 20 module-attribute

The rows read_discrepancies returns when a call names no limit.

INSTRUCTIONS = "Veridelta compares two datasets under the rules in a YAML configuration file. Check a file with validate_config, and fix each error it reports, before you run it with run_comparison. Use describe_schema to list a side's columns when a rule must name one. Report counts and column names, and leave row values out of a reply unless the user asks for them. Three tools return row values, only when the server allows them. This server reads files only from the folders it was started with." module-attribute

What the server tells an agent's host about itself when the host connects.

sdk = None module-attribute

The SDK, imported by the first build_server() rather than here, since it brings a web stack that the rest of Veridelta never loads. Tests set this attribute.

DiscrepancyReport

Bases: TypedDict

What read_discrepancies returns: rows of one kind, up to the cap.

Attributes:

Name Type Description
kind Literal['added', 'removed', 'changed']

added, rows only in the target; removed, rows only in the source; or changed, rows in both with a column that differs.

total int

How many rows of that kind the run found.

rows list[dict[str, Any]]

The first of them, each a JSON object. A local run's changed row holds {column}_source, {column}_target, and {column}_is_match for each compared column. A date or a time is ISO 8601 text, a decimal is text, and bytes are hex.

truncated bool

Whether total is more than rows holds.

keys_only bool

Whether the pair was compared in place, such as two warehouse tables, which brings back primary keys alone.

Source code in src/veridelta/mcp_server.py
class DiscrepancyReport(TypedDict):
    """What `read_discrepancies` returns: rows of one kind, up to the cap.

    Attributes:
        kind: `added`, rows only in the target; `removed`, rows only in the
            source; or `changed`, rows in both with a column that differs.
        total: How many rows of that kind the run found.
        rows: The first of them, each a JSON object. A local run's `changed`
            row holds `{column}_source`, `{column}_target`, and
            `{column}_is_match` for each compared column. A date or a time is
            ISO 8601 text, a decimal is text, and bytes are hex.
        truncated: Whether `total` is more than `rows` holds.
        keys_only: Whether the pair was compared in place, such as two
            warehouse tables, which brings back primary keys alone.
    """

    kind: Literal["added", "removed", "changed"]
    total: int
    rows: list[dict[str, Any]]
    truncated: bool
    keys_only: bool

ProposalReport

Bases: TypedDict

What propose_value_maps returns: the proposals, up to the cap.

Attributes:

Name Type Description
proposals list[dict[str, Any]]

Each proposal as veridelta crosswalk --json prints it. A proposal comes back whole or not at all, and the proposals stop before the one whose value_map would pass the server's cap.

total int

How many proposals there are.

truncated bool

Whether total is more than proposals holds.

Source code in src/veridelta/mcp_server.py
class ProposalReport(TypedDict):
    """What `propose_value_maps` returns: the proposals, up to the cap.

    Attributes:
        proposals: Each proposal as `veridelta crosswalk --json` prints it. A
            proposal comes back whole or not at all, and the proposals stop
            before the one whose `value_map` would pass the server's cap.
        total: How many proposals there are.
        truncated: Whether `total` is more than `proposals` holds.
    """

    proposals: list[dict[str, Any]]
    total: int
    truncated: bool

RunReport

Bases: TypedDict

What run_comparison returns: the summary veridelta run --json prints, with its verdict.

verdict is match when the comparison falls within threshold, and drift otherwise, and exit_code is what veridelta run exits with, 0 or 1. artifacts_written says whether the rows that differ were written to output_path. The fields between them are DiffSummary's.

Source code in src/veridelta/mcp_server.py
class RunReport(TypedDict):
    """What `run_comparison` returns: the summary `veridelta run --json` prints, with its verdict.

    `verdict` is `match` when the comparison falls within `threshold`, and
    `drift` otherwise, and `exit_code` is what `veridelta run` exits with, 0
    or 1. `artifacts_written` says whether the rows that differ were written
    to `output_path`. The fields between them are `DiffSummary`'s.
    """

    verdict: Literal["match", "drift"]
    exit_code: int
    total_rows_source: int
    total_rows_target: int
    added_count: int
    removed_count: int
    changed_count: int
    column_mismatches: dict[str, int]
    is_match: bool
    accepted_count: int
    total_mismatches: int
    mismatch_ratio: float
    match_rate_percentage: float
    is_perfect_match: bool
    volume_shift: int
    report_summary: str
    artifacts_written: bool

SchemaReport

Bases: TypedDict

What describe_schema returns: one side's columns, and none of its rows.

Attributes:

Name Type Description
side Literal['source', 'target']

source or target.

columns dict[str, str]

Each column's name, as stored and before normalize_column_names or a rename_to, mapped to its type as Polars names it, such as Int64, in the stored order.

Source code in src/veridelta/mcp_server.py
class SchemaReport(TypedDict):
    """What `describe_schema` returns: one side's columns, and none of its rows.

    Attributes:
        side: `source` or `target`.
        columns: Each column's name, as stored and before
            `normalize_column_names` or a `rename_to`, mapped to its type as
            Polars names it, such as `Int64`, in the stored order.
    """

    side: Literal["source", "target"]
    columns: dict[str, str]

Settings dataclass

What the person who starts the server allows, which no tool call can change.

Attributes:

Name Type Description
roots tuple[Path, ...]

The folders a tool may read a configuration file from, resolved on creation. A relative path in a tool call is read against the first. Every tool reads data on this machine only from them too.

allow_row_values bool

Whether read_discrepancies, propose_value_maps, and suggest_rules may return values from the data. Defaults to False.

max_rows int

The most rows, value map entries, or example keys one of those calls returns. At least 1. Defaults to 50.

allow_queries bool

Whether a tool may run a side's query, which runs as written with the configuration's credentials. A DuckDB file still reads other files only from under the roots. Defaults to False.

Source code in src/veridelta/mcp_server.py
@dataclass(frozen=True)
class Settings:
    """What the person who starts the server allows, which no tool call can change.

    Attributes:
        roots (tuple[Path, ...]): The folders a tool may read a configuration
            file from, resolved on creation. A relative path in a tool call is
            read against the first. Every tool reads data on this machine only
            from them too.
        allow_row_values (bool): Whether `read_discrepancies`,
            `propose_value_maps`, and `suggest_rules` may return values from the
            data. Defaults to False.
        max_rows (int): The most rows, value map entries, or example keys one of
            those calls returns. At least 1. Defaults to 50.
        allow_queries (bool): Whether a tool may run a side's `query`, which
            runs as written with the configuration's credentials. A DuckDB file
            still reads other files only from under the roots. Defaults to
            False.
    """

    roots: tuple[Path, ...]
    allow_row_values: bool = False
    max_rows: int = DEFAULT_ROW_CAP
    allow_queries: bool = False

    def __post_init__(self) -> None:
        """Resolve every root, so a path compares with them as the filesystem does.

        Raises:
            ConfigError: If no root is given, or `max_rows` is below 1.
        """
        if not self.roots:
            raise ConfigError("The MCP server needs at least one folder to read from.")
        if self.max_rows < 1:
            raise ConfigError(f"The MCP server's row cap must be at least 1, got {self.max_rows}.")
        object.__setattr__(self, "roots", tuple(Path(root).resolve() for root in self.roots))

__post_init__()

Resolve every root, so a path compares with them as the filesystem does.

Raises:

Type Description
ConfigError

If no root is given, or max_rows is below 1.

Source code in src/veridelta/mcp_server.py
def __post_init__(self) -> None:
    """Resolve every root, so a path compares with them as the filesystem does.

    Raises:
        ConfigError: If no root is given, or `max_rows` is below 1.
    """
    if not self.roots:
        raise ConfigError("The MCP server needs at least one folder to read from.")
    if self.max_rows < 1:
        raise ConfigError(f"The MCP server's row cap must be at least 1, got {self.max_rows}.")
    object.__setattr__(self, "roots", tuple(Path(root).resolve() for root in self.roots))

SuggestionReport

Bases: TypedDict

What suggest_rules returns: the suggestions, up to the cap.

Attributes:

Name Type Description
suggestions list[dict[str, Any]]

Each suggestion as veridelta suggest --json prints it. A suggestion comes back whole or not at all, and the suggestions stop before the one whose example keys would pass the server's cap.

total int

How many suggestions there are.

truncated bool

Whether total is more than suggestions holds.

Source code in src/veridelta/mcp_server.py
class SuggestionReport(TypedDict):
    """What `suggest_rules` returns: the suggestions, up to the cap.

    Attributes:
        suggestions: Each suggestion as `veridelta suggest --json` prints it. A
            suggestion comes back whole or not at all, and the suggestions stop
            before the one whose example keys would pass the server's cap.
        total: How many suggestions there are.
        truncated: Whether `total` is more than `suggestions` holds.
    """

    suggestions: list[dict[str, Any]]
    total: int
    truncated: bool

build_server(settings)

Build the server and register its tools.

Parameters:

Name Type Description Default
settings Settings

The roots every tool is held to.

required

Returns:

Name Type Description
MCPServer MCPServer

The SDK's server, ready for run().

Raises:

Type Description
VerideltaError

If the mcp extra is not installed.

Source code in src/veridelta/mcp_server.py
def build_server(settings: Settings) -> "MCPServer":
    """Build the server and register its tools.

    Args:
        settings (Settings): The roots every tool is held to.

    Returns:
        MCPServer: The SDK's server, ready for `run()`.

    Raises:
        VerideltaError: If the `mcp` extra is not installed.
    """
    loaded = _load_sdk()
    root = logging.getLogger()
    handlers, level = root.handlers[:], root.level
    try:
        server: MCPServer = loaded.server(
            "veridelta", instructions=INSTRUCTIONS, version=__version__
        )
    finally:
        # The SDK's constructor configures the root logger, which belongs to the
        # program. The command line sets Veridelta's own logging with --verbose.
        root.handlers[:] = handlers
        root.setLevel(level)

    def validate_config(
        path: Annotated[str, Field(description="The configuration file.")],
        schemas: Annotated[
            bool, Field(description="Also read each side's columns, but no rows.")
        ] = False,
        allow_missing_env: Annotated[
            bool, Field(description="Report an unset ${NAME} as a warning, not an error.")
        ] = False,
    ) -> ValidationReport:
        """Check a Veridelta configuration file for what would stop a run.

        Reads no rows. Returns what `veridelta validate --json` prints: its errors and warnings.
        """
        return _answer(
            loaded.tool_error,
            lambda: check_configuration(
                settings, path, schemas=schemas, allow_missing_env=allow_missing_env
            ),
            environment_values(settings, path),
        )

    def run_comparison(
        path: Annotated[str, Field(description="The configuration file.")],
    ) -> RunReport:
        """Compare the two datasets a Veridelta configuration file names.

        Returns what `veridelta run --json` prints, with the verdict and the exit code a run gives.
        """
        return _answer(
            loaded.tool_error,
            lambda: run_configuration(settings, path),
            environment_values(settings, path),
        )

    def describe_schema(
        path: Annotated[str, Field(description="The configuration file.")],
        side: Annotated[Literal["source", "target"], Field(description="The side to describe.")],
    ) -> SchemaReport:
        """List the columns of one side a Veridelta configuration file names, with their types.

        Reads no rows. Names are as stored, before normalize_column_names or a rename_to.
        """
        return _answer(
            loaded.tool_error,
            lambda: describe_side(settings, path, side),
            environment_values(settings, path),
        )

    def read_discrepancies(
        path: Annotated[str, Field(description="The configuration file.")],
        kind: Annotated[
            Literal["added", "removed", "changed"],
            Field(
                description=(
                    "added: rows only in the target. removed: rows only in the source. "
                    "changed: rows in both that differ."
                )
            ),
        ],
        limit: Annotated[
            int, Field(ge=1, description="The most rows to return, within the server's cap.")
        ] = DEFAULT_ROW_LIMIT,
    ) -> DiscrepancyReport:
        """Return the rows that differ between the datasets a Veridelta configuration file names.

        Runs the comparison. Needs a server started with --allow-row-values.
        """
        return _answer(
            loaded.tool_error,
            lambda: read_rows(settings, path, kind, limit),
            environment_values(settings, path),
        )

    def propose_value_maps(
        path: Annotated[str, Field(description="The configuration file.")],
        min_confidence: Annotated[
            float,
            Field(gt=0.5, le=1, description="Share of a value's rows that must agree."),
        ] = DEFAULT_MIN_CONFIDENCE,
        min_support: Annotated[
            int, Field(ge=1, description="Agreeing rows an entry needs.")
        ] = DEFAULT_MIN_SUPPORT,
        sample_fraction: Annotated[
            float,
            Field(gt=0, le=1, description="Share of source rows to read, chosen by primary key."),
        ] = 1.0,
    ) -> ProposalReport:
        """Propose value_map entries for columns that hold the same values in two encodings.

        Returns what veridelta crosswalk --json prints. Needs a server started with
        --allow-row-values.
        """
        return _answer(
            loaded.tool_error,
            lambda: propose_maps(
                settings,
                path,
                min_confidence=min_confidence,
                min_support=min_support,
                sample_fraction=sample_fraction,
            ),
            environment_values(settings, path),
        )

    def suggest_rules(
        path: Annotated[str, Field(description="The configuration file.")],
        max_share: Annotated[
            float,
            Field(
                gt=0,
                le=1,
                description="Largest gap a tolerance may explain, as a share of the larger value.",
            ),
        ] = DEFAULT_MAX_SHARE,
    ) -> SuggestionReport:
        """Suggest rules that would explain the differences, each with its evidence.

        Returns what veridelta suggest --json prints. Needs a server started with
        --allow-row-values.
        """
        return _answer(
            loaded.tool_error,
            lambda: suggest(settings, path, max_share=max_share),
            environment_values(settings, path),
        )

    for tool in (
        validate_config,
        run_comparison,
        describe_schema,
        read_discrepancies,
        propose_value_maps,
        suggest_rules,
    ):
        server.tool(description=inspect.getdoc(tool))(tool)
    return server

check_configuration(settings, path, *, schemas=False, allow_missing_env=False)

Check a configuration file for what would stop a run, as validate --json does.

Parameters:

Name Type Description Default
settings Settings

The roots the server was started with.

required
path str

The configuration file, under a root.

required
schemas bool

Whether to also connect and check the rules against each side's columns, which reads no rows.

False
allow_missing_env bool

Whether an unset ${NAME} is a warning rather than an error, so a file can be checked without its secrets.

False

Returns:

Name Type Description
ValidationReport ValidationReport

The errors and warnings, with the resolved path.

Raises:

Type Description
ConfigError

If the path is outside the roots, or, with schemas, the data a side reads on this machine is.

Source code in src/veridelta/mcp_server.py
def check_configuration(
    settings: Settings, path: str, *, schemas: bool = False, allow_missing_env: bool = False
) -> ValidationReport:
    """Check a configuration file for what would stop a run, as `validate --json` does.

    Args:
        settings (Settings): The roots the server was started with.
        path (str): The configuration file, under a root.
        schemas (bool): Whether to also connect and check the rules against each
            side's columns, which reads no rows.
        allow_missing_env (bool): Whether an unset `${NAME}` is a warning rather
            than an error, so a file can be checked without its secrets.

    Returns:
        ValidationReport: The errors and warnings, with the resolved path.

    Raises:
        ConfigError: If the path is outside the roots, or, with `schemas`, the
            data a side reads on this machine is.
    """
    resolved = resolve_path(settings, path)
    if schemas:
        try:
            _, source, target = load_config(resolved, unset_env=[] if allow_missing_env else None)
        except ConfigError:
            # The check reports why the file does not load, and then reads no columns.
            pass
        else:
            check_data_paths(settings, source, target)
    with sandboxed(settings.roots):
        findings = DiffEngine.check_config_file(
            resolved, schemas=schemas, allow_missing_env=allow_missing_env
        )
    return validation_report(str(resolved), findings)

check_data_paths(settings, source, target)

Refuse a side whose data on this machine lies outside the roots.

Every tool that opens a side calls this first, since a column name or an error message can carry a file's text as a row does. The paths are expanded and resolved as the readers do, so neither ~ nor a link leads out. Data on another machine, such as an object store, a database server, or a warehouse, is read as the command line reads it.

Parameters:

Name Type Description Default
settings Settings

The roots the server was started with.

required
source SourceRef

The source configuration.

required
target SourceRef

The target configuration.

required

Raises:

Type Description
ConfigError

If a file or folder a side reads is outside every root.

Source code in src/veridelta/mcp_server.py
def check_data_paths(settings: Settings, source: SourceRef, target: SourceRef) -> None:
    """Refuse a side whose data on this machine lies outside the roots.

    Every tool that opens a side calls this first, since a column name or an
    error message can carry a file's text as a row does. The paths are
    expanded and resolved as the readers do, so neither `~` nor a link leads
    out. Data on another machine, such as an object store, a database server,
    or a warehouse, is read as the command line reads it.

    Args:
        settings (Settings): The roots the server was started with.
        source (SourceRef): The source configuration.
        target (SourceRef): The target configuration.

    Raises:
        ConfigError: If a file or folder a side reads is outside every root.
    """
    for side, config in (("source", source), ("target", target)):
        for setting, location in _data_locations(config):
            if not _opened_inside(settings, location):
                raise ConfigError(
                    f"The {side} {setting} '{location}' is outside the folders this server "
                    f"reads from: {_roots(settings)}. A tool reads data on this machine only "
                    "from under them, so move the data into one, or ask the person who started "
                    "the server to add its folder with --root."
                )

describe_side(settings, path, side)

List one side's columns and their types, as a run reads them before its first row.

Parameters:

Name Type Description Default
settings Settings

The roots the server was started with.

required
path str

The configuration file, under a root.

required
side Literal['source', 'target']

The side to describe.

required

Returns:

Name Type Description
SchemaReport SchemaReport

The side, and each of its columns mapped to its type.

Raises:

Type Description
ConfigError

If the file or its data is outside the roots, the file does not load, or it reads the side through a query, which would have to run in full.

ConnectorError

If the side cannot be reached or read.

Source code in src/veridelta/mcp_server.py
def describe_side(settings: Settings, path: str, side: Literal["source", "target"]) -> SchemaReport:
    """List one side's columns and their types, as a run reads them before its first row.

    Args:
        settings (Settings): The roots the server was started with.
        path (str): The configuration file, under a root.
        side (Literal["source", "target"]): The side to describe.

    Returns:
        SchemaReport: The side, and each of its columns mapped to its type.

    Raises:
        ConfigError: If the file or its data is outside the roots, the file
            does not load, or it reads the side through a `query`, which would
            have to run in full.
        ConnectorError: If the side cannot be reached or read.
    """
    _, source, target = _load(settings, path)
    with sandboxed(settings.roots):
        schema = DiffEngine.read_schema(source if side == "source" else target)
    return SchemaReport(side=side, columns={name: str(dtype) for name, dtype in schema.items()})

environment_values(settings, path)

Return the value of each environment variable a configuration file references.

A tool masks each in its answer, since a path, a name, or an error built from the configuration carries the values it took. A file outside the roots, or one that cannot be read, gives none, and so does a variable that is unset or shorter than four characters.

Parameters:

Name Type Description Default
settings Settings

The roots the server was started with.

required
path str

The configuration file from the tool call.

required

Returns:

Type Description
tuple[str, ...]

tuple[str, ...]: The values to mask.

Source code in src/veridelta/mcp_server.py
def environment_values(settings: Settings, path: str) -> tuple[str, ...]:
    """Return the value of each environment variable a configuration file references.

    A tool masks each in its answer, since a path, a name, or an error built
    from the configuration carries the values it took. A file outside the
    roots, or one that cannot be read, gives none, and so does a variable
    that is unset or shorter than four characters.

    Args:
        settings (Settings): The roots the server was started with.
        path (str): The configuration file from the tool call.

    Returns:
        tuple[str, ...]: The values to mask.
    """
    try:
        text = resolve_path(settings, path).read_text(encoding="utf-8")
    except (ConfigError, OSError, UnicodeDecodeError):
        return ()
    values = (os.environ.get(name, "") for name in referenced_variables(text))
    return tuple(value for value in values if len(value) >= _SHORTEST_MASKED)

propose_maps(settings, path, *, min_confidence=DEFAULT_MIN_CONFIDENCE, min_support=DEFAULT_MIN_SUPPORT, sample_fraction=1.0)

Propose value_map entries as veridelta crosswalk --json does, up to the server's cap.

Parameters:

Name Type Description Default
settings Settings

The roots, the permission, and the cap.

required
path str

The configuration file, under a root.

required
min_confidence float

Share of a source value's rows that must agree on one target value.

DEFAULT_MIN_CONFIDENCE
min_support int

Agreeing rows an entry needs.

DEFAULT_MIN_SUPPORT
sample_fraction float

Share of source rows to read, chosen by primary key.

1.0

Returns:

Name Type Description
ProposalReport ProposalReport

The proposals, how many there are, and whether some were left out.

Raises:

Type Description
ConfigError

If the server does not allow row values, the file or its data is outside the roots, a side reads through a query the server does not allow, a threshold is out of range, or the configuration cannot be proposed from.

ConnectorError

If a source cannot be read.

Source code in src/veridelta/mcp_server.py
def propose_maps(
    settings: Settings,
    path: str,
    *,
    min_confidence: float = DEFAULT_MIN_CONFIDENCE,
    min_support: int = DEFAULT_MIN_SUPPORT,
    sample_fraction: float = 1.0,
) -> ProposalReport:
    """Propose `value_map` entries as `veridelta crosswalk --json` does, up to the server's cap.

    Args:
        settings (Settings): The roots, the permission, and the cap.
        path (str): The configuration file, under a root.
        min_confidence (float): Share of a source value's rows that must agree
            on one target value.
        min_support (int): Agreeing rows an entry needs.
        sample_fraction (float): Share of source rows to read, chosen by
            primary key.

    Returns:
        ProposalReport: The proposals, how many there are, and whether some
            were left out.

    Raises:
        ConfigError: If the server does not allow row values, the file or its
            data is outside the roots, a side reads through a `query` the
            server does not allow, a threshold is out of range, or the
            configuration cannot be proposed from.
        ConnectorError: If a source cannot be read.
    """
    _allow_rows(settings, "propose_value_maps")
    diff, source, target = _load(settings, path)
    _refuse_queries(settings, source, target)
    with sandboxed(settings.roots):
        proposals = DiffEngine.propose_value_maps_from_configs(
            diff,
            source,
            target,
            min_confidence=min_confidence,
            min_support=min_support,
            sample_fraction=sample_fraction,
        )
    kept: list[dict[str, Any]] = []
    entries = 0
    for proposal in proposals:
        entries += len(proposal.value_map)
        if entries > settings.max_rows:
            break
        kept.append(proposal.model_dump(mode="json"))
    return ProposalReport(
        proposals=kept, total=len(proposals), truncated=len(kept) < len(proposals)
    )

read_rows(settings, path, kind, limit=DEFAULT_ROW_LIMIT)

Run the comparison and return the first rows of one kind, up to the server's cap.

Parameters:

Name Type Description Default
settings Settings

The roots, the permission, and the cap.

required
path str

The configuration file, under a root.

required
kind Literal['added', 'removed', 'changed']

Which rows to return.

required
limit int

The most rows the call asks for. The server's max_rows caps it.

DEFAULT_ROW_LIMIT

Returns:

Name Type Description
DiscrepancyReport DiscrepancyReport

The rows, how many there are, and whether more were left out.

Raises:

Type Description
ConfigError

If the server does not allow row values, limit is below 1, the file, its data, or its output_path is outside the roots, a side reads through a query the server does not allow, or the configuration cannot run.

ConnectorError

If a source cannot be read.

DataIntegrityError

If a primary key repeats on either side.

Source code in src/veridelta/mcp_server.py
def read_rows(
    settings: Settings,
    path: str,
    kind: Literal["added", "removed", "changed"],
    limit: int = DEFAULT_ROW_LIMIT,
) -> DiscrepancyReport:
    """Run the comparison and return the first rows of one kind, up to the server's cap.

    Args:
        settings (Settings): The roots, the permission, and the cap.
        path (str): The configuration file, under a root.
        kind (Literal["added", "removed", "changed"]): Which rows to return.
        limit (int): The most rows the call asks for. The server's `max_rows`
            caps it.

    Returns:
        DiscrepancyReport: The rows, how many there are, and whether more were
            left out.

    Raises:
        ConfigError: If the server does not allow row values, `limit` is below
            1, the file, its data, or its `output_path` is outside the roots,
            a side reads through a `query` the server does not allow, or the
            configuration cannot run.
        ConnectorError: If a source cannot be read.
        DataIntegrityError: If a primary key repeats on either side.
    """
    _allow_rows(settings, "read_discrepancies")
    if limit < 1:
        raise ConfigError(f"limit must be at least 1, got {limit}.")
    with sandboxed(settings.roots):
        result = DiffEngine.run_from_configs(*_runnable(settings, path))
    found = {
        "added": (result.added, result.summary.added_count),
        "removed": (result.removed, result.summary.removed_count),
        "changed": (result.changed, result.summary.changed_count),
    }
    frame, total = found[kind]
    shown = frame.head(min(limit, settings.max_rows))
    return DiscrepancyReport(
        kind=kind,
        total=total,
        rows=_records(shown),
        truncated=total > shown.height,
        keys_only=result.keys_only,
    )

resolve_path(settings, path)

Resolve a path from a tool call, and refuse one outside the roots.

A relative path is read against the first root. Resolving follows symbolic links and folds .., so neither leads out of a root.

Parameters:

Name Type Description Default
settings Settings

The roots the server was started with.

required
path str

The path from the tool call.

required

Returns:

Name Type Description
Path Path

The resolved path, under one of the roots.

Raises:

Type Description
ConfigError

If the path resolves outside every root.

Source code in src/veridelta/mcp_server.py
def resolve_path(settings: Settings, path: str) -> Path:
    """Resolve a path from a tool call, and refuse one outside the roots.

    A relative path is read against the first root. Resolving follows symbolic
    links and folds `..`, so neither leads out of a root.

    Args:
        settings (Settings): The roots the server was started with.
        path (str): The path from the tool call.

    Returns:
        Path: The resolved path, under one of the roots.

    Raises:
        ConfigError: If the path resolves outside every root.
    """
    resolved = _inside(settings, path)
    if resolved is None:
        raise ConfigError(
            f"'{path}' is outside the folders this server reads from: {_roots(settings)}. "
            "Ask the person who started it to add the folder with --root."
        )
    return resolved

run_configuration(settings, path)

Compare the two datasets a configuration file names, as veridelta run --json does.

A run writes the rows that differ to output_path, so a configuration whose output_path or data on this machine lies outside the roots is refused before any row is read, as is a side's query unless the server allows queries.

Parameters:

Name Type Description Default
settings Settings

The roots the server was started with.

required
path str

The configuration file, under a root.

required

Returns:

Name Type Description
RunReport RunReport

The summary, with the verdict and the exit code.

Raises:

Type Description
ConfigError

If the file, its data, or its output_path is outside the roots, a side reads through a query the server does not allow, or the configuration cannot run as written.

ConnectorError

If a source cannot be read.

DataIntegrityError

If a primary key repeats on either side.

Source code in src/veridelta/mcp_server.py
def run_configuration(settings: Settings, path: str) -> RunReport:
    """Compare the two datasets a configuration file names, as `veridelta run --json` does.

    A run writes the rows that differ to `output_path`, so a configuration
    whose `output_path` or data on this machine lies outside the roots is
    refused before any row is read, as is a side's `query` unless the server
    allows queries.

    Args:
        settings (Settings): The roots the server was started with.
        path (str): The configuration file, under a root.

    Returns:
        RunReport: The summary, with the verdict and the exit code.

    Raises:
        ConfigError: If the file, its data, or its `output_path` is outside
            the roots, a side reads through a `query` the server does not
            allow, or the configuration cannot run as written.
        ConnectorError: If a source cannot be read.
        DataIntegrityError: If a primary key repeats on either side.
    """
    with sandboxed(settings.roots):
        summary = DiffEngine.run_from_configs(*_runnable(settings, path)).summary
    return cast(
        "RunReport",
        {
            "verdict": "match" if summary.is_match else "drift",
            "exit_code": 0 if summary.is_match else 1,
            **summary.model_dump(mode="json"),
            "artifacts_written": summary.artifacts_written,
        },
    )

serve(settings)

Answer an agent's host over stdio until it disconnects.

A configuration's relative paths resolve against the working directory, as on the command line, and the tools check them there. veridelta mcp runs the server in its first root; a program that calls this from another folder has its relative paths checked against that folder.

Parameters:

Name Type Description Default
settings Settings

The roots every tool is held to.

required

Raises:

Type Description
VerideltaError

If the mcp extra is not installed.

Source code in src/veridelta/mcp_server.py
def serve(settings: Settings) -> None:
    """Answer an agent's host over stdio until it disconnects.

    A configuration's relative paths resolve against the working directory,
    as on the command line, and the tools check them there. `veridelta mcp`
    runs the server in its first root; a program that calls this from
    another folder has its relative paths checked against that folder.

    Args:
        settings (Settings): The roots every tool is held to.

    Raises:
        VerideltaError: If the `mcp` extra is not installed.
    """
    build_server(settings).run(transport="stdio")

suggest(settings, path, *, max_share=DEFAULT_MAX_SHARE)

Suggest rules as veridelta suggest --json does, up to the server's cap.

Parameters:

Name Type Description Default
settings Settings

The roots, the permission, and the cap.

required
path str

The configuration file, under a root.

required
max_share float

Largest gap a tolerance may explain, as a share of the larger of its two values.

DEFAULT_MAX_SHARE

Returns:

Name Type Description
SuggestionReport SuggestionReport

The suggestions, how many there are, and whether some were left out.

Raises:

Type Description
ConfigError

If the server does not allow row values, the file or its data is outside the roots, a side reads through a query the server does not allow, max_share is out of range, or the pair is compared where it is stored.

ConnectorError

If a source cannot be read.

DataIntegrityError

If either dataset repeats a normalized primary key.

Source code in src/veridelta/mcp_server.py
def suggest(
    settings: Settings, path: str, *, max_share: float = DEFAULT_MAX_SHARE
) -> SuggestionReport:
    """Suggest rules as `veridelta suggest --json` does, up to the server's cap.

    Args:
        settings (Settings): The roots, the permission, and the cap.
        path (str): The configuration file, under a root.
        max_share (float): Largest gap a tolerance may explain, as a share of
            the larger of its two values.

    Returns:
        SuggestionReport: The suggestions, how many there are, and whether
            some were left out.

    Raises:
        ConfigError: If the server does not allow row values, the file or its
            data is outside the roots, a side reads through a `query` the
            server does not allow, `max_share` is out of range, or the pair is
            compared where it is stored.
        ConnectorError: If a source cannot be read.
        DataIntegrityError: If either dataset repeats a normalized primary key.
    """
    _allow_rows(settings, "suggest_rules")
    diff, source, target = _load(settings, path)
    _refuse_queries(settings, source, target)
    with sandboxed(settings.roots):
        suggestions = DiffEngine.suggest_rules_from_configs(
            diff, source, target, max_share=max_share
        )
    kept: list[dict[str, Any]] = []
    examples = 0
    for suggestion in suggestions:
        examples += len(suggestion.examples)
        if examples > settings.max_rows:
            break
        kept.append(suggestion.model_dump(mode="json"))
    return SuggestionReport(
        suggestions=kept, total=len(suggestions), truncated=len(kept) < len(suggestions)
    )

Reports

Standalone HTML reports and Markdown summaries, rendered from a DiffResult.

Standalone HTML reports and Markdown summaries for comparison results.

Renders a DiffResult into one self-contained file: no CDN reference, no build step, no runtime dependency. Veridelta runs in CI, and CI runners are often air-gapped, where a report that fetches a stylesheet from the internet renders as unstyled text at exactly the moment someone needs to read it.

The Markdown summary is the short form CI posts to a job summary or a pull request comment: the verdict, the counts, and the columns that drifted. It lists changed values only when asked, and ends with the same counts as JSON in an HTML comment, which a reader never sees and a script can parse.

DEFAULT_MAX_ROWS = 1000 module-attribute

Rows embedded per table before truncation.

A diff of ten million rows would otherwise produce an HTML file nobody can open. The report states when it has truncated, so a reader never mistakes a capped table for the whole story.

render_html(result, *, max_rows=DEFAULT_MAX_ROWS)

Render a comparison result as a standalone HTML document.

Parameters:

Name Type Description Default
result DiffResult

Completed comparison.

required
max_rows int

Rows to embed per table before truncating.

DEFAULT_MAX_ROWS

Returns:

Name Type Description
str str

A complete HTML document with no external references.

Raises:

Type Description
ConfigError

If max_rows is negative.

Source code in src/veridelta/report.py
def render_html(result: DiffResult, *, max_rows: int = DEFAULT_MAX_ROWS) -> str:
    """Render a comparison result as a standalone HTML document.

    Args:
        result (DiffResult): Completed comparison.
        max_rows (int): Rows to embed per table before truncating.

    Returns:
        str: A complete HTML document with no external references.

    Raises:
        ConfigError: If `max_rows` is negative.
    """
    if max_rows < 0:
        raise ConfigError(f"max_rows must be zero or more, got {max_rows}.")
    summary = result.summary
    verdict = "PASSED" if summary.is_match else "FAILED"
    generated = datetime.now(UTC).strftime("%Y-%m-%d %H:%M:%S UTC")

    cards = "".join(
        [
            _card("Match rate", f"{summary.match_rate_percentage}%"),
            _card("Source rows", f"{summary.total_rows_source:,}"),
            _card("Target rows", f"{summary.total_rows_target:,}"),
            _card("Added", f"{summary.added_count:,}"),
            _card("Removed", f"{summary.removed_count:,}"),
            _card("Changed", f"{summary.changed_count:,}"),
            *([_card("Accepted", f"{summary.accepted_count:,}")] if summary.accepted_count else []),
        ]
    )

    changed = result.changed
    keys_note = ""
    if result.keys_only and result.changed_sample is not None:
        changed = result.changed_sample
        keys_note = (
            "<p class='note'>This comparison ran as pushdown, inside the database that "
            f"stores both tables. Changed rows show values for the first {changed.height:,} "
            f"of {summary.changed_count:,}, in key order, fetched because "
            "<code>pushdown_sample_rows</code> is set. Added and removed rows list "
            "primary keys.</p>"
        )
    elif result.keys_only:
        keys_note = (
            "<p class='note'>This comparison ran as pushdown, inside the database that "
            "stores both tables, and never extracts rows. The tables below list "
            "primary keys rather than values.</p>"
        )

    drift = "<p class='empty'>No column-level drift.</p>"
    if summary.column_mismatches:
        ranked = sorted(summary.column_mismatches.items(), key=lambda item: -item[1])
        rows = "".join(
            f"<tr><td>{_escape(col)}</td><td>{count:,}</td></tr>" for col, count in ranked
        )
        drift = _scroll_region(
            "column-level-drift",
            "<table><thead><tr><th scope='col'>Column</th><th scope='col'>Mismatches</th>"
            f"</tr></thead><tbody>{rows}</tbody></table>",
        )

    tables = "\n".join(
        [
            _table("Changed rows", changed, max_rows),
            _table("Added rows", result.added, max_rows),
            _table("Removed rows", result.removed, max_rows),
        ]
    )

    return f"""<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Veridelta Report &mdash; {verdict}</title>
<style>{_STYLE}</style>
</head>
<body>
<main>
<h1>Veridelta Report</h1>
<p class="sub">
  <span class="verdict {verdict.lower()}">{verdict}</span> &middot; generated {generated}
</p>
{keys_note}
<div class="cards">{cards}</div>
<h2 id="column-level-drift">Column-level drift</h2>
{drift}
{tables}
</main>
<script>{_SCRIPT}</script>
</body>
</html>
"""

render_markdown(result, *, max_rows=0)

Render a comparison result as a short Markdown summary.

Parameters:

Name Type Description Default
result DiffResult

Completed comparison.

required
max_rows int

Changed values to list, lowest keys first. 0, the default, lists none, since CI posts the summary where more people may read it than may read the data.

0

Returns:

Name Type Description
str str

The verdict, a table of counts, and the top drifting columns, limited to the configured report_top_columns_limit, then any changed values asked for.

Raises:

Type Description
ConfigError

If max_rows is negative.

Source code in src/veridelta/report.py
def render_markdown(result: DiffResult, *, max_rows: int = 0) -> str:
    """Render a comparison result as a short Markdown summary.

    Args:
        result (DiffResult): Completed comparison.
        max_rows (int): Changed values to list, lowest keys first. 0, the
            default, lists none, since CI posts the summary where more people
            may read it than may read the data.

    Returns:
        str: The verdict, a table of counts, and the top drifting columns,
            limited to the configured `report_top_columns_limit`, then any
            changed values asked for.

    Raises:
        ConfigError: If `max_rows` is negative.
    """
    if max_rows < 0:
        raise ConfigError(f"max_rows must be zero or more, got {max_rows}.")
    summary = result.summary
    verdict = "PASSED" if summary.is_match else "FAILED"
    perfect = " (Perfect Match)" if summary.is_perfect_match else ""
    lines = [
        f"### Veridelta: {verdict}{perfect}",
        "",
        "| Metric | Value |",
        "| :--- | ---: |",
        f"| Match rate | {summary.match_rate_percentage}% |",
        f"| Source rows | {summary.total_rows_source:,} |",
        f"| Target rows | {summary.total_rows_target:,} |",
        f"| Volume shift | {summary.volume_shift:+,} |",
        f"| Added | {summary.added_count:,} |",
        f"| Removed | {summary.removed_count:,} |",
        f"| Changed | {summary.changed_count:,} |",
    ]
    if summary.accepted_count:
        lines.append(f"| Accepted by the baseline | {summary.accepted_count:,} |")
    if result.keys_only:
        lines += [
            "",
            "> Pushdown compared these tables inside the database that stores them, so its "
            "artifacts list primary keys only.",
        ]
    if summary.report_limit > 0:
        lines += ["", "#### Column-level drift", ""]
        lines += _drift_lines(summary)
    comment = ["", *_summary_comment(summary)]
    if max_rows > 0 and summary.changed_count > 0:
        lines += ["", "#### Changed values", ""]
        used = len("\n".join([*lines, *comment]).encode()) + 1
        lines += _value_lines(result, max_rows, _MARKDOWN_BUDGET - used)
    return "\n".join([*lines, *comment]) + "\n"

write_html(result, path, *, max_rows=DEFAULT_MAX_ROWS)

Write a standalone HTML report to disk.

Parameters:

Name Type Description Default
result DiffResult

Completed comparison.

required
path str | Path

Destination file. Parent directories are created.

required
max_rows int

Rows to embed per table before truncating.

DEFAULT_MAX_ROWS

Returns:

Name Type Description
Path Path

The file that was written.

Raises:

Type Description
ConfigError

If max_rows is negative.

Examples:

>>> import polars as pl
>>> from veridelta.engine import DiffEngine
>>> from veridelta.models import DiffConfig
>>> source = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 20.0]})
>>> target = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 21.5]})
>>> result = DiffEngine(DiffConfig(primary_keys=["id"]), source, target).run()
>>> render_html(result).startswith("<!DOCTYPE html>")
True
>>> path = write_html(result, "reports/orders.html")
Source code in src/veridelta/report.py
def write_html(result: DiffResult, path: str | Path, *, max_rows: int = DEFAULT_MAX_ROWS) -> Path:
    """Write a standalone HTML report to disk.

    Args:
        result (DiffResult): Completed comparison.
        path (str | Path): Destination file. Parent directories are created.
        max_rows (int): Rows to embed per table before truncating.

    Returns:
        Path: The file that was written.

    Raises:
        ConfigError: If `max_rows` is negative.

    Examples:
        >>> import polars as pl
        >>> from veridelta.engine import DiffEngine
        >>> from veridelta.models import DiffConfig
        >>> source = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 20.0]})
        >>> target = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 21.5]})
        >>> result = DiffEngine(DiffConfig(primary_keys=["id"]), source, target).run()
        >>> render_html(result).startswith("<!DOCTYPE html>")
        True
        >>> path = write_html(result, "reports/orders.html")  # doctest: +SKIP
    """
    destination = Path(path)
    destination.parent.mkdir(parents=True, exist_ok=True)
    destination.write_text(render_html(result, max_rows=max_rows), encoding="utf-8")
    return destination

write_markdown(result, path, *, max_rows=0)

Write the Markdown summary to disk.

Parameters:

Name Type Description Default
result DiffResult

Completed comparison.

required
path str | Path

Destination file. Parent directories are created.

required
max_rows int

Changed values to list, as for render_markdown.

0

Returns:

Name Type Description
Path Path

The file that was written.

Raises:

Type Description
ConfigError

If max_rows is negative.

Examples:

>>> import polars as pl
>>> from veridelta.engine import DiffEngine
>>> from veridelta.models import DiffConfig
>>> source = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 20.0]})
>>> target = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 21.5]})
>>> result = DiffEngine(DiffConfig(primary_keys=["id"]), source, target).run()
>>> print(render_markdown(result).splitlines()[0])
### Veridelta: FAILED
>>> path = write_markdown(result, "summary.md")
Source code in src/veridelta/report.py
def write_markdown(result: DiffResult, path: str | Path, *, max_rows: int = 0) -> Path:
    """Write the Markdown summary to disk.

    Args:
        result (DiffResult): Completed comparison.
        path (str | Path): Destination file. Parent directories are created.
        max_rows (int): Changed values to list, as for `render_markdown`.

    Returns:
        Path: The file that was written.

    Raises:
        ConfigError: If `max_rows` is negative.

    Examples:
        >>> import polars as pl
        >>> from veridelta.engine import DiffEngine
        >>> from veridelta.models import DiffConfig
        >>> source = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 20.0]})
        >>> target = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 21.5]})
        >>> result = DiffEngine(DiffConfig(primary_keys=["id"]), source, target).run()
        >>> print(render_markdown(result).splitlines()[0])
        ### Veridelta: FAILED
        >>> path = write_markdown(result, "summary.md")  # doctest: +SKIP
    """
    destination = Path(path)
    destination.parent.mkdir(parents=True, exist_ok=True)
    destination.write_text(render_markdown(result, max_rows=max_rows), encoding="utf-8")
    return destination

OpenTelemetry metrics

A run's counts, column drift, and verdict as OTLP/JSON metrics. See OpenTelemetry metrics.

OpenTelemetry metrics for comparison results.

Renders a DiffResult as one OTLP/JSON metrics export: the body an OTLP/HTTP endpoint accepts at /v1/metrics, written on a single line so the OpenTelemetry Collector's otlpjsonfile receiver can read the file as well. No OpenTelemetry SDK is needed. The JSON follows the protobuf JSON mapping that OTLP specifies, so 64-bit integers are written as strings.

Every value is a gauge: a snapshot of one run, which a backend graphs over time rather than adds up. The export carries counts, column names, and what was compared, never row values, connection URIs, credentials, or query text, since metrics usually end up in a third-party backend.

The standard OTEL_RESOURCE_ATTRIBUTES and OTEL_SERVICE_NAME variables add resource attributes, as they do for an OpenTelemetry SDK.

send_otlp_metrics posts the same export to an OTLP/HTTP endpoint, which the standard OTEL_EXPORTER_OTLP_* variables configure. Their headers often carry an API key, so no header value reaches a log line or an error, and a redirect is refused rather than followed to a second host.

SERVICE_NAME = 'veridelta' module-attribute

service.name resource attribute and instrumentation scope name.

render_otlp_metrics(result, *, config_path=None, source=None, target=None, time_unix_nano=None)

Render a comparison as an OTLP/JSON metrics export.

Parameters:

Name Type Description Default
result DiffResult

Completed comparison.

required
config_path str | Path | None

Configuration file the run used, recorded as veridelta.config.path.

None
source SourceRef | None

Source configuration, recorded by type and by table or path.

None
target SourceRef | None

Target configuration, recorded likewise.

None
time_unix_nano int | None

Observation time, in nanoseconds since the epoch. Defaults to now.

None

Returns:

Name Type Description
str str

One line of JSON holding an ExportMetricsServiceRequest.

Source code in src/veridelta/telemetry.py
def render_otlp_metrics(
    result: DiffResult,
    *,
    config_path: str | Path | None = None,
    source: SourceRef | None = None,
    target: SourceRef | None = None,
    time_unix_nano: int | None = None,
) -> str:
    """Render a comparison as an OTLP/JSON metrics export.

    Args:
        result (DiffResult): Completed comparison.
        config_path (str | Path | None): Configuration file the run used,
            recorded as `veridelta.config.path`.
        source (SourceRef | None): Source configuration, recorded by type and
            by table or path.
        target (SourceRef | None): Target configuration, recorded likewise.
        time_unix_nano (int | None): Observation time, in nanoseconds since
            the epoch. Defaults to now.

    Returns:
        str: One line of JSON holding an `ExportMetricsServiceRequest`.
    """
    observed = time.time_ns() if time_unix_nano is None else time_unix_nano
    document = {
        "resourceMetrics": [
            {
                "resource": {"attributes": _resource_attributes(config_path, source, target)},
                "scopeMetrics": [
                    {
                        "scope": {"name": SERVICE_NAME, "version": __version__},
                        "metrics": _metrics(result, observed),
                    }
                ],
            }
        ]
    }
    # ASCII escapes keep the export on one line whatever a column is called.
    return json.dumps(document, separators=(",", ":"), ensure_ascii=True, allow_nan=False)

send_otlp_metrics(result, *, config_path=None, source=None, target=None, time_unix_nano=None)

Send a comparison's OTLP/JSON metrics export to an OTLP/HTTP endpoint.

The standard variables configure the send. OTEL_EXPORTER_OTLP_METRICS_ENDPOINT is the URL as written; otherwise /v1/metrics follows OTEL_EXPORTER_OTLP_ENDPOINT, which defaults to http://localhost:4318. The HEADERS, TIMEOUT, and PROTOCOL variables follow the same pattern, with the metrics variable winning over the general one. The send is one attempt, with no retry.

Parameters:

Name Type Description Default
result DiffResult

Completed comparison.

required
config_path str | Path | None

As for render_otlp_metrics.

None
source SourceRef | None

As for render_otlp_metrics.

None
target SourceRef | None

As for render_otlp_metrics.

None
time_unix_nano int | None

As for render_otlp_metrics.

None

Returns:

Name Type Description
str str

The endpoint the export reached, with no user part or query.

Raises:

Type Description
ConfigError

If a variable holds an unusable value, or names a protocol other than http/json.

ConnectorError

If the endpoint cannot be reached, redirects, or answers with an HTTP error.

Examples:

>>> import polars as pl
>>> from veridelta.engine import DiffEngine
>>> from veridelta.models import DiffConfig
>>> source = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 20.0]})
>>> target = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 21.5]})
>>> result = DiffEngine(DiffConfig(primary_keys=["id"]), source, target).run()
>>> send_otlp_metrics(result, config_path="veridelta.yaml")
'http://localhost:4318/v1/metrics'
Source code in src/veridelta/telemetry.py
def send_otlp_metrics(
    result: DiffResult,
    *,
    config_path: str | Path | None = None,
    source: SourceRef | None = None,
    target: SourceRef | None = None,
    time_unix_nano: int | None = None,
) -> str:
    """Send a comparison's OTLP/JSON metrics export to an OTLP/HTTP endpoint.

    The standard variables configure the send. `OTEL_EXPORTER_OTLP_METRICS_ENDPOINT`
    is the URL as written; otherwise `/v1/metrics` follows `OTEL_EXPORTER_OTLP_ENDPOINT`,
    which defaults to `http://localhost:4318`. The `HEADERS`, `TIMEOUT`, and
    `PROTOCOL` variables follow the same pattern, with the metrics variable
    winning over the general one. The send is one attempt, with no retry.

    Args:
        result (DiffResult): Completed comparison.
        config_path (str | Path | None): As for `render_otlp_metrics`.
        source (SourceRef | None): As for `render_otlp_metrics`.
        target (SourceRef | None): As for `render_otlp_metrics`.
        time_unix_nano (int | None): As for `render_otlp_metrics`.

    Returns:
        str: The endpoint the export reached, with no user part or query.

    Raises:
        ConfigError: If a variable holds an unusable value, or names a
            protocol other than `http/json`.
        ConnectorError: If the endpoint cannot be reached, redirects, or
            answers with an HTTP error.

    Examples:
        >>> import polars as pl
        >>> from veridelta.engine import DiffEngine
        >>> from veridelta.models import DiffConfig
        >>> source = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 20.0]})
        >>> target = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 21.5]})
        >>> result = DiffEngine(DiffConfig(primary_keys=["id"]), source, target).run()
        >>> send_otlp_metrics(result, config_path="veridelta.yaml")  # doctest: +SKIP
        'http://localhost:4318/v1/metrics'
    """
    _check_otlp_protocol()
    endpoint = _otlp_endpoint()
    headers = _otlp_headers()
    timeout = _otlp_timeout()
    export = render_otlp_metrics(
        result,
        config_path=config_path,
        source=source,
        target=target,
        time_unix_nano=time_unix_nano,
    )
    # A query can hold a credential too, so messages name the endpoint without one.
    where = cast("str", redacted_location(endpoint))
    parts = urlsplit(endpoint)
    if headers and parts.scheme == "http" and not _on_this_machine(parts.hostname or ""):
        # Headers often carry an API key, which plain http sends readable on the way.
        logger.warning(
            "Sending OTLP headers to %s over plain http, so anyone on the network path can "
            "read them. Use an https:// endpoint.",
            where,
        )
    request = urllib.request.Request(
        endpoint,
        data=export.encode("utf-8"),
        headers={**headers, "Content-Type": "application/json"},
        method="POST",
    )
    opener = urllib.request.build_opener(_RefuseRedirects())
    started = time.perf_counter()
    try:
        with opener.open(request, timeout=timeout) as response:
            response.read()
    except urllib.error.HTTPError as exc:
        # The body is left out: a server can echo the request it refused.
        answer = (
            "a redirect, which Veridelta does not follow; set the final URL instead"
            if 300 <= exc.code < 400
            else f"HTTP {exc.code} {exc.reason}"
        )
        failure = f"the endpoint answered with {answer}"
        # The error holds the answer open, and Python 3.14 warns when one is never closed.
        exc.close()
    except TimeoutError:
        failure = f"no answer came within {timeout:g} seconds"
    except urllib.error.URLError as exc:
        reason = exc.reason
        failure = (
            f"no answer came within {timeout:g} seconds"
            if isinstance(reason, TimeoutError)
            else str(reason)
        )
    except (OSError, http.client.HTTPException) as exc:
        # Raised while reading the answer, such as a connection closed without one.
        failure = str(exc) or type(exc).__name__
    else:
        logger.info("Sent metrics to %s in %.3fs", where, time.perf_counter() - started)
        return where
    logger.warning("Sending metrics to %s failed after %.3fs", where, time.perf_counter() - started)
    raise ConnectorError(f"Sending metrics to {where} failed: {failure}.")

write_otlp_metrics(result, path, *, config_path=None, source=None, target=None, time_unix_nano=None)

Write a comparison's OTLP/JSON metrics export to disk.

The file holds one line, ending in a newline, so it can be sent as is to an OTLP/HTTP endpoint or read by the Collector's otlpjsonfile receiver.

Parameters:

Name Type Description Default
result DiffResult

Completed comparison.

required
path str | Path

Destination file. Parent directories are created.

required
config_path str | Path | None

As for render_otlp_metrics.

None
source SourceRef | None

As for render_otlp_metrics.

None
target SourceRef | None

As for render_otlp_metrics.

None
time_unix_nano int | None

As for render_otlp_metrics.

None

Returns:

Name Type Description
Path Path

The file that was written.

Examples:

>>> import polars as pl
>>> from veridelta.engine import DiffEngine
>>> from veridelta.models import DiffConfig
>>> source = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 20.0]})
>>> target = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 21.5]})
>>> result = DiffEngine(DiffConfig(primary_keys=["id"]), source, target).run()
>>> import json
>>> export = json.loads(render_otlp_metrics(result, time_unix_nano=0))
>>> [
...     metric["name"]
...     for metric in export["resourceMetrics"][0]["scopeMetrics"][0]["metrics"]
... ]
['veridelta.dataset.rows', 'veridelta.diff.rows',
 'veridelta.column.mismatched_rows', 'veridelta.diff.mismatch_ratio',
 'veridelta.diff.match']
>>> path = write_otlp_metrics(result, "metrics.json")
Source code in src/veridelta/telemetry.py
def write_otlp_metrics(
    result: DiffResult,
    path: str | Path,
    *,
    config_path: str | Path | None = None,
    source: SourceRef | None = None,
    target: SourceRef | None = None,
    time_unix_nano: int | None = None,
) -> Path:
    """Write a comparison's OTLP/JSON metrics export to disk.

    The file holds one line, ending in a newline, so it can be sent as is to an
    OTLP/HTTP endpoint or read by the Collector's `otlpjsonfile` receiver.

    Args:
        result (DiffResult): Completed comparison.
        path (str | Path): Destination file. Parent directories are created.
        config_path (str | Path | None): As for `render_otlp_metrics`.
        source (SourceRef | None): As for `render_otlp_metrics`.
        target (SourceRef | None): As for `render_otlp_metrics`.
        time_unix_nano (int | None): As for `render_otlp_metrics`.

    Returns:
        Path: The file that was written.

    Examples:
        >>> import polars as pl
        >>> from veridelta.engine import DiffEngine
        >>> from veridelta.models import DiffConfig
        >>> source = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 20.0]})
        >>> target = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 21.5]})
        >>> result = DiffEngine(DiffConfig(primary_keys=["id"]), source, target).run()
        >>> import json
        >>> export = json.loads(render_otlp_metrics(result, time_unix_nano=0))
        >>> [
        ...     metric["name"]
        ...     for metric in export["resourceMetrics"][0]["scopeMetrics"][0]["metrics"]
        ... ]
        ['veridelta.dataset.rows', 'veridelta.diff.rows',
         'veridelta.column.mismatched_rows', 'veridelta.diff.mismatch_ratio',
         'veridelta.diff.match']
        >>> path = write_otlp_metrics(result, "metrics.json")  # doctest: +SKIP
    """
    destination = Path(path)
    destination.parent.mkdir(parents=True, exist_ok=True)
    export = render_otlp_metrics(
        result,
        config_path=config_path,
        source=source,
        target=target,
        time_unix_nano=time_unix_nano,
    )
    # A bare newline everywhere, so Windows writes the same bytes as Linux.
    destination.write_text(export + "\n", encoding="utf-8", newline="\n")
    return destination

Datasets

Sample datasets for the tutorials, downloaded once and cached.

Sample datasets for the tutorials and documentation examples.

load_nyc_taxi()

Load the NYC Taxi sample dataset.

The file is downloaded from the Veridelta repository once and cached under ~/.cache/veridelta/datasets. A download must match the sample's pinned SHA-256, and stay under 1 MiB, before it is cached, and a cached copy that no longer matches, such as a corrupt one, is downloaded again. A download times out after 15 seconds.

Returns:

Type Description
DataFrame

pl.DataFrame: The sample trips.

Raises:

Type Description
DatasetError

If the download fails, or is not the published sample.

Source code in src/veridelta/datasets.py
def load_nyc_taxi() -> pl.DataFrame:
    """Load the NYC Taxi sample dataset.

    The file is downloaded from the Veridelta repository once and cached under
    `~/.cache/veridelta/datasets`. A download must match the sample's pinned
    SHA-256, and stay under 1 MiB, before it is cached, and a cached copy that
    no longer matches, such as a corrupt one, is downloaded again. A download
    times out after 15 seconds.

    Returns:
        pl.DataFrame: The sample trips.

    Raises:
        DatasetError: If the download fails, or is not the published sample.
    """
    cache_path = _get_cache_dir() / "sample_taxi_data.parquet"
    if not _is_the_sample(cache_path):
        if cache_path.exists():
            logger.warning(
                "The cached NYC taxi sample differs from the published one; it is evicted "
                "and downloaded again."
            )
            cache_path.unlink()
        _download(cache_path)
    return pl.read_parquet(cache_path)