API reference
The public Python interface, generated from its docstrings.
Configuration models
Pydantic models for a comparison and its sources. Build them in Python, or read them from YAML with load_config. Baseline and AcceptedChange hold the drift a run accepts; see Accepting drift.
Configuration and result models.
The Pydantic models here define every configuration setting, in YAML or in Python, and the results a comparison returns.
ArtifactFormat = Literal['csv', 'parquet', 'json', 'ndjson', 'arrow']
module-attribute
File format for exported discrepancy artifacts.
A closed set that matches the engine's writers, so a typo fails when the config loads instead of after the comparison runs. Excel can be read but not written: writing a workbook needs another dependency.
BIGQUERY_PROJECT_PATTERN = '^[a-z][a-z0-9-]{4,28}[a-z0-9]$'
module-attribute
Google Cloud project id: six to thirty lowercase letters, digits, or hyphens,
starting with a letter and not ending with a hyphen. Older domain-scoped ids,
such as example.com:project, are refused rather than guessed at.
BIGQUERY_TABLE_PATTERN = '^[A-Za-z_][A-Za-z0-9_]*(\\.[A-Za-z_][A-Za-z0-9_]*)?$'
module-attribute
BigQuery table path: a table, or a dataset and table. The project is set separately, since project ids may hold hyphens, which an identifier may not.
CastTarget = Literal['Int64', 'Float64', 'String', 'Boolean', 'Date', 'Datetime']
module-attribute
Polars type a column may be cast to before comparison.
A closed set rather than a free-form name, for two reasons. A typo fails when
the config loads instead of skipping the cast. And the SQL compiler renders
this field into CAST(x AS <type>), where a type name cannot be quoted or
bound, so only a closed set keeps configuration text out of the SQL. Every
member maps to a fixed keyword in each dialect.
DATABASE_PUSHDOWN_SCHEMES = POSTGRES_SCHEMES
module-attribute
URI schemes whose tables a database source can compare inside the database.
FindingSeverity = Literal['error', 'warning']
module-attribute
How sure a configuration check is: an error stops the run, and a warning
stops it only if the tables hold what the setting cannot handle.
POSTGRES_SCHEMES = frozenset({'postgres', 'postgresql'})
module-attribute
The URI schemes that name a Postgres database.
SQL_IDENTIFIER_SEGMENT = re.compile(SQL_IDENTIFIER_SEGMENT_PATTERN)
module-attribute
Compiled allowlist applied by the warehouse SQL compiler before quoting.
SQL_IDENTIFIER_SEGMENT_PATTERN = '^[A-Za-z_][A-Za-z0-9_]*$'
module-attribute
Unquoted SQL identifier: letter or underscore, then alphanumeric or underscore.
SQL_RELATION_PATTERN = '^[A-Za-z_][A-Za-z0-9_]*(\\.[A-Za-z_][A-Za-z0-9_]*){0,2}$'
module-attribute
Warehouse table path: one to three identifier segments joined by dots.
SchemaMode = Literal['exact', 'allow_additions', 'allow_removals', 'intersection']
module-attribute
How strictly the two datasets' columns must agree.
exact: both sides have the same columns, in any order.allow_additions: the target may add columns, but keeps every source column.allow_removals: the target may drop columns, but adds none.intersection, the default: compare only the columns on both sides.
SentinelValue = str | int | float | bool
module-attribute
One null_values entry, which keeps the type written in the config.
In YAML, -999 is an integer sentinel and "-999" a text one. Each applies only
to columns of a matching type.
SourceRef = Annotated[SourceConfig | SnowflakeConfig | DatabricksConfig | BigQueryConfig | DeltaLakeConfig | IcebergConfig | DatabaseConfig | DuckDBConfig, Field(discriminator='type')]
module-attribute
YAML/Python source or target: file, warehouse, lakehouse, database, or DuckDB.
SourceType = Literal['csv', 'parquet', 'json', 'ndjson', 'arrow', 'avro', 'excel']
module-attribute
File formats Veridelta can read.
Exactly the set LoaderFactory implements. Delta Lake is not here: it is a
table format, read through the delta source type rather than as a file.
SuggestedSetting = float | bool | str | list[SentinelValue]
module-attribute
The value of one rule setting a suggestion adds, as DiffRule takes it.
WhitespaceMode = Literal['none', 'left', 'right', 'both']
module-attribute
Which ends of a string to strip whitespace from.
none: strip nothing.left: strip leading whitespace.right: strip trailing whitespace.both: strip both ends.
AcceptedChange
Bases: BaseModel
One changed row a baseline accepts, and the columns whose drift it accepts.
Attributes:
| Name | Type | Description |
|---|---|---|
key |
dict[str, Any]
|
The row's primary key, by key column. |
columns |
list[str]
|
Compared columns whose drift on this row is accepted. |
Source code in src/veridelta/models.py
Baseline
Bases: BaseModel
Drift a run accepts: rows by kind and primary key, and changed columns by row.
veridelta run --baseline FILE reads one and leaves what it lists out of the
counts and the verdict, so a run fails only on drift the file does not list.
A changed row is accepted only in the columns its entry names: drift in any
other column still counts.
Attributes:
| Name | Type | Description |
|---|---|---|
primary_keys |
list[str]
|
The key columns every entry names, as the
configuration's |
added |
list[dict[str, Any]]
|
Keys of accepted rows only in the target. |
removed |
list[dict[str, Any]]
|
Keys of accepted rows only in the source. |
changed |
list[AcceptedChange]
|
Accepted changed rows, with their columns. |
Source code in src/veridelta/models.py
1071 1072 1073 1074 1075 1076 1077 1078 1079 1080 1081 1082 1083 1084 1085 1086 1087 1088 1089 1090 1091 1092 1093 1094 1095 1096 1097 1098 1099 1100 1101 1102 1103 1104 1105 1106 1107 1108 1109 1110 1111 1112 1113 1114 1115 1116 1117 1118 1119 1120 1121 1122 1123 1124 1125 1126 1127 1128 1129 1130 1131 1132 1133 1134 1135 1136 1137 1138 1139 1140 1141 1142 1143 1144 1145 1146 1147 1148 1149 1150 1151 1152 1153 1154 1155 1156 1157 1158 1159 1160 1161 1162 1163 1164 1165 1166 1167 1168 1169 1170 1171 1172 1173 1174 1175 1176 1177 1178 1179 1180 1181 1182 1183 1184 1185 1186 1187 1188 1189 1190 1191 1192 1193 1194 1195 1196 1197 1198 1199 1200 1201 1202 1203 1204 1205 1206 1207 1208 1209 1210 1211 1212 1213 1214 1215 1216 1217 1218 1219 1220 1221 1222 1223 1224 1225 1226 1227 1228 1229 1230 1231 1232 1233 1234 1235 1236 1237 1238 1239 1240 1241 1242 1243 1244 1245 1246 1247 1248 1249 1250 1251 1252 1253 1254 1255 1256 1257 1258 1259 1260 1261 1262 1263 1264 1265 1266 1267 1268 1269 1270 1271 1272 1273 1274 1275 1276 1277 1278 1279 1280 1281 1282 1283 | |
merged(other)
Return the entries of both, with each changed row's columns joined, in key order.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
other
|
Baseline
|
Entries to add. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
Baseline |
Baseline
|
Both, without repeats. |
Source code in src/veridelta/models.py
of(result)
classmethod
Accept every row of drift a run found, with what its baseline accepted.
Saved and read back with run --baseline, it makes the same run match.
An entry of the run's own baseline that matched nothing, such as a row
fixed since, is left out.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
result
|
DiffResult
|
A local run's result. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
Baseline |
Baseline
|
The run's drift, by kind and key. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If the run compared its data where it is stored, which returns no differing columns to accept. |
Source code in src/veridelta/models.py
of_rows(primary_keys, frames, compared_columns)
classmethod
Accept the drift in a local run's frames, in key order.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
primary_keys
|
Sequence[str]
|
The run's key columns. |
required |
frames
|
tuple[DataFrame, DataFrame, DataFrame]
|
The added,
removed, and changed rows, as |
required |
compared_columns
|
Sequence[str]
|
The columns the run compared. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
Baseline |
Baseline
|
Added and removed rows by key, and changed rows with the columns that differ on each. |
Source code in src/veridelta/models.py
read(path)
classmethod
Read a baseline from a JSON file.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
str
|
The file, such as |
required |
Returns:
| Name | Type | Description |
|---|---|---|
Baseline |
Baseline
|
The drift it accepts. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If the file cannot be read or is not a valid baseline. |
Source code in src/veridelta/models.py
without(other)
Return the entries here that other does not hold.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
other
|
Baseline
|
Entries to take away, by key, and by column on a changed row. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
Baseline |
Baseline
|
What is left. |
Source code in src/veridelta/models.py
write(path)
Write the baseline as JSON, for run --baseline to read.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
str
|
The file, such as |
required |
Returns:
| Name | Type | Description |
|---|---|---|
Path |
Path
|
The file written. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If the file cannot be written. |
Source code in src/veridelta/models.py
BigQueryConfig
Bases: BaseModel
Immutable connection settings for BigQuery warehouse pushdown.
Attributes:
| Name | Type | Description |
|---|---|---|
type |
Literal['bigquery']
|
Source kind, which selects this model. |
table |
str
|
Table to compare, as |
project |
str
|
Google Cloud project that runs the queries and holds the data. It never appears in SQL. |
dataset |
str | None
|
Default dataset for a table named without one. |
location |
str | None
|
Location the queries run in, such as |
credentials_path |
str | None
|
Service account key file. Application
Default Credentials are used when unset. Left out when the config
is printed, but kept by |
maximum_bytes_billed |
int | None
|
Fail any statement that would bill more bytes than this, instead of running it. |
Source code in src/veridelta/models.py
1426 1427 1428 1429 1430 1431 1432 1433 1434 1435 1436 1437 1438 1439 1440 1441 1442 1443 1444 1445 1446 1447 1448 1449 1450 1451 1452 1453 1454 1455 1456 1457 1458 1459 1460 1461 1462 1463 1464 1465 1466 1467 1468 1469 1470 1471 1472 1473 1474 1475 1476 1477 1478 1479 1480 1481 1482 1483 1484 1485 1486 1487 1488 1489 1490 1491 1492 1493 | |
validate_dataset()
Require a default dataset when the table is named without one.
Returns:
| Name | Type | Description |
|---|---|---|
BigQueryConfig |
BigQueryConfig
|
The validated configuration. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Source code in src/veridelta/models.py
ConfigFinding
Bases: BaseModel
One problem found by checking a configuration without running it.
Attributes:
| Name | Type | Description |
|---|---|---|
severity |
FindingSeverity
|
|
message |
str
|
What is wrong and what to do about it. |
Source code in src/veridelta/models.py
DatabaseConfig
Bases: BaseModel
Immutable settings for reading a database table or query into a local comparison.
The rows are read through ConnectorX into Polars and compared by the local
engine, so a database pairs with files, lakehouse tables, or another
database. Requires the database extra (uv add 'veridelta[database]').
Two Postgres tables on one connection can instead be compared inside the
database, without reading their rows, when both sides set pushdown.
Attributes:
| Name | Type | Description |
|---|---|---|
type |
Literal['database']
|
Source kind, which selects this model. |
uri |
str
|
ConnectorX connection URI, such as
|
password |
str | None
|
Optional password, percent-encoded into the
URI's user information when the connector reads. Left out when the
config is printed, but kept by |
table |
str | None
|
Table or view to read whole, as one to three unquoted identifier segments. |
query |
str | None
|
SQL statement to run instead, sent to the database
exactly as written. Set exactly one of |
pushdown |
bool
|
Whether to compare inside the database instead of
reading the rows. Postgres |
partition_on |
str | None
|
Integer column to split a |
partitions |
int | None
|
How many ranges to split the read into: 2
or more. Set it with |
Source code in src/veridelta/models.py
1561 1562 1563 1564 1565 1566 1567 1568 1569 1570 1571 1572 1573 1574 1575 1576 1577 1578 1579 1580 1581 1582 1583 1584 1585 1586 1587 1588 1589 1590 1591 1592 1593 1594 1595 1596 1597 1598 1599 1600 1601 1602 1603 1604 1605 1606 1607 1608 1609 1610 1611 1612 1613 1614 1615 1616 1617 1618 1619 1620 1621 1622 1623 1624 1625 1626 1627 1628 1629 1630 1631 1632 1633 1634 1635 1636 1637 1638 1639 1640 1641 1642 1643 1644 1645 1646 1647 1648 1649 1650 1651 1652 1653 1654 1655 1656 1657 1658 1659 1660 1661 1662 1663 1664 1665 1666 1667 1668 1669 1670 1671 1672 1673 1674 1675 1676 1677 1678 1679 1680 1681 1682 1683 1684 1685 1686 1687 1688 1689 1690 1691 1692 1693 1694 1695 1696 1697 1698 1699 1700 1701 1702 1703 1704 1705 1706 1707 1708 1709 1710 1711 1712 1713 1714 1715 1716 1717 1718 1719 1720 1721 1722 1723 1724 1725 1726 | |
redacted_uri
property
Return the URI with any password in it replaced by ***, and no query.
The query goes whole, since a parameter such as ?password= can carry
a credential too.
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
The URI, safe to print or log. |
__repr_args__()
Print the URI with its password masked.
repr(), str(), and rich displays all read from here, while
model_dump() and the connector get the URI as written.
Yields:
| Type | Description |
|---|---|
Iterable[tuple[str | None, Any]]
|
tuple[str | None, Any]: Each field name with the value to print. |
Source code in src/veridelta/models.py
validate_connection()
Reject a source that is ambiguous about its rows or its password.
Returns:
| Name | Type | Description |
|---|---|---|
DatabaseConfig |
DatabaseConfig
|
The validated instance. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If both or neither of |
Source code in src/veridelta/models.py
validate_partitions()
Reject a split read that names half its settings or has nothing to split.
Returns:
| Name | Type | Description |
|---|---|---|
DatabaseConfig |
DatabaseConfig
|
The validated instance. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If only one of |
Source code in src/veridelta/models.py
DatabricksConfig
Bases: BaseModel
Immutable connection settings for Databricks SQL warehouse pushdown.
Attributes:
| Name | Type | Description |
|---|---|---|
type |
Literal['databricks']
|
Source kind, which selects this model. |
table |
str
|
Fully qualified table or view to compare. |
server_hostname |
str
|
Workspace hostname for the SQL warehouse. |
http_path |
str
|
HTTP path of the SQL warehouse or cluster. |
access_token |
str | None
|
Optional personal access token. Left out
when the config is printed, but kept by |
catalog |
str | None
|
Optional Unity Catalog name. |
schema_name |
str | None
|
Optional default schema name. |
Source code in src/veridelta/models.py
DeltaLakeConfig
Bases: BaseModel
Immutable settings for a Delta Lake table scan.
Attributes:
| Name | Type | Description |
|---|---|---|
type |
Literal['delta']
|
Source kind, which selects this model. |
table_uri |
str
|
Filesystem path or object-store URI of the table. |
version |
int | None
|
Optional table version to time-travel. |
storage_options |
dict[str, str]
|
Object-store credentials and options.
Left out when the config is printed, but kept by |
Source code in src/veridelta/models.py
DiffConfig
Bases: BaseModel
Settings and rules for one comparison.
Attributes:
| Name | Type | Description |
|---|---|---|
primary_keys |
list[str]
|
Columns that join the datasets. At least one is required, and together they must be unique in each dataset. |
schema_mode |
SchemaMode
|
How strictly the two sides' columns must agree.
Defaults to |
strict_types |
bool
|
Whether a column stored as different types on the two
sides fails every row. Defaults to False, which compares such a column
anyway: two numeric types compare by value, so an integer |
normalize_column_names |
bool
|
Whether to strip whitespace from column names and lowercase them before alignment. Defaults to False. |
default_absolute_tolerance |
float
|
Absolute tolerance for numeric columns whose rule sets none. Defaults to 0. |
default_relative_tolerance |
float
|
Relative tolerance for numeric columns whose rule sets none. Defaults to 0. |
default_treat_null_as_equal |
bool
|
Whether two NULLs match in columns whose rule does not say. Defaults to True. |
default_whitespace_mode |
WhitespaceMode
|
Whitespace mode for text columns
whose rule sets none. Defaults to |
default_null_values |
list[SentinelValue]
|
Values to read as NULL in every column. Each applies only to columns whose type can hold it, and the rest are skipped without an error. |
rules |
list[DiffRule]
|
Per-column overrides. A rule naming a column wins
over a |
threshold |
float
|
Largest share of mismatched rows, from 0 to 1, that still counts as a match. Defaults to 0. |
report_top_columns_limit |
int
|
Most drifted columns to list in
|
pushdown_sample_rows |
int
|
Changed rows a pushdown run fetches with both sides' values, so the HTML report and the result can show values instead of keys. Defaults to 0, which fetches none, so no value leaves the warehouse. A local run holds every row already. |
output_path |
str | None
|
Folder to write the added, removed, and changed rows to. None, the default, writes nothing. |
output_format |
ArtifactFormat
|
File format of those rows. Defaults to
|
Source code in src/veridelta/models.py
520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 | |
apply_schema_normalization()
Lowercase and strip configured column names when normalization is enabled.
Keys, rule column_names, and rename_to are normalized the same way
the engine normalizes headers, so every name still refers to a column.
A pattern is left alone: lowercasing a regex changes what it means
(\D is not \d), so patterns are written against the lowercase names.
Rules are copied rather than edited, since the caller may still hold them.
Returns:
| Name | Type | Description |
|---|---|---|
DiffConfig |
DiffConfig
|
The configuration with normalized names. |
Source code in src/veridelta/models.py
validate_default_null_values(v)
classmethod
Reject a default sentinel that can never match, such as NaN.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
v
|
list[SentinelValue]
|
Configured default sentinels. |
required |
Returns:
| Type | Description |
|---|---|
list[SentinelValue]
|
list[SentinelValue]: The sentinels, unchanged. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If any entry is a non-finite float. |
Source code in src/veridelta/models.py
DiffResult
dataclass
A completed comparison: the counts plus the rows behind them.
DiffSummary serializes to JSON, so it cannot carry frames. This class pairs it
with the discrepancy rows the engine already materialized.
Attributes:
| Name | Type | Description |
|---|---|---|
summary |
DiffSummary
|
Counts, ratios, and the formatted report. |
added |
DataFrame
|
Rows present only in the target. |
removed |
DataFrame
|
Rows present only in the source. |
changed |
DataFrame
|
Rows present in both with at least one differing
column. A local run carries |
primary_keys |
tuple[str, ...]
|
Join keys, in configured order. |
compared_columns |
tuple[str, ...]
|
Columns compared, after renames and
exclusions. Recorded so |
keys_only |
bool
|
Whether the run was pushdown, which returns primary keys instead of rows. |
changed_sample |
DataFrame | None
|
Pushdown only, when
|
accepted |
Baseline | None
|
What a baseline accepted in this run: the rows it matched, and on each changed row the columns it accepted. None when the run had no baseline. |
Source code in src/veridelta/models.py
825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 916 917 918 919 920 921 | |
get_mismatches(column)
Isolate the rows where one column disagreed.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
column
|
str
|
Compared column to isolate, named as it appears after
any |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
pl.DataFrame: Primary keys alongside the source and target values, |
DataFrame
|
restricted to rows where this column differed. Pushdown runs return |
DataFrame
|
every changed primary key instead, since the warehouse never |
DataFrame
|
projected the values and cannot attribute a row to one column. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If the column was not part of the comparison. |
Examples:
>>> import polars as pl
>>> from veridelta.engine import DiffEngine
>>> source = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 20.0]})
>>> target = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 21.5]})
>>> result = DiffEngine(DiffConfig(primary_keys=["id"]), source, target).run()
>>> result.get_mismatches("amount")["id"].to_list()
[2]
Source code in src/veridelta/models.py
to_pandas()
Convert the changed rows to pandas for notebook use.
Returns:
| Type | Description |
|---|---|
DataFrame
|
pd.DataFrame: |
DataFrame
|
convert the same way through Polars' own |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If pandas or pyarrow is not installed. |
Source code in src/veridelta/models.py
DiffRule
Bases: BaseModel
Overrides for one or more columns, chosen by exact name or by pattern.
Each setting runs at a fixed stage of the transform order, the same in a local run and in a warehouse.
Attributes:
| Name | Type | Description |
|---|---|---|
column_names |
list[str]
|
Exact source column names the rule governs. |
pattern |
str | None
|
Regular expression that selects columns by name, such
as |
absolute_tolerance |
float | None
|
Largest absolute difference that still matches, for numeric columns. Must be finite. |
relative_tolerance |
float | None
|
Largest difference relative to the source
value that still matches, such as |
max_levenshtein_distance |
int | None
|
Most single-character insertions,
deletions, and substitutions that still match, for columns compared as
text. Needs the |
min_jaro_winkler_similarity |
float | None
|
Lowest Jaro-Winkler similarity,
above 0 and at most 1, that still matches, for columns compared as text.
Needs the |
case_insensitive |
bool | None
|
Whether to ignore case in text. |
whitespace_mode |
WhitespaceMode | None
|
Which ends of text to strip whitespace from. |
regex_replace |
dict[str, str] | None
|
|
pad_zeros |
int | None
|
Width to left-pad values to with zeros, such as |
value_map |
dict[str, str] | None
|
Source values to translate to target
values before comparison, such as |
null_values |
list[SentinelValue] | None
|
Values to read as NULL, such as
|
treat_null_as_equal |
bool | None
|
Whether two NULLs match. |
datetime_format |
str | None
|
Format for |
timezone |
str | None
|
Timezone to convert timestamps to, such as |
cast_to |
CastTarget | None
|
Polars type to cast the column to, such as
|
ignore |
bool
|
Whether to leave the column out of the comparison. Defaults to False. |
rename_to |
str | None
|
Target column name, when it differs from the source.
Valid only when |
Source code in src/veridelta/models.py
295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 | |
validate_null_values(v)
classmethod
Reject a sentinel that can never match, such as NaN.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
v
|
list[SentinelValue] | None
|
Configured sentinels. |
required |
Returns:
| Type | Description |
|---|---|
list[SentinelValue] | None
|
list[SentinelValue] | None: The sentinels, unchanged. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If any entry is a non-finite float. |
Source code in src/veridelta/models.py
validate_pattern(v)
classmethod
Reject a pattern that is not a valid regular expression.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
v
|
str | None
|
Configured pattern. |
required |
Returns:
| Type | Description |
|---|---|
str | None
|
str | None: The pattern, unchanged. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the pattern does not compile. |
Source code in src/veridelta/models.py
validate_regex_replace(v)
classmethod
Reject a regex_replace key that is not a valid regular expression.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
v
|
dict[str, str] | None
|
Patterns and their replacements. |
required |
Returns:
| Type | Description |
|---|---|
dict[str, str] | None
|
dict[str, str] | None: The mapping, unchanged. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If a pattern does not compile. |
Source code in src/veridelta/models.py
validate_similarity_measure()
Reject a rule that sets both text similarity limits.
Returns:
| Name | Type | Description |
|---|---|---|
DiffRule |
DiffRule
|
The validated rule. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If both |
Source code in src/veridelta/models.py
DiffSummary
Bases: BaseModel
Counts and verdict for one comparison.
Attributes:
| Name | Type | Description |
|---|---|---|
total_rows_source |
int
|
Rows in the source dataset. |
total_rows_target |
int
|
Rows in the target dataset. |
added_count |
int
|
Rows only in the target. |
removed_count |
int
|
Rows only in the source. |
changed_count |
int
|
Rows in both datasets with at least one differing column. |
column_mismatches |
dict[str, int]
|
Mismatched rows per compared column. |
is_match |
bool
|
Whether the mismatch ratio is within |
accepted_count |
int
|
Rows of drift a baseline accepted, which the counts
above leave out. Zero without |
total_mismatches |
int
|
Added, removed, and changed rows together. |
mismatch_ratio |
float
|
|
match_rate_percentage |
float
|
Match rate as a percentage, such as |
is_perfect_match |
bool
|
Whether nothing mismatched. |
volume_shift |
int
|
Target rows minus source rows. |
report_summary |
str
|
Plain-text report for CI logs. |
report_limit |
int
|
Most columns |
artifacts_written |
bool
|
Whether artifacts were written to |
Source code in src/veridelta/models.py
687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 | |
is_perfect_match
property
Return whether nothing mismatched under the configured rules.
Returns:
| Name | Type | Description |
|---|---|---|
bool |
bool
|
Whether |
match_rate_percentage
property
Express the match rate as a percentage.
Returns:
| Name | Type | Description |
|---|---|---|
float |
float
|
The percentage, rounded to two decimal places, such as |
mismatch_ratio
property
Divide the mismatches by the source row count.
Returns:
| Name | Type | Description |
|---|---|---|
float |
float
|
The ratio, which can exceed 1. |
report_summary
property
Format the counts and the most drifted columns as plain text for a CI log.
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
The report. |
total_mismatches
property
Count added, removed, and changed rows together.
Returns:
| Name | Type | Description |
|---|---|---|
int |
int
|
The total. |
volume_shift
property
Subtract the source row count from the target's.
Returns:
| Name | Type | Description |
|---|---|---|
int |
int
|
Target rows minus source rows. |
DuckDBConfig
Bases: BaseModel
Immutable settings for reading a DuckDB or MotherDuck table or query into a local comparison.
The rows are read through DuckDB into Polars and compared by the local
engine, so a DuckDB source pairs with files, lakehouse tables, databases,
or another DuckDB source. Requires the duckdb extra
(uv add 'veridelta[duckdb]'). Two tables in one database can instead be
compared inside DuckDB, without reading their rows, when both sides set
pushdown.
Attributes:
| Name | Type | Description |
|---|---|---|
type |
Literal['duckdb']
|
Source kind, which selects this model. |
database |
str
|
Path to a DuckDB file, which opens read-only, or a
MotherDuck database written as |
table |
str | None
|
Table or view to read whole, as one to three
unquoted identifier segments, such as |
query |
str | None
|
SQL statement to run instead, sent to DuckDB
exactly as written. Set exactly one of |
motherduck_token |
str | None
|
Token for a MotherDuck database.
When unset, the |
pushdown |
bool
|
Whether to compare inside DuckDB instead of reading
the rows. |
Source code in src/veridelta/models.py
1733 1734 1735 1736 1737 1738 1739 1740 1741 1742 1743 1744 1745 1746 1747 1748 1749 1750 1751 1752 1753 1754 1755 1756 1757 1758 1759 1760 1761 1762 1763 1764 1765 1766 1767 1768 1769 1770 1771 1772 1773 1774 1775 1776 1777 1778 1779 1780 1781 1782 1783 1784 1785 1786 1787 1788 1789 1790 1791 1792 1793 1794 1795 1796 1797 1798 1799 1800 1801 1802 1803 1804 1805 1806 1807 1808 1809 1810 1811 1812 1813 1814 1815 1816 1817 1818 1819 1820 1821 1822 1823 1824 1825 1826 1827 1828 1829 1830 1831 | |
is_motherduck
property
Return whether database names a MotherDuck database rather than a file.
Returns:
| Name | Type | Description |
|---|---|---|
bool |
bool
|
Whether |
validate_connection()
Reject a source that is ambiguous about its rows or would expose its token.
Returns:
| Name | Type | Description |
|---|---|---|
DuckDBConfig |
DuckDBConfig
|
The validated instance. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If both or neither of |
Source code in src/veridelta/models.py
IcebergConfig
Bases: BaseModel
Immutable settings for an Apache Iceberg table scan.
Attributes:
| Name | Type | Description |
|---|---|---|
type |
Literal['iceberg']
|
Source kind, which selects this model. |
table_uri |
str
|
Catalog identifier or filesystem URI of the table. |
snapshot_id |
int | None
|
Optional snapshot to time-travel. |
storage_options |
dict[str, str]
|
Object-store credentials and options.
Left out when the config is printed, but kept by |
Source code in src/veridelta/models.py
RuleSuggestion
Bases: BaseModel
A rule suggest proposes for one column, with the evidence for it.
Attributes:
| Name | Type | Description |
|---|---|---|
column |
str
|
Compared column, named as it appears after any |
settings |
dict[str, SuggestedSetting]
|
The |
differing |
int
|
Joined rows whose values in the column differ under the declared rules. |
explained |
int
|
Of those, the rows that match once the rule is added, counted by running the comparison again with it. |
largest_gap |
float | None
|
For a tolerance, the largest absolute difference among the rows it explains, and None for other settings. |
examples |
tuple[dict[str, Any], ...]
|
The primary keys of up to three rows the rule explains, lowest first. |
rule |
DiffRule
|
The rule to add: the settings of the rule that governs
the column today, if any, with |
governing_rule_index |
int | None
|
Position in |
Source code in src/veridelta/models.py
SnowflakeConfig
Bases: BaseModel
Immutable connection settings for Snowflake warehouse pushdown.
Attributes:
| Name | Type | Description |
|---|---|---|
type |
Literal['snowflake']
|
Source kind, which selects this model. |
table |
str
|
Fully qualified table or view to compare. |
account |
str
|
Snowflake account identifier. |
user |
str
|
Login name used to authenticate the session. |
warehouse |
str
|
Virtual warehouse that executes pushdown SQL. |
database |
str
|
Default database for unqualified object names. |
schema_name |
str
|
Default schema for unqualified object names. |
password |
str | None
|
Optional password or programmatic access
token, unset for a key pair or SSO. Left out when the config is
printed, but kept by |
private_key_path |
str | None
|
Optional path to a PEM private key, for key-pair sign-in instead of a password. Left out when printed. |
private_key_passphrase |
str | None
|
Passphrase of an encrypted
|
role |
str | None
|
Optional role assumed after authentication. |
Source code in src/veridelta/models.py
1320 1321 1322 1323 1324 1325 1326 1327 1328 1329 1330 1331 1332 1333 1334 1335 1336 1337 1338 1339 1340 1341 1342 1343 1344 1345 1346 1347 1348 1349 1350 1351 1352 1353 1354 1355 1356 1357 1358 1359 1360 1361 1362 1363 1364 1365 1366 1367 1368 1369 1370 1371 1372 1373 1374 1375 1376 1377 1378 1379 1380 1381 1382 1383 1384 1385 1386 1387 1388 1389 1390 1391 | |
validate_sign_in()
Reject credentials that leave unclear how the session signs in.
Returns:
| Name | Type | Description |
|---|---|---|
SnowflakeConfig |
SnowflakeConfig
|
The validated instance. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If both |
Source code in src/veridelta/models.py
SourceConfig
Bases: BaseModel
Settings for a file source.
Attributes:
| Name | Type | Description |
|---|---|---|
type |
Literal['file']
|
Source kind. A YAML file source may omit it. |
path |
str
|
Local path or URI of the file. |
format |
SourceType
|
File format, such as |
options |
dict[str, Any]
|
Keyword arguments for the Polars reader, such as
|
Source code in src/veridelta/models.py
__repr_args__()
Leave object-store credentials out of the printed reader options.
A cloud path's credentials travel in a nested storage_options map.
Printing drops that one key and keeps the other options, which are the
useful part when debugging. repr(), str(), and rich displays all
read from here, while model_dump() and the reader get the full map.
Yields:
| Type | Description |
|---|---|
Iterable[tuple[str | None, Any]]
|
tuple[str | None, Any]: Each field name with the value to print. |
Source code in src/veridelta/models.py
ValueMapEntry
Bases: BaseModel
One proposed value_map entry and the rows that support it.
Attributes:
| Name | Type | Description |
|---|---|---|
source_value |
str
|
Source text as the |
target_value |
str
|
The target value those rows compare against, as text. |
rows |
int
|
Joined rows whose source holds |
agreeing_rows |
int
|
Those rows whose target is |
confidence |
float
|
|
Source code in src/veridelta/models.py
confidence
property
Return the share of the source value's rows that agree.
Returns:
| Name | Type | Description |
|---|---|---|
float |
float
|
|
validate_counts()
Reject more agreeing rows than rows.
Returns:
| Name | Type | Description |
|---|---|---|
ValueMapEntry |
ValueMapEntry
|
The validated entry. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Source code in src/veridelta/models.py
ValueMapProposal
Bases: BaseModel
A proposed value_map for one column, with the evidence for each new entry.
Attributes:
| Name | Type | Description |
|---|---|---|
column |
str
|
Compared column, named as it appears after any |
value_map |
dict[str, str]
|
The governing rule's existing entries plus the proposed ones. |
entries |
tuple[ValueMapEntry, ...]
|
The proposed entries alone, most agreeing rows first. |
governing_rule_index |
int | None
|
Position in |
Source code in src/veridelta/models.py
to_rule()
Build a rule for the column carrying the proposed map.
When governing_rule_index is set, merge value_map into that rule
instead, since a second rule for the column would not apply.
Returns:
| Name | Type | Description |
|---|---|---|
DiffRule |
DiffRule
|
Rule naming the column, with the full proposed map. |
Source code in src/veridelta/models.py
mismatch_ratio_of(mismatches, source_rows)
Divide the mismatches by the source row count, which the threshold applies to.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
mismatches
|
int
|
Added, removed, and changed rows together. |
required |
source_rows
|
int
|
The rows the source holds. An empty source counts as one. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
float |
float
|
The ratio, which can exceed 1. |
Source code in src/veridelta/models.py
normalize_column_name(name)
Strip and lowercase a column name, as normalize_column_names asks.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str
|
A header, or a name a configuration gives. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
The name the engine compares by. |
Source code in src/veridelta/models.py
redacted_location(location)
Return a path or URL with anything secret left out, safe to print or log.
A URL keeps its scheme, host, and path. Its query goes, since it can hold a token or the signature of a pre-signed link, and so does its user part, which can hold a login, unless an Azure scheme names a container there. A path on this machine comes back as it is.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
location
|
str
|
A file path, or the URL of a file or a table. |
required |
Returns:
| Type | Description |
|---|---|
str | None
|
str | None: The location without its secrets, or None when it does not parse as a URL. |
Examples:
>>> redacted_location("s3://bucket/events.parquet?X-Amz-Signature=abc")
's3://bucket/events.parquet'
>>> redacted_location("abfss://lake@account.dfs.core.windows.net/events")
'abfss://lake@account.dfs.core.windows.net/events'
Source code in src/veridelta/models.py
Engine
DiffEngine loads, aligns, and compares two datasets, on Polars or inside a warehouse.
Load, align, and compare datasets.
Holds the DiffEngine that loads, aligns, and compares the two sides with
Polars. Its helpers live in private modules beside it, and __all__ lists the
public names, the file loaders from _reading.py among them.
DEFAULT_MAX_SHARE = 0.01
module-attribute
Largest gap a suggested tolerance may explain, as a share of the larger of its two values. A gap past it is a change, not noise, so the column gets no suggestion.
DEFAULT_MIN_CONFIDENCE = 0.95
module-attribute
Share of a source value's rows that must agree on one target value before it is proposed. Above one half, at most one target can qualify, and 5% leaves room for noise in legacy data.
DEFAULT_MIN_SUPPORT = 5
module-attribute
Agreeing rows a proposal needs, so a one-off coincidence is never offered.
ArrowLoader
Bases: BaseLoader
Streaming loader for Arrow IPC (Feather v2) files over pl.scan_ipc.
IPC files carry their schema, so no type inference runs and the dtypes the engine compares are exactly the ones the writer stored.
Source code in src/veridelta/_reading.py
load(config)
Scan an Arrow IPC file.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
SourceConfig
|
Source configuration, whose options go to |
required |
Returns:
| Type | Description |
|---|---|
LazyFrame
|
pl.LazyFrame: The unevaluated rows. |
Source code in src/veridelta/_reading.py
AvroLoader
Bases: BaseLoader
Loader for Avro object container files, read eagerly with pl.read_avro.
Polars has no lazy Avro reader, so the file is read whole and wrapped, like JSON and Excel. Avro carries its schema, so the dtypes compared are the writer's, with no inference. The reader takes a local path only.
Source code in src/veridelta/_reading.py
load(config)
Read an Avro file.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
SourceConfig
|
Source configuration, whose |
required |
Returns:
| Type | Description |
|---|---|
LazyFrame
|
pl.LazyFrame: A lazy wrapper over the rows read. |
Source code in src/veridelta/_reading.py
BaseLoader
Bases: ABC
Base class for the file loaders, each turning a SourceConfig into a LazyFrame.
Each file format has a subclass registered in LoaderFactory._loaders, keyed by
its SourceType. A loader prefers a Polars scan_* reader, so the comparison
stays lazy end to end; the eager loaders say why in their own docstrings.
SourceConfig.options reach the reader unchanged, so any keyword the Polars
function accepts is valid.
Source code in src/veridelta/_reading.py
load(config)
abstractmethod
Load a source into a LazyFrame.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
SourceConfig
|
Path, format, and reader options. |
required |
Returns:
| Type | Description |
|---|---|
LazyFrame
|
pl.LazyFrame: The unevaluated rows. |
CSVLoader
Bases: BaseLoader
Streaming CSV loader over pl.scan_csv.
Delimiters, encodings, and header handling are all controlled through
SourceConfig.options, for example {"separator": ";"}.
Source code in src/veridelta/_reading.py
load(config)
Scan a CSV file.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
SourceConfig
|
Source configuration, whose options go to |
required |
Returns:
| Type | Description |
|---|---|
LazyFrame
|
pl.LazyFrame: The unevaluated rows. |
Source code in src/veridelta/_reading.py
DiffEngine
Compare two datasets on their primary keys and report what differs.
The engine takes two Polars LazyFrame inputs and a DiffConfig, applies
the nine-stage DiffRule pipeline to each side, then joins on the primary
keys to classify rows as added (target-only), removed (source-only), or
changed (present on both sides with at least one compared column
differing). Nothing is materialized until run() collects the joins, so
the source frames can be scan_* graphs over files far larger than memory.
These entry points cover the usual situations:
DiffEngine(config, source, target).run()for frames you already hold. In-memoryDataFrameinputs must be wrapped with.lazy()first.DiffEngine.run_from_configs(diff, source, target)forSourceRefpairs from YAML. File, lakehouse, database, and DuckDB pairs load throughLoaderFactoryand run locally. A same-warehouse pair, or a Postgres or DuckDB pair that setspushdown, compiles to SQL and runs where it is stored.DiffEngine.validate_schemas(...)to enforceschema_modeand primary key presence on metadata alone, before any rows are read.DiffEngine.validate_rules(...)to also resolve every rule and build each column's comparison, still on metadata alone.DiffEngine(config, source, target).propose_value_maps(), orDiffEngine.propose_value_maps_from_configs(...)for anySourceRefpair, counted in the warehouse for same-warehouse pairs, to suggestvalue_mapentries from how the two sides' values line up, without running the comparison.
Attributes:
| Name | Type | Description |
|---|---|---|
config |
DiffConfig
|
Primary keys, rules, defaults, and threshold. |
source |
LazyFrame
|
Source side; mutated in place as alignment and normalization stages run. |
target |
LazyFrame
|
Target side, treated as the authoritative schema. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
From |
DataIntegrityError
|
From |
Source code in src/veridelta/engine.py
134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 945 946 947 948 949 950 951 952 953 954 955 956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 973 974 975 976 977 978 979 980 981 982 983 984 985 986 987 988 989 990 991 992 993 994 995 996 997 998 999 1000 1001 1002 1003 1004 1005 1006 1007 1008 1009 1010 1011 1012 1013 1014 1015 1016 1017 1018 1019 1020 1021 1022 1023 1024 1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 1052 1053 1054 1055 1056 1057 1058 1059 1060 1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 1077 1078 1079 1080 1081 1082 1083 1084 1085 1086 1087 1088 1089 1090 1091 1092 1093 1094 1095 1096 1097 1098 1099 1100 1101 1102 | |
__init__(config, source_df, target_df)
Hold the configuration and the two datasets to compare.
run() aligns the datasets and compares them.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
DiffConfig
|
Keys, rules, and settings for the comparison. |
required |
source_df
|
LazyFrame
|
Source rows, as stored. |
required |
target_df
|
LazyFrame
|
Target rows, as stored. |
required |
Source code in src/veridelta/engine.py
check_config_file(path, *, schemas=False, allow_missing_env=False)
staticmethod
Load a configuration file and check it for what would stop a run.
A file that does not load is one error finding, with the loader's
message, so every problem is reported the same way. A file that loads
gets the checks of check_configs. veridelta validate and the MCP
server's validate_config tool both report through this method.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
str | Path
|
The configuration file. |
required |
schemas
|
bool
|
Whether to also connect and check the rules against the stored columns. |
False
|
allow_missing_env
|
bool
|
Whether to read an unset |
False
|
Returns:
| Type | Description |
|---|---|
list[ConfigFinding]
|
list[ConfigFinding]: A warning for each unset variable first, then
the findings of |
Source code in src/veridelta/engine.py
check_configs(diff, source, target, *, schemas=False)
staticmethod
Check a loaded configuration for what would stop a run.
By default nothing connects and no rows are read, so the checks need only the configuration and the installed extras:
- the pair is one an engine can compare: both local, or two tables on one warehouse connection;
- every extra a side reads through, or a local run scores with, is installed;
- a database
tableuses a URI scheme Veridelta can quote for; - each
regex_replacepattern compiles in Polars; - on a warehouse pair, settings the warehouse refuses for some stored names or types, reported as warnings.
With schemas, and no errors so far, each side's columns are read too,
but never its rows. Local sides are checked with validate_rules; a
database table is read with a zero-row probe, and a query is not
run at all. A warehouse pair runs the schema probes a run starts with,
then compiles every comparison statement without executing it, which
settles the warnings above one way or the other. Either way, a name in a
rule's column_names that neither side has is a warning, since that rule
does nothing.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
diff
|
DiffConfig
|
Comparison settings and rules. |
required |
source
|
SourceRef
|
Source configuration. |
required |
target
|
SourceRef
|
Target configuration. |
required |
schemas
|
bool
|
Whether to also connect and check the rules against the stored columns. |
False
|
Returns:
| Type | Description |
|---|---|
list[ConfigFinding]
|
list[ConfigFinding]: Errors and warnings, empty when nothing is wrong. A configuration with no errors is expected to start. |
Source code in src/veridelta/engine.py
propose_value_maps(*, min_confidence=DEFAULT_MIN_CONFIDENCE, min_support=DEFAULT_MIN_SUPPORT, sample_fraction=1.0)
Propose value_map entries from how source and target values line up.
Rows are aligned, normalized, and joined as run() does, on a copy, so this
engine can still run afterward. For each compared text column that stage 4 can
map, a source value is proposed for the target value it lines up with in at
least min_confidence of its joined rows, provided at least min_support rows
agree. Values are read as the value_map stage sees them, so a
case_insensitive column gets lowercase keys. Rows a column's existing map
already translates are left out, so a raw value equal to one of that map's
outputs cannot receive an entry.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
min_confidence
|
float
|
Share of rows that must agree, above 0.5. |
DEFAULT_MIN_CONFIDENCE
|
min_support
|
int
|
Agreeing rows a proposal needs. |
DEFAULT_MIN_SUPPORT
|
sample_fraction
|
float
|
Share of source rows to read, picked by a hash of the primary keys, so the same data samples the same rows under one Polars version. |
1.0
|
Returns:
| Type | Description |
|---|---|
list[ValueMapProposal]
|
list[ValueMapProposal]: One proposal per column with new entries, in source column order. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If a threshold is out of range, or the configuration fails as it would in a run. |
DataIntegrityError
|
If either dataset repeats a normalized primary key. |
Examples:
>>> import polars as pl
>>> from veridelta.models import DiffConfig
>>> source = pl.LazyFrame({"id": range(6), "sex": ["M"] * 6})
>>> target = pl.LazyFrame({"id": range(6), "sex": ["Male"] * 6})
>>> engine = DiffEngine(DiffConfig(primary_keys=["id"]), source, target)
>>> proposals = engine.propose_value_maps()
>>> proposals[0].to_rule().value_map
{'M': 'Male'}
Source code in src/veridelta/engine.py
propose_value_maps_from_configs(diff, source, target, *, min_confidence=DEFAULT_MIN_CONFIDENCE, min_support=DEFAULT_MIN_SUPPORT, sample_fraction=1.0)
classmethod
Propose value_map entries for a SourceRef pair, wherever it lives.
File, lakehouse, and database pairs are loaded and proposed locally. A pair of tables on one warehouse connection is counted in the warehouse instead, in one statement, for columns stored as text on both sides; rows never leave it. A sampled warehouse run reads a different, though equally repeatable, set of keys than a local one.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
diff
|
DiffConfig
|
Comparison settings and rules. |
required |
source
|
SourceRef
|
Source configuration. |
required |
target
|
SourceRef
|
Target configuration. |
required |
min_confidence
|
float
|
Share of rows that must agree, above 0.5. |
DEFAULT_MIN_CONFIDENCE
|
min_support
|
int
|
Agreeing rows a proposal needs. |
DEFAULT_MIN_SUPPORT
|
sample_fraction
|
float
|
Share of source rows to read. |
1.0
|
Returns:
| Type | Description |
|---|---|
list[ValueMapProposal]
|
list[ValueMapProposal]: One proposal per column with new entries. |
Raises:
| Type | Description |
|---|---|
ConnectorError
|
If only one side is a warehouse table, the sides use different warehouses or connections, or a warehouse returns a malformed result. |
ConfigError
|
If a threshold is out of range, or the configuration fails as it would in a run. |
DataIntegrityError
|
If either dataset repeats a normalized primary key. |
Source code in src/veridelta/engine.py
read_schema(config)
staticmethod
Read one side's columns and their types, and return none of its rows.
Each side is read as a run reads it before its first row. A file is
read by its loader, and a CSV file's types come from its first rows; a
lakehouse table gives its schema; a database or DuckDB table is read
with a probe that returns no rows. A warehouse table, or a table with
pushdown, gets the probe a pushdown run starts with, in a session of
its own that is closed after.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
SourceRef
|
One side of a configuration. |
required |
Returns:
| Type | Description |
|---|---|
Schema
|
pl.Schema: The columns in their stored order and with their stored
names, before |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If the side reads a database or DuckDB |
ConnectorError
|
If the side cannot be reached or read. |
Source code in src/veridelta/engine.py
run(*, baseline=None)
Compare the two datasets and return the result.
The comparison stays lazy until it collects the joins, so the inputs can be scans over files larger than memory. A run takes these steps, in order:
- Align the columns: apply renames, drop ignored columns, and treat the target as the authoritative schema.
- Check that the primary keys exist and that
schema_modeholds. - Apply stages 1 to 7 of the
DiffRuletransform order to each side, and build each column's stage 8 and 9 comparison, so a rule the run cannot honor fails before any rows move. - Check that the normalized primary keys are unique on each side.
- Find the added, removed, and changed rows, leave out the drift
baselineaccepts, and count the mismatches. - Write the artifacts, when
output_pathis set.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
baseline
|
Baseline | None
|
Drift to accept: rows by kind and key, and
changed columns by row. The counts, the verdict, the artifacts,
and the reports leave it out, and |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
DiffResult |
DiffResult
|
Counts, column-level drift, and the differing rows. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If a primary key is missing or holds two types a join
cannot pair, |
DataIntegrityError
|
If either dataset repeats a normalized primary key. |
Examples:
>>> import polars as pl
>>> from veridelta.models import DiffConfig
>>> source = pl.LazyFrame({"id": [1, 2, 3], "amount": [10.0, 20.0, 30.0]})
>>> target = pl.LazyFrame({"id": [2, 3, 4], "amount": [20.0, 31.0, 40.0]})
>>> result = DiffEngine(DiffConfig(primary_keys=["id"]), source, target).run()
>>> result.summary.added_count, result.summary.removed_count
(1, 1)
>>> result.summary.column_mismatches
{'amount': 1}
Source code in src/veridelta/engine.py
run_from_configs(diff, source, target, *, baseline=None)
classmethod
Route a comparison to pushdown or local Polars evaluation.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
diff
|
DiffConfig
|
Comparison settings and rules. |
required |
source
|
SourceRef
|
Source file, lakehouse, database, DuckDB, or warehouse config. |
required |
target
|
SourceRef
|
Target file, lakehouse, database, DuckDB, or warehouse config. |
required |
baseline
|
Baseline | None
|
Drift to accept, which a local run leaves out of the counts and the verdict. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
DiffResult |
DiffResult
|
The result. A pushdown pair returns counts and keys, and a local pair also returns the differing rows. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If primary keys are missing, |
DataIntegrityError
|
If either dataset repeats a normalized primary key. |
ConnectorError
|
If warehouse backends are mixed or connections differ. |
Examples:
>>> from veridelta.models import DiffConfig, SourceConfig
>>> diff = DiffConfig(primary_keys=["order_id"])
>>> source = SourceConfig(path="legacy/orders.parquet", format="parquet")
>>> target = SourceConfig(path="modern/orders.parquet", format="parquet")
>>> result = DiffEngine.run_from_configs(diff, source, target)
Source code in src/veridelta/engine.py
suggest_rules(*, max_share=DEFAULT_MAX_SHARE)
Suggest rules that would explain the differences in each compared column.
The comparison runs as run() runs it, on a copy, without writing artifacts,
so this engine can still run afterward. A numeric column gets a tolerance when
every gap between its differing values is at most max_share of the larger
of the two: an absolute tolerance when the gaps stay about one size, and a
relative one when they grow with the values. Each tolerance is the round value
just above the largest gap, and the comparison runs again with it to count
the rows it explains. No model is called.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
max_share
|
float
|
Largest gap a tolerance may explain, as a share of the larger of its two values, above 0 and at most 1. |
DEFAULT_MAX_SHARE
|
Returns:
| Type | Description |
|---|---|
list[RuleSuggestion]
|
list[RuleSuggestion]: One suggestion per column it can explain, in compared column order. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If |
DataIntegrityError
|
If either dataset repeats a normalized primary key. |
Examples:
>>> import polars as pl
>>> from veridelta.models import DiffConfig
>>> source = pl.LazyFrame({"id": [1, 2, 3], "fare": [10.0, 20.0, 30.0]})
>>> target = pl.LazyFrame({"id": [1, 2, 3], "fare": [10.004, 20.004, 30.003]})
>>> engine = DiffEngine(DiffConfig(primary_keys=["id"]), source, target)
>>> [(s.column, s.settings) for s in engine.suggest_rules()]
[('fare', {'absolute_tolerance': 0.005})]
Source code in src/veridelta/engine.py
632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 | |
suggest_rules_from_configs(diff, source, target, *, max_share=DEFAULT_MAX_SHARE)
classmethod
Suggest rules for a SourceRef pair, read locally.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
diff
|
DiffConfig
|
Comparison settings and rules. |
required |
source
|
SourceRef
|
Source configuration. |
required |
target
|
SourceRef
|
Target configuration. |
required |
max_share
|
float
|
Largest gap a tolerance may explain, as a share of the larger of its two values, above 0 and at most 1. |
DEFAULT_MAX_SHARE
|
Returns:
| Type | Description |
|---|---|
list[RuleSuggestion]
|
list[RuleSuggestion]: One suggestion per column it can explain. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If |
DataIntegrityError
|
If either dataset repeats a normalized primary key. |
Source code in src/veridelta/engine.py
validate_rules(config, source_df, target_df)
classmethod
Check everything a run checks before it reads a row.
Goes past validate_schemas: every rule is resolved against the aligned
columns, both schemas are normalized, and each column's comparison is built. A
rule the run could not honor fails here, such as a null sentinel its column's
type cannot hold, or a similarity limit without the fuzzy extra. Operates on
schema metadata only, so callers may pass zero-row frames. Repeated keys and
invalid regular expressions surface only when rows are read.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
DiffConfig
|
Comparison settings and rules. |
required |
source_df
|
LazyFrame
|
Source frame or column probe. |
required |
target_df
|
LazyFrame
|
Target frame or column probe. |
required |
Returns:
| Type | Description |
|---|---|
list[str]
|
list[str]: The columns a run would compare, in source order, under their target names. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If primary keys are missing or hold two types a join cannot pair, schema constraints are violated, or a rule cannot apply as configured. |
Source code in src/veridelta/engine.py
validate_schemas(config, source_df, target_df)
classmethod
Align structure and enforce SchemaMode without comparing any rows.
Operates on schema metadata only, so callers may pass zero-row frames.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
DiffConfig
|
Comparison settings and rules. |
required |
source_df
|
LazyFrame
|
Source frame or column probe. |
required |
target_df
|
LazyFrame
|
Target frame or column probe. |
required |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If primary keys are missing or schema constraints are violated. |
Source code in src/veridelta/engine.py
EffectiveRule
Bases: TypedDict
Per-column settings after specific, pattern, and global rules are merged.
Source code in src/veridelta/_resolution.py
ExcelLoader
Bases: BaseLoader
Loader for Excel workbooks, backed by the optional excel extra.
Like JSON, this is eager: a spreadsheet is a random-access container with no streaming reader.
Source code in src/veridelta/_reading.py
load(config)
Read one worksheet.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
SourceConfig
|
Source configuration, whose options go to
|
required |
Returns:
| Type | Description |
|---|---|
LazyFrame
|
pl.LazyFrame: A lazy wrapper over the rows read. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If the |
Source code in src/veridelta/_reading.py
JSONLoader
Bases: BaseLoader
Loader for a JSON file holding one array of records.
Polars has no lazy JSON reader: a JSON array cannot be parsed incrementally the
way newline-delimited records can. The file is read whole and wrapped. Prefer
ndjson for large files.
Source code in src/veridelta/_reading.py
load(config)
Read a JSON file.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
SourceConfig
|
Source configuration, whose options go to |
required |
Returns:
| Type | Description |
|---|---|
LazyFrame
|
pl.LazyFrame: A lazy wrapper over the rows read. |
Source code in src/veridelta/_reading.py
LoaderFactory
Resolve a file, lakehouse, database, or DuckDB SourceRef to a LazyFrame.
A file source goes to the loader for its format, and its schema is read
before it returns, so a file that cannot be read fails here, by name. A
Delta Lake or Iceberg source returns its connector's lazy scan, and a database or DuckDB source is
read once through its connector, which then closes. A warehouse source is
refused: its comparison runs as SQL pushdown through
DiffEngine.run_from_configs.
Attributes:
| Name | Type | Description |
|---|---|---|
_loaders |
ClassVar[dict[str, BaseLoader]]
|
Format name to loader. It is
the one list of formats, and |
Source code in src/veridelta/_reading.py
257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 | |
get_loader(source_type)
classmethod
Return the loader for a file format.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source_type
|
str
|
Format name, such as |
required |
Returns:
| Name | Type | Description |
|---|---|---|
BaseLoader |
BaseLoader
|
The loader. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If the format has no loader. |
Source code in src/veridelta/_reading.py
load(config)
classmethod
Load a file, lakehouse, database, or DuckDB source into a LazyFrame.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
SourceRef
|
File, Delta, Iceberg, database, or DuckDB configuration. |
required |
Returns:
| Type | Description |
|---|---|
LazyFrame
|
pl.LazyFrame: Unevaluated scan graph, or a lazy wrapper over the rows a database or DuckDB source read. |
Raises:
| Type | Description |
|---|---|
ConnectorError
|
If |
ConfigError
|
If the file format has no loader, or a database
|
Source code in src/veridelta/_reading.py
NDJSONLoader
Bases: BaseLoader
Streaming loader for newline-delimited JSON over pl.scan_ndjson.
One record per line is the JSON shape Polars can read incrementally, so
this is the format to prefer over json for large exports.
Source code in src/veridelta/_reading.py
load(config)
Scan a newline-delimited JSON file.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
SourceConfig
|
Source configuration, whose options go to |
required |
Returns:
| Type | Description |
|---|---|
LazyFrame
|
pl.LazyFrame: The unevaluated rows. |
Source code in src/veridelta/_reading.py
ParquetLoader
Bases: BaseLoader
Streaming Parquet loader over pl.scan_parquet.
Accepts whatever path or glob the Polars scanner accepts. Because the scan
stays lazy, columns dropped by ignore rules are never read from disk.
Source code in src/veridelta/_reading.py
load(config)
Scan a Parquet file.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
SourceConfig
|
Source configuration, whose options go to |
required |
Returns:
| Type | Description |
|---|---|
LazyFrame
|
pl.LazyFrame: The unevaluated rows. |
Source code in src/veridelta/_reading.py
Configuration loading
Functions that read and check a YAML file, or build a configuration from two files, and return configuration models. The models themselves live in veridelta.models, which Configuration models documents.
Configuration files: loading, checking, and their JSON Schema.
load_config reads a YAML file into the models the engine runs on, and
config_json_schema describes the same files for editors and validators.
SCHEMA_URL = 'https://veridelta.github.io/veridelta/schema/veridelta.schema.json'
module-attribute
Where the docs site publishes the configuration schema for editors to fetch.
config_json_schema()
Return a JSON Schema for configuration files, for editors and validators.
It is generated from the same models load_config validates with, then
adjusted where the loader does something before validating:
typeis required in every warehouse, lakehouse, and database block, because the loader reads a block without one as a file source.- Patterned and enumerated strings inside
sourceandtarget, such astable, also accept a${NAME}reference, which the loader expands.
The schema is stricter than the loader in one way: it does not model
Pydantic's lax coercion, so a quoted number such as threshold: "0.1" is
flagged even though it loads.
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
dict[str, Any]: A Draft 2020-12 JSON Schema. |
Source code in src/veridelta/config.py
files_config(source, target, primary_keys)
Build the configuration that compares two files on their primary keys.
It holds what a file with only primary_keys and a path for each side holds,
and is checked the same way, so each file's format follows its suffix. The
paths are read as given: a shell has already expanded its own variables, so a
${NAME} in one is left as it is.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source
|
str | Path
|
The source file. |
required |
target
|
str | Path
|
The target file. |
required |
primary_keys
|
Sequence[str]
|
The columns that identify a row on both sides. |
required |
Returns:
| Type | Description |
|---|---|
tuple[DiffConfig, SourceConfig, SourceConfig]
|
tuple[DiffConfig, SourceConfig, SourceConfig]: The comparison settings, then
the source and the target, as |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If |
Examples:
>>> diff, source, target = files_config("legacy.csv", "modern.parquet", ["id"])
>>> diff.primary_keys, source.path, target.format
(['id'], 'legacy.csv', 'parquet')
Source code in src/veridelta/config.py
load_config(path, *, unset_env=None)
Load and validate a configuration file.
The source and target blocks become source configurations, and every other
root key belongs to the DiffConfig. A file source may omit type, which
defaults to file.
Strings inside source and target may reference environment variables as
${NAME}, or ${NAME:-default} to fall back when the variable is unset or
empty, so credentials can stay out of the file. $${ writes a literal ${.
Root settings and rules are read verbatim, which keeps a ${1} in a regex
replacement intact.
Passing a list as unset_env checks a file without its secrets: an unset
variable with no default then reads as its own name, so ${TABLE} becomes
TABLE, and its name is appended to the list once. A block that fails
validation after such a guess says which variables it guessed.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
str | Path
|
Path to the YAML file. |
required |
unset_env
|
list[str] | None
|
Collects unset variables instead of raising for them. None, the default, raises. |
None
|
Returns:
| Type | Description |
|---|---|
tuple[DiffConfig, SourceRef, SourceRef]
|
tuple[DiffConfig, SourceRef, SourceRef]: The comparison settings, then the source and the target. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If the file is missing or not valid YAML, lacks a |
Examples:
>>> from pathlib import Path
>>> from tempfile import TemporaryDirectory
>>> text = "{primary_keys: [id], source: {path: a.csv}, target: {path: b.csv}}"
>>> with TemporaryDirectory() as folder:
... path = Path(folder, "veridelta.yaml")
... _ = path.write_text(text)
... diff, source, target = load_config(path)
>>> diff.primary_keys, source.path
(['id'], 'a.csv')
Source code in src/veridelta/config.py
referenced_variables(text)
Name each environment variable a configuration's text references, once, in order.
The text is read as written, so a file that does not load names its
variables too, and a reference outside source and target, which the
loader leaves as it is, counts as well.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
The text of a configuration file. |
required |
Returns:
| Type | Description |
|---|---|
list[str]
|
list[str]: The name in each |
Examples:
>>> referenced_variables("{password: ${PASSWORD}, role: '${ROLE:-ANALYST}$${X}'}")
['PASSWORD', 'ROLE']
Source code in src/veridelta/config.py
Exceptions
The errors Veridelta raises. Each derives from VerideltaError, so one except clause catches them all without hiding Python's own errors.
The errors Veridelta raises.
Each derives from VerideltaError, so one except clause catches them all.
ConfigError
Bases: VerideltaError
Raised when a configuration is invalid or the data breaks its rules.
Such failures include a primary key missing from a dataset, a rule the
column types cannot satisfy, and a column that schema_mode forbids.
Source code in src/veridelta/exceptions.py
ConnectorError
Bases: VerideltaError
Raised when a source connector cannot complete an operation.
Such failures include a missing optional extra, a failed query or scan, a pair of backends that cannot be compared, and a call before a session opens.
Source code in src/veridelta/exceptions.py
DataIntegrityError
Bases: VerideltaError
Raised when the data breaks an assumption the comparison relies on.
Repeated primary keys in either dataset raise it before any join runs, since a repeated key multiplies the joined rows.
DatasetError
Bases: VerideltaError
Raised when a sample dataset cannot be downloaded.
veridelta.datasets fetches tutorial data over the network, and a failed
or interrupted download raises this error.
VerideltaError
Bases: Exception
Base class for every error Veridelta raises.
Catch it to handle any Veridelta failure without also catching Python's own
errors, such as MemoryError or ValueError.
missing_extra(extra, needed_for)
Return the one message for an optional extra that is not installed.
The error that carries it stays the one its caller raises: a reader's
ConfigError or a connector's ConnectorError.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
extra
|
str
|
The extra, such as |
required |
needed_for
|
str
|
What needs it, as the subject of a sentence, such as
|
required |
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
What needs the extra, and the command that installs it. |
Source code in src/veridelta/exceptions.py
Connectors
Warehouse sessions, lakehouse scanners, and the database and DuckDB readers. VerideltaConnector is the lifecycle they share, ReaderConnector and PushdownSession are the two kinds the engine drives, and SQLPushdownCompiler writes each dialect's comparison SQL.
Warehouse pushdown, lakehouse-native, database, and DuckDB connector abstractions.
PushdownQueryType = Literal['mismatch', 'added', 'missing', 'count', 'duplicates', 'columns', 'schema', 'value_maps', 'settings', 'samples']
module-attribute
Warehouse pushdown round-trip: comparison rows, tallies, totals, key checks, probes, value map evidence, a check of the server's settings, or a sample of changed rows with their values.
BigQueryConnector
Bases: PushdownSession
BigQuery warehouse connector backed by the optional BigQuery extra.
connect() builds a bigquery.Client for the configured project, from a
service account key file when credentials_path is set and from
Application Default Credentials otherwise. Every statement runs as a
GoogleSQL job under one QueryJobConfig, which names the default dataset
and the maximum_bytes_billed cap. Install the client with
uv add 'veridelta[bigquery]'; it is imported on first connect.
Attributes:
| Name | Type | Description |
|---|---|---|
compiler |
SQLPushdownCompiler
|
BigQuery-dialect compiler (backtick quoting, GoogleSQL type names) the engine uses for every statement. |
Source code in src/veridelta/connectors/warehouse.py
358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 | |
__init__(config)
Initialize the connector with validated BigQuery settings.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
BigQueryConfig
|
Frozen project, table, and job settings. |
required |
Source code in src/veridelta/connectors/warehouse.py
close()
Close the BigQuery client, if one is open.
Idempotent. Afterwards execute_pushdown raises ConnectorError until
connect() is called again.
Source code in src/veridelta/connectors/warehouse.py
connect()
Create the BigQuery client and the job settings every statement uses.
Raises:
| Type | Description |
|---|---|
ConnectorError
|
If the BigQuery extra is missing or the client cannot be created, such as when no credentials are found. |
Source code in src/veridelta/connectors/warehouse.py
execute_pushdown(statement, query_type='mismatch')
Execute compiler SQL as a BigQuery job and return a LazyFrame.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
statement
|
str
|
SQL produced by |
required |
query_type
|
PushdownQueryType
|
Which comparison round-trip this statement represents; recorded in the log line for the call. |
'mismatch'
|
Returns:
| Type | Description |
|---|---|
LazyFrame
|
pl.LazyFrame: Unevaluated frame wrapped around the Arrow result. |
Raises:
| Type | Description |
|---|---|
ConnectorError
|
If the connector is not connected, the job fails, or it returns no Arrow batches. |
Source code in src/veridelta/connectors/warehouse.py
DatabaseConnector
Bases: ReaderConnector
Read one database table or query into Polars through ConnectorX.
connect() reads eagerly and keeps the frame, and lazyframe() hands it to
the local engine as a LazyFrame over those rows. The read is the one place
the rows are fetched, so it happens once per connect(). With
partition_on set, ConnectorX splits that read into ranges over parallel
connections, after Veridelta confirms the column holds no NULL and reads
its lowest and highest values. The comparison runs in Polars, never in the
database.
Source code in src/veridelta/connectors/database.py
74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 | |
__init__(config, *, probe=False)
Initialize the connector with validated database settings.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
DatabaseConfig
|
Frozen URI, credentials, and table or query. |
required |
probe
|
bool
|
Whether to read the table's columns and no rows, for a
schema check. Only a |
False
|
Source code in src/veridelta/connectors/database.py
connect()
Read the configured table or query into memory.
Raises:
| Type | Description |
|---|---|
ConnectorError
|
If the |
ConfigError
|
If |
Source code in src/veridelta/connectors/database.py
DatabricksConnector
Bases: _CursorSession
Databricks SQL warehouse connector backed by the optional Databricks extra.
connect() opens a databricks.sql session against the configured SQL
warehouse HTTP path; execute_pushdown runs compiler SQL on a fresh cursor
and fetches the result as Arrow. Install the driver with
uv add 'veridelta[databricks]'; without it, connect() raises
ConnectorError with that hint instead of an ImportError.
Attributes:
| Name | Type | Description |
|---|---|---|
compiler |
SQLPushdownCompiler
|
Databricks-dialect compiler (backtick quoting, Spark type names) the engine uses for every statement. |
Source code in src/veridelta/connectors/warehouse.py
__init__(config)
Initialize the connector with validated Databricks settings.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
DatabricksConfig
|
Frozen workspace hostname and HTTP path. |
required |
Source code in src/veridelta/connectors/warehouse.py
connect()
Open a Databricks SQL session for subsequent pushdown statements.
Raises:
| Type | Description |
|---|---|
ConnectorError
|
If the Databricks extra is missing or authentication fails. |
Source code in src/veridelta/connectors/warehouse.py
DeltaLakeConnector
Bases: ReaderConnector
Delta Lake scanner backed by pl.scan_delta.
connect() opens a lazy scan of DeltaLakeConfig.table_uri, pinned to
version when one is set, and lazyframe() hands that scan to the local
engine. No SQL is involved: the comparison runs in Polars. Requires the
delta extra (uv add 'veridelta[delta]').
Source code in src/veridelta/connectors/lakehouse.py
__init__(config)
Initialize the connector with validated Delta Lake settings.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
DeltaLakeConfig
|
Frozen table URI and optional version. |
required |
connect()
Open a lazy pl.scan_delta of the configured table and read its log.
Raises:
| Type | Description |
|---|---|
ConnectorError
|
If the |
Source code in src/veridelta/connectors/lakehouse.py
DuckDBConnector
Bases: ReaderConnector
Read one DuckDB or MotherDuck table or query into Polars.
connect() reads eagerly and keeps the frame, and lazyframe() hands it to
the local engine as a LazyFrame over those rows. The read is the one place
the rows are fetched, so it happens once per connect(). The comparison
runs in Polars, never in DuckDB.
Source code in src/veridelta/connectors/duckdb.py
107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 | |
__init__(config, *, probe=False)
Initialize the connector with validated DuckDB settings.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
DuckDBConfig
|
Frozen database, token, and table or query. |
required |
probe
|
bool
|
Whether to read the table's columns and no rows, for a
schema check. Only a |
False
|
Source code in src/veridelta/connectors/duckdb.py
connect()
Read the configured table or query into memory.
Raises:
| Type | Description |
|---|---|
ConnectorError
|
If the |
ConfigError
|
If a probe was asked of a |
Source code in src/veridelta/connectors/duckdb.py
DuckDBPushdownSession
Bases: PushdownSession
Run compiled comparison SQL inside DuckDB or MotherDuck.
Opened for two tables in one database that both set pushdown. It holds
one connection from connect() to close(), opened as a read opens it:
a file read-only, MotherDuck read-write with its token, both in UTC. Only
counts and keys come back, and a row sample when one is asked for.
Source code in src/veridelta/connectors/duckdb.py
189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 | |
__init__(config)
Initialize the session for one side's connection settings.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
DuckDBConfig
|
A DuckDB table that sets |
required |
Source code in src/veridelta/connectors/duckdb.py
close()
Close the connection. Idempotent; connect() opens it again.
Source code in src/veridelta/connectors/duckdb.py
connect()
Open the database for the statements to come.
Raises:
| Type | Description |
|---|---|
ConnectorError
|
If the |
Source code in src/veridelta/connectors/duckdb.py
execute_pushdown(statement, query_type='mismatch')
Run one compiled statement and return its rows lazily.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
statement
|
str
|
SQL from the DuckDB compiler. |
required |
query_type
|
PushdownQueryType
|
Which round-trip this is, for logs and errors. |
'mismatch'
|
Returns:
| Type | Description |
|---|---|
LazyFrame
|
pl.LazyFrame: The statement's result. |
Raises:
| Type | Description |
|---|---|
ConnectorError
|
If the session is not connected, a result column has no Polars type, or the statement fails. |
Source code in src/veridelta/connectors/duckdb.py
IcebergConnector
Bases: ReaderConnector
Apache Iceberg scanner backed by pl.scan_iceberg.
connect() opens a lazy scan of IcebergConfig.table_uri, pinned to
snapshot_id when one is set, and lazyframe() hands that scan to the
local engine. As with Delta Lake, the comparison runs in Polars. Requires
the iceberg extra (uv add 'veridelta[iceberg]').
Source code in src/veridelta/connectors/lakehouse.py
__init__(config)
Initialize the connector with validated Iceberg settings.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
IcebergConfig
|
Frozen table URI and storage options. |
required |
connect()
Open a lazy pl.scan_iceberg of the configured table and read its metadata.
Raises:
| Type | Description |
|---|---|
ConnectorError
|
If the |
Source code in src/veridelta/connectors/lakehouse.py
PostgresPushdownSession
Bases: PushdownSession
Run compiled comparison SQL inside Postgres, through ConnectorX.
Opened for two database sources on one Postgres connection that both set
pushdown. Each statement is one ConnectorX read, which opens its own
connection and returns Arrow, so nothing is held between statements and
close() only forgets the connection. Only counts and keys come back.
connect() checks that the server reads string literals by the SQL
standard, as the Postgres dialect writes them: with
standard_conforming_strings off, a backslash would escape the closing
quote of a value that ends in one.
Source code in src/veridelta/connectors/database.py
244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 | |
__init__(config)
Initialize the session for one side's connection settings.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
DatabaseConfig
|
A Postgres table that sets |
required |
Source code in src/veridelta/connectors/database.py
close()
connect()
Check the server and keep the connection URI for the statements to come.
Raises:
| Type | Description |
|---|---|
ConnectorError
|
If the |
Source code in src/veridelta/connectors/database.py
declared_types(table)
Return the declared precision and scale of a table's numeric columns.
ConnectorX describes every numeric as Decimal(38, 10), so a schema
probe alone cannot tell numeric(10, 2) from numeric(12, 4).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
table
|
str
|
The table, as configured. |
required |
Returns:
| Type | Description |
|---|---|
dict[str, Decimal]
|
dict[str, pl.Decimal]: Each |
Raises:
| Type | Description |
|---|---|
ConnectorError
|
If the session is not connected or the query fails. |
Source code in src/veridelta/connectors/database.py
execute_pushdown(statement, query_type='mismatch')
Run one compiled statement and return its rows lazily.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
statement
|
str
|
SQL from the Postgres compiler. |
required |
query_type
|
PushdownQueryType
|
Which round-trip this is, for logs and errors. |
'mismatch'
|
Returns:
| Type | Description |
|---|---|
LazyFrame
|
pl.LazyFrame: The statement's result. |
Raises:
| Type | Description |
|---|---|
ConnectorError
|
If the session is not connected or the statement fails. |
Source code in src/veridelta/connectors/database.py
PushdownSession
Bases: VerideltaConnector
A connector that runs compiled comparison SQL where the data lives.
The engine compiles every statement with compiler and runs it through
execute_pushdown, one round-trip per PushdownQueryType. Only what a
statement returns comes back: counts, keys, and the samples asked for,
never the tables themselves.
Attributes:
| Name | Type | Description |
|---|---|---|
compiler |
SQLPushdownCompiler
|
The dialect's compiler, which the
subclass sets before |
Source code in src/veridelta/connectors/base.py
execute_pushdown(statement, query_type='mismatch')
abstractmethod
Execute compiled SQL and return an unevaluated result graph.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
statement
|
str
|
Compiler-produced SQL. Never user text. |
required |
query_type
|
PushdownQueryType
|
Which round-trip the SQL represents
( |
'mismatch'
|
Returns:
| Type | Description |
|---|---|
LazyFrame
|
pl.LazyFrame: Unevaluated result graph. Must not be collected here. |
Raises:
| Type | Description |
|---|---|
ConnectorError
|
If the session is not connected or the statement fails. |
Source code in src/veridelta/connectors/base.py
ReaderConnector
Bases: VerideltaConnector
A connector the local engine reads through lazyframe().
connect() opens a scan or reads the rows and keeps them as _frame,
lazyframe() hands them to the engine, and the comparison runs in Polars.
close() drops them. A reader has no execute_pushdown: nothing is pushed
into a source that is compared locally.
Source code in src/veridelta/connectors/base.py
close()
lazyframe()
Return what connect() opened, as an unevaluated LazyFrame.
Returns:
| Type | Description |
|---|---|
LazyFrame
|
pl.LazyFrame: A lazy scan, or a lazy wrapper over the rows read. |
Raises:
| Type | Description |
|---|---|
ConnectorError
|
If |
Source code in src/veridelta/connectors/base.py
SQLDialect
Bases: StrEnum
Warehouse SQL dialects supported by the pushdown compiler.
POSTGRES runs where two database sources on one Postgres connection opt
into pushdown, and DUCKDB where two DuckDB tables in one database do.
The differential test harness also runs DUCKDB output, so the compiled
SQL itself, not a rewrite of it, is checked against the local engine.
Source code in src/veridelta/connectors/sql.py
SQLPushdownCompiler
Compile DiffRule semantics into dialect-specific SQL strings.
One instance targets one SQLDialect. The engine drives a warehouse run
with these statements, in this order:
compile_schema_probe_queryper side, to learn column names and types.compile_duplicate_key_queryper side, to assert normalized keys are unique before anything joins on them.compile_count_queryper side, for thethresholddenominator.compile_queryfor inner-join rows where a compared column differs.compile_added_queryandcompile_missing_queryfor the anti-joins.compile_column_mismatch_queryfor the per-column tally.
A value map proposal runs steps 1 and 2, then compile_value_map_query.
Every join reads keys through the same stages 1-7 as compared columns,
driven by key_rules, so rows match on the keys the local engine sees.
Rules reach the compiler already folded over the configuration's
default_* settings, so every compared column arrives as one fully
specified DiffRule.
The compiler is also the security boundary for warehouse SQL. Identifiers
are allowlisted segment by segment and then dialect-quoted, data literals
are escaped through _literal, and every dialect keyword comes from a
module-level table keyed by SQLDialect, so an unsupported combination
raises rather than borrowing another dialect's spelling.
Attributes:
| Name | Type | Description |
|---|---|---|
dialect |
SQLDialect
|
Target dialect for quoting, casts, and functions. |
Source code in src/veridelta/connectors/sql.py
699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 945 946 947 948 949 950 951 952 953 954 955 956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 973 974 975 976 977 978 979 980 981 982 983 984 985 986 987 988 989 990 991 992 993 994 995 996 997 998 999 1000 1001 1002 1003 1004 1005 1006 1007 1008 1009 1010 1011 1012 1013 1014 1015 1016 1017 1018 1019 1020 1021 1022 1023 1024 1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 1052 1053 1054 1055 1056 1057 1058 1059 1060 1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 1077 1078 1079 1080 1081 1082 1083 1084 1085 1086 1087 1088 1089 1090 1091 1092 1093 1094 1095 1096 1097 1098 1099 1100 1101 1102 1103 1104 1105 1106 1107 1108 1109 1110 1111 1112 1113 1114 1115 1116 1117 1118 1119 1120 1121 1122 1123 1124 1125 1126 1127 1128 1129 1130 1131 1132 1133 1134 1135 1136 1137 1138 1139 1140 1141 1142 1143 1144 1145 1146 1147 1148 1149 1150 1151 1152 1153 1154 1155 1156 1157 1158 1159 1160 1161 1162 1163 1164 1165 1166 1167 1168 1169 1170 1171 1172 1173 1174 1175 1176 1177 1178 1179 1180 1181 1182 1183 1184 1185 1186 1187 1188 1189 1190 1191 1192 1193 1194 1195 1196 1197 1198 1199 1200 1201 1202 1203 1204 1205 1206 1207 1208 1209 1210 1211 1212 1213 1214 1215 1216 1217 1218 1219 1220 1221 1222 1223 1224 1225 1226 1227 1228 1229 1230 1231 1232 1233 1234 1235 1236 1237 1238 1239 1240 1241 1242 1243 1244 1245 1246 1247 1248 1249 1250 1251 1252 1253 1254 1255 1256 1257 1258 1259 1260 1261 1262 1263 1264 1265 1266 1267 1268 1269 1270 1271 1272 1273 1274 1275 1276 1277 1278 1279 1280 1281 1282 1283 1284 1285 1286 1287 1288 1289 1290 1291 1292 1293 1294 1295 1296 1297 1298 1299 1300 1301 1302 1303 1304 1305 1306 1307 1308 1309 1310 1311 1312 1313 1314 1315 1316 1317 1318 1319 1320 1321 1322 1323 1324 1325 1326 1327 1328 1329 1330 1331 1332 1333 1334 1335 1336 1337 1338 1339 1340 1341 1342 1343 1344 1345 1346 1347 1348 1349 1350 1351 1352 1353 1354 1355 1356 1357 1358 1359 1360 1361 1362 1363 1364 1365 1366 1367 1368 1369 1370 1371 1372 1373 1374 1375 1376 1377 1378 1379 1380 1381 1382 1383 1384 1385 1386 1387 1388 1389 1390 1391 1392 1393 1394 1395 1396 1397 1398 1399 1400 1401 1402 1403 1404 1405 1406 1407 1408 1409 1410 1411 1412 1413 1414 1415 1416 1417 1418 1419 1420 1421 1422 1423 1424 1425 1426 1427 1428 1429 1430 1431 1432 1433 1434 1435 1436 1437 1438 1439 1440 1441 1442 1443 1444 1445 1446 1447 1448 1449 1450 1451 1452 1453 1454 1455 1456 1457 1458 1459 1460 1461 1462 1463 1464 1465 1466 1467 1468 1469 1470 1471 1472 1473 1474 1475 1476 1477 1478 1479 1480 1481 1482 1483 1484 1485 1486 1487 1488 1489 1490 1491 1492 1493 1494 1495 1496 1497 1498 1499 1500 1501 1502 1503 1504 1505 1506 1507 1508 1509 1510 1511 1512 1513 1514 1515 1516 1517 1518 1519 1520 1521 1522 1523 1524 1525 1526 1527 1528 1529 1530 1531 1532 1533 1534 1535 1536 1537 1538 1539 1540 1541 1542 1543 1544 1545 1546 1547 1548 1549 1550 1551 1552 1553 1554 1555 1556 1557 1558 1559 1560 1561 1562 1563 1564 1565 1566 1567 1568 1569 1570 1571 1572 1573 1574 1575 1576 1577 1578 1579 1580 1581 1582 1583 1584 1585 1586 1587 1588 1589 1590 1591 1592 1593 1594 1595 1596 1597 1598 1599 1600 1601 1602 1603 1604 1605 1606 1607 1608 1609 1610 1611 1612 1613 1614 1615 1616 1617 1618 1619 1620 1621 1622 1623 1624 1625 1626 1627 1628 1629 1630 1631 1632 1633 1634 1635 1636 1637 1638 1639 1640 1641 1642 1643 1644 1645 1646 1647 1648 1649 1650 1651 1652 1653 1654 1655 1656 1657 1658 1659 1660 1661 1662 1663 1664 1665 1666 1667 1668 1669 1670 1671 1672 1673 1674 1675 1676 1677 1678 1679 1680 1681 1682 1683 1684 1685 1686 1687 1688 1689 1690 1691 1692 1693 1694 1695 1696 1697 1698 1699 1700 1701 1702 1703 1704 1705 1706 1707 1708 1709 1710 1711 1712 1713 1714 1715 1716 1717 1718 1719 1720 1721 1722 1723 1724 1725 1726 1727 1728 1729 1730 1731 1732 1733 1734 1735 1736 1737 1738 1739 1740 1741 1742 1743 1744 1745 1746 1747 1748 1749 1750 1751 1752 1753 1754 1755 1756 1757 1758 1759 1760 1761 1762 1763 1764 1765 1766 1767 1768 1769 1770 1771 1772 1773 1774 1775 1776 1777 1778 1779 1780 1781 1782 1783 1784 1785 1786 1787 1788 1789 1790 1791 1792 1793 1794 1795 1796 1797 1798 1799 1800 1801 1802 1803 1804 1805 1806 1807 1808 1809 1810 1811 1812 1813 1814 1815 1816 1817 1818 1819 1820 1821 1822 1823 1824 1825 1826 1827 1828 1829 1830 1831 1832 1833 1834 1835 1836 1837 1838 1839 1840 1841 1842 1843 1844 1845 1846 1847 1848 1849 1850 1851 1852 1853 1854 1855 1856 1857 1858 1859 1860 1861 1862 1863 1864 1865 1866 1867 1868 1869 1870 1871 1872 1873 1874 1875 1876 1877 1878 1879 1880 1881 1882 1883 1884 1885 1886 1887 1888 | |
__init__(dialect)
Initialize a compiler for a single warehouse dialect.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dialect
|
SQLDialect
|
Target SQL dialect. |
required |
compile_added_query(source_table, target_table, primary_keys, *, source_types=None, target_types=None, key_rules=None)
Assemble a RIGHT JOIN anti-join for rows present only in the target.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source_table
|
str
|
Source relation (optionally dotted catalog path). |
required |
target_table
|
str
|
Target relation (optionally dotted catalog path). |
required |
primary_keys
|
list[str]
|
Join keys, spelled as the target stores them. |
required |
source_types
|
ColumnTypes | None
|
Probed source dtypes. |
None
|
target_types
|
ColumnTypes | None
|
Probed target dtypes. |
None
|
key_rules
|
Sequence[DiffRule] | None
|
Key normalization, as for
|
None
|
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Target keys whose normalized value has no source counterpart. |
Raises:
| Type | Description |
|---|---|
ConnectorError
|
If tables or keys are empty, or a key rule does not name exactly one primary key. |
Source code in src/veridelta/connectors/sql.py
compile_changed_sample_query(source_table, target_table, primary_keys, rules, *, limit, source_types=None, target_types=None, key_rules=None, wide_integers=frozenset(), type_drift=frozenset())
Assemble a query for the first changed rows, with both sides' values.
This is compile_query with values: the same normalized CTEs, join,
and WHERE clause, so every sampled row is one compile_query reports
as changed, and each match flag is the predicate that decided it. Each
output column takes a positional alias, so a long column name cannot
pass an identifier limit and a key cannot clash with a suffixed column;
SampleQuery.renames gives the local engine's names back. Rows come in
key order, so the same tables give the same sample.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source_table
|
str
|
Source relation (optionally dotted catalog path). |
required |
target_table
|
str
|
Target relation (optionally dotted catalog path). |
required |
primary_keys
|
list[str]
|
Join keys, spelled as the target stores them. |
required |
rules
|
list[DiffRule]
|
Per-column semantic overrides. |
required |
limit
|
int
|
Most rows to return, at least 1. |
required |
source_types
|
ColumnTypes | None
|
Probed source dtypes, as for
|
None
|
target_types
|
ColumnTypes | None
|
Probed target dtypes. |
None
|
key_rules
|
Sequence[DiffRule] | None
|
Key normalization, as for
|
None
|
wide_integers
|
frozenset[str]
|
Integer columns to measure in a wide
type, as for |
frozenset()
|
type_drift
|
frozenset[str]
|
Columns |
frozenset()
|
Returns:
| Type | Description |
|---|---|
SampleQuery | None
|
SampleQuery | None: The statement and its alias names, or None when |
SampleQuery | None
|
no column is compared, since then no row can have changed. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If a rule sets |
ConnectorError
|
If |
Source code in src/veridelta/connectors/sql.py
876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 945 946 947 948 949 950 951 952 953 954 955 956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 973 974 | |
compile_column_mismatch_query(source_table, target_table, primary_keys, rules, *, source_types=None, target_types=None, key_rules=None, wide_integers=frozenset(), type_drift=frozenset())
Assemble a per-column mismatch tally over the joined rows.
Each column contributes one SUM(CASE ...) term, so a single round trip
fills DiffSummary.column_mismatches the way the local engine does.
COALESCE(pred, FALSE) is load-bearing: under three-valued logic a NULL
predicate is neither true nor false, and without the coalesce those rows
would silently count as matches instead of mismatches. The local engine
resolves the same case with val_match.fill_null(False).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source_table
|
str
|
Source relation (optionally dotted catalog path). |
required |
target_table
|
str
|
Target relation (optionally dotted catalog path). |
required |
primary_keys
|
list[str]
|
Join keys present on both relations. |
required |
rules
|
list[DiffRule]
|
Per-column semantic overrides. |
required |
source_types
|
ColumnTypes | None
|
Probed source dtypes, used to drop null sentinels the column cannot hold. When None, sentinels are emitted unfiltered. |
None
|
target_types
|
ColumnTypes | None
|
Probed target dtypes. |
None
|
key_rules
|
Sequence[DiffRule] | None
|
Key normalization, as for
|
None
|
wide_integers
|
frozenset[str]
|
Integer columns to measure in a wide
type, as for |
frozenset()
|
type_drift
|
frozenset[str]
|
Columns |
frozenset()
|
Returns:
| Type | Description |
|---|---|
str | None
|
str | None: Single-row aggregate statement, or None when no rule |
str | None
|
yields a comparable column, mirroring the local engine's decision to |
str | None
|
skip the tally when there are no match expressions. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If a rule sets |
ConnectorError
|
If tables or keys are empty, a rule is pattern-only,
|
Source code in src/veridelta/connectors/sql.py
1052 1053 1054 1055 1056 1057 1058 1059 1060 1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 1077 1078 1079 1080 1081 1082 1083 1084 1085 1086 1087 1088 1089 1090 1091 1092 1093 1094 1095 1096 1097 1098 1099 1100 1101 1102 1103 1104 1105 1106 1107 1108 1109 1110 1111 1112 1113 1114 1115 1116 1117 1118 1119 1120 1121 1122 1123 1124 1125 1126 1127 | |
compile_column_predicate(rule, source_column, target_column=None, *, source_dtype=None, target_dtype=None)
Compile a boolean match predicate for one source/target column pair.
Follows the canonical transform order documented on DiffRule, which is
the single source of truth shared with the local engine. So a pushdown
run and a local run evaluate the same pipeline. A setting a dialect
cannot reproduce raises ConfigError instead of compiling to something
that differs:
min_jaro_winkler_similarity, on every dialect;max_levenshtein_distance, on Postgres and DuckDB;datetime_formaton Postgres, or with a directive the dialect lacks;- a
regex_replacereplacement that refers to a group by name.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
rule
|
DiffRule
|
Semantic comparison overrides for the column. |
required |
source_column
|
str
|
Column name on the source relation. |
required |
target_column
|
str | None
|
Column name on the target relation. Defaults
to |
None
|
source_dtype
|
DataType | None
|
Probed source dtype, used to drop sentinels the column cannot hold. Each side is filtered separately because the two relations can disagree on a type. When None, sentinels are emitted unfiltered. |
None
|
target_dtype
|
DataType | None
|
Probed target dtype. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Boolean SQL expression that is true when the column values match. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If the rule sets one of the settings listed above that this dialect refuses. |
ConnectorError
|
If identifiers are empty or not allowlisted. |
Source code in src/veridelta/connectors/sql.py
compile_count_query(table)
Assemble a total row count query for one relation.
The count supplies the denominator for DiffSummary.mismatch_ratio, so
warehouse runs honor threshold the same way local comparisons do.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
table
|
str
|
Relation to count (optionally dotted catalog path). |
required |
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
|
str
|
for the active dialect. |
Raises:
| Type | Description |
|---|---|
ConnectorError
|
If the relation name is empty or not allowlisted. |
Source code in src/veridelta/connectors/sql.py
compile_duplicate_key_query(table, primary_keys, *, is_source, key_rules=None, types=None)
Assemble a count of the rows whose normalized key is not unique.
The local engine asserts uniqueness after normalization and reports
every row that shares its key with another. Summing the size of each
key group larger than one is that same number, and GROUP BY puts NULL
keys in one group, as Polars counts them as duplicates of each other.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
table
|
str
|
Relation to check (optionally dotted catalog path). |
required |
primary_keys
|
list[str]
|
Keys, spelled as the target stores them. |
required |
is_source
|
bool
|
Whether the relation is the source, which reads a
renamed key under its stored name and applies |
required |
key_rules
|
Sequence[DiffRule] | None
|
Key normalization, as for
|
None
|
types
|
ColumnTypes | None
|
Probed dtypes for this relation. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
A single-row statement whose |
str
|
every normalized key is unique. |
Raises:
| Type | Description |
|---|---|
ConnectorError
|
If the table or keys are empty, or a key rule does not name exactly one primary key. |
Source code in src/veridelta/connectors/sql.py
compile_missing_query(source_table, target_table, primary_keys, *, source_types=None, target_types=None, key_rules=None)
Assemble a LEFT JOIN anti-join for rows present only in the source.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source_table
|
str
|
Source relation (optionally dotted catalog path). |
required |
target_table
|
str
|
Target relation (optionally dotted catalog path). |
required |
primary_keys
|
list[str]
|
Join keys, spelled as the target stores them. |
required |
source_types
|
ColumnTypes | None
|
Probed source dtypes. |
None
|
target_types
|
ColumnTypes | None
|
Probed target dtypes. |
None
|
key_rules
|
Sequence[DiffRule] | None
|
Key normalization, as for
|
None
|
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
Source keys whose normalized value has no target counterpart. |
Raises:
| Type | Description |
|---|---|
ConnectorError
|
If tables or keys are empty, or a key rule does not name exactly one primary key. |
Source code in src/veridelta/connectors/sql.py
compile_query(source_table, target_table, primary_keys, rules, *, source_types=None, target_types=None, key_rules=None, wide_integers=frozenset(), type_drift=frozenset())
Assemble a changed-row inner-join query from tables, keys, and rules.
Stages 1 through 7 run once per column in a pair of CTEs, keys included. The join and the match predicates then read those projected values, so each stage appears once in the statement however many predicates read it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source_table
|
str
|
Source relation (optionally dotted catalog path). |
required |
target_table
|
str
|
Target relation (optionally dotted catalog path). |
required |
primary_keys
|
list[str]
|
Join keys, spelled as the target stores them. |
required |
rules
|
list[DiffRule]
|
Per-column semantic overrides. |
required |
source_types
|
ColumnTypes | None
|
Probed source dtypes, used to drop null sentinels the column cannot hold. When None, sentinels are emitted unfiltered. |
None
|
target_types
|
ColumnTypes | None
|
Probed target dtypes. |
None
|
key_rules
|
Sequence[DiffRule] | None
|
One rule per key that needs
normalizing, naming the stored source column and, for a renamed
key, its |
None
|
wide_integers
|
frozenset[str]
|
Compared columns, by target name,
that hold integers on both sides after normalization. Their
tolerance is measured in |
frozenset()
|
type_drift
|
frozenset[str]
|
Compared columns, by target name, that
|
frozenset()
|
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
|
str
|
keeping the joined rows where at least one compared column differs. |
|
str
|
Each predicate is wrapped in |
|
str
|
NULL reads as a mismatch rather than as an unknown that |
|
str
|
drops, matching the local engine's |
|
str
|
column is compared the statement selects no rows, since without |
|
str
|
match expressions the local engine reports nothing as changed. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If a rule sets |
ConnectorError
|
If tables or keys are empty, a rule is pattern-only,
|
Source code in src/veridelta/connectors/sql.py
796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 | |
compile_schema_probe_query(table)
Assemble a zero-row projection used to read a relation's columns.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
table
|
str
|
Relation to probe (optionally dotted catalog path). |
required |
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
|
str
|
metadata without scanning rows. |
Raises:
| Type | Description |
|---|---|
ConnectorError
|
If the relation name is empty or not allowlisted. |
Source code in src/veridelta/connectors/sql.py
compile_value_map_query(source_table, target_table, primary_keys, rules, *, min_support, sample_fraction=1.0, source_types=None, target_types=None, key_rules=None)
Assemble one statement that counts how source and target values line up.
Keys and candidate columns are normalized in the usual CTE pair and
joined once. Each column then contributes one UNION ALL branch,
labeled by its position in rules rather than by name, so no column
name becomes a string literal. A branch leaves out NULL sources and
values the column's existing map produced, as the local engine does.
Only exact predicates run here: the target differs from the source, at
least min_support rows agree, and the agreeing rows are more than
half of the source value's rows. Every confidence floor is above one
half, so this keeps a superset of what qualifies, and the engine
applies the floor itself, in floating point exactly as locally.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
source_table
|
str
|
Source relation (optionally dotted catalog path). |
required |
target_table
|
str
|
Target relation (optionally dotted catalog path). |
required |
primary_keys
|
list[str]
|
Join keys, spelled as the target stores them. |
required |
rules
|
list[DiffRule]
|
One rule per candidate column, naming its stored source column and, when renamed, the target's name. |
required |
min_support
|
int
|
Agreeing rows a pair needs. |
required |
sample_fraction
|
float
|
Share of source keys to read, chosen by a hash of the normalized keys. 1 reads every row. |
1.0
|
source_types
|
ColumnTypes | None
|
Probed source dtypes. |
None
|
target_types
|
ColumnTypes | None
|
Probed target dtypes. |
None
|
key_rules
|
Sequence[DiffRule] | None
|
Key normalization, as for
|
None
|
Returns:
| Type | Description |
|---|---|
str | None
|
str | None: A statement returning |
str | None
|
|
str | None
|
|
str | None
|
None when there is no candidate column. |
Raises:
| Type | Description |
|---|---|
ConnectorError
|
If tables or keys are empty, a rule is pattern-only,
a key rule does not name exactly one primary key, or
|
Source code in src/veridelta/connectors/sql.py
1129 1130 1131 1132 1133 1134 1135 1136 1137 1138 1139 1140 1141 1142 1143 1144 1145 1146 1147 1148 1149 1150 1151 1152 1153 1154 1155 1156 1157 1158 1159 1160 1161 1162 1163 1164 1165 1166 1167 1168 1169 1170 1171 1172 1173 1174 1175 1176 1177 1178 1179 1180 1181 1182 1183 1184 1185 1186 1187 1188 1189 1190 1191 1192 1193 1194 1195 1196 1197 1198 1199 1200 1201 | |
SnowflakeConnector
Bases: _CursorSession
Snowflake SQL warehouse connector backed by the optional Snowflake extra.
connect() opens a snowflake.connector session from the frozen
SnowflakeConfig; execute_pushdown runs compiler SQL on a fresh cursor
and fetches the result as Arrow. Install the driver with
uv add 'veridelta[snowflake]'; without it, connect() raises
ConnectorError with that hint instead of an ImportError.
Attributes:
| Name | Type | Description |
|---|---|---|
compiler |
SQLPushdownCompiler
|
Snowflake-dialect compiler the engine uses to build every statement this connector executes. |
Source code in src/veridelta/connectors/warehouse.py
__init__(config)
Initialize the connector with validated Snowflake settings.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
SnowflakeConfig
|
Frozen account, warehouse, and database settings. |
required |
Source code in src/veridelta/connectors/warehouse.py
connect()
Open a Snowflake session for subsequent pushdown statements.
Raises:
| Type | Description |
|---|---|
ConnectorError
|
If the Snowflake extra is missing or authentication fails. |
Source code in src/veridelta/connectors/warehouse.py
VerideltaConnector
Bases: ABC
The lifecycle every connector shares: connect(), close(), and the context manager.
A connector is one of two kinds, and the engine routes each source to one by its configuration:
- A
ReaderConnectorreads a source for the local engine. The lakehouse connectors (DeltaLakeConnector,IcebergConnector) open a Polarsscan_*handle, and the database and DuckDB connectors (DatabaseConnector,DuckDBConnector) read one table or query when they connect. Each hands the rows to the engine throughlazyframe(), and the diff runs in Polars. - A
PushdownSessionruns compiled SQL where the data lives. The warehouse connectors (SnowflakeConnector,DatabricksConnector,BigQueryConnector) hold a driver session, and the pushdown sessions (PostgresPushdownSession,DuckDBPushdownSession) serve two tables that both setpushdown. The engine compiles comparison SQL with the session'scompilerand callsexecute_pushdownfor each round-trip; results come back as Arrow wrapped in a LazyFrame.
Call connect() before anything else and close() when finished; the
connector is also a context manager whose exit calls close(). After
close() the connector is back in its unconnected state, so any further
call raises ConnectorError until connect() runs again.
Source code in src/veridelta/connectors/base.py
__enter__()
Return the connector unchanged; connect() stays an explicit call.
Returns:
| Type | Description |
|---|---|
Self
|
The connector itself, so |
__exit__(exc_type, exc, traceback)
Close the connector when leaving the context, error or not.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
exc_type
|
type[BaseException] | None
|
Pending exception type. |
required |
exc
|
BaseException | None
|
Pending exception. |
required |
traceback
|
TracebackType | None
|
Pending traceback. |
required |
Source code in src/veridelta/connectors/base.py
close()
Release what connect() opened.
Safe to call before connect() and safe to call twice. The default
holds no resources; connectors that open a driver session, a scan, or
a read override it. It is not abstract, so a subclass that holds
nothing need not define it.
Source code in src/veridelta/connectors/base.py
connect()
abstractmethod
Open the driver session, the scan, or the read that later calls use.
Raises:
| Type | Description |
|---|---|
ConnectorError
|
If the backend cannot be reached or its extra is missing. |
MCP server
The tools veridelta mcp serves, and the folder guard each one goes through. The SDK they run on comes with the mcp extra, and this module imports without it.
Serve Veridelta to an AI agent as Model Context Protocol tools.
veridelta mcp calls serve, which answers an agent's host over stdio through
the official MCP SDK, from the mcp extra. Each tool calls a function here. A
tool with a command returns the object that command prints with --json, so an
agent that knows the command line knows the tools.
The person who starts the server names the folders it may read configuration
files and data on this machine from, and a tool refuses a path outside them. A
tool returns findings, counts, and column names, never a value from the file.
Two tools return values from the data, and only when the person who starts the
server allows it: then at most a set number of rows. A side's query runs only
when that person allows queries too. The SDK is imported by the first
build_server(), not with this module, so veridelta imports without the
extra.
DEFAULT_ROW_CAP = 50
module-attribute
The most rows, or value map entries, one call returns unless the server is
started with another --max-rows.
DEFAULT_ROW_LIMIT = 20
module-attribute
The rows read_discrepancies returns when a call names no limit.
INSTRUCTIONS = "Veridelta compares two datasets under the rules in a YAML configuration file. Check a file with validate_config, and fix each error it reports, before you run it with run_comparison. Use describe_schema to list a side's columns when a rule must name one. Report counts and column names, and leave row values out of a reply unless the user asks for them. Three tools return row values, only when the server allows them. This server reads files only from the folders it was started with."
module-attribute
What the server tells an agent's host about itself when the host connects.
sdk = None
module-attribute
The SDK, imported by the first build_server() rather than here, since it
brings a web stack that the rest of Veridelta never loads. Tests set this
attribute.
DiscrepancyReport
Bases: TypedDict
What read_discrepancies returns: rows of one kind, up to the cap.
Attributes:
| Name | Type | Description |
|---|---|---|
kind |
Literal['added', 'removed', 'changed']
|
|
total |
int
|
How many rows of that kind the run found. |
rows |
list[dict[str, Any]]
|
The first of them, each a JSON object. A local run's |
truncated |
bool
|
Whether |
keys_only |
bool
|
Whether the pair was compared in place, such as two warehouse tables, which brings back primary keys alone. |
Source code in src/veridelta/mcp_server.py
ProposalReport
Bases: TypedDict
What propose_value_maps returns: the proposals, up to the cap.
Attributes:
| Name | Type | Description |
|---|---|---|
proposals |
list[dict[str, Any]]
|
Each proposal as |
total |
int
|
How many proposals there are. |
truncated |
bool
|
Whether |
Source code in src/veridelta/mcp_server.py
RunReport
Bases: TypedDict
What run_comparison returns: the summary veridelta run --json prints, with its verdict.
verdict is match when the comparison falls within threshold, and
drift otherwise, and exit_code is what veridelta run exits with, 0
or 1. artifacts_written says whether the rows that differ were written
to output_path. The fields between them are DiffSummary's.
Source code in src/veridelta/mcp_server.py
SchemaReport
Bases: TypedDict
What describe_schema returns: one side's columns, and none of its rows.
Attributes:
| Name | Type | Description |
|---|---|---|
side |
Literal['source', 'target']
|
|
columns |
dict[str, str]
|
Each column's name, as stored and before
|
Source code in src/veridelta/mcp_server.py
Settings
dataclass
What the person who starts the server allows, which no tool call can change.
Attributes:
| Name | Type | Description |
|---|---|---|
roots |
tuple[Path, ...]
|
The folders a tool may read a configuration file from, resolved on creation. A relative path in a tool call is read against the first. Every tool reads data on this machine only from them too. |
allow_row_values |
bool
|
Whether |
max_rows |
int
|
The most rows, value map entries, or example keys one of those calls returns. At least 1. Defaults to 50. |
allow_queries |
bool
|
Whether a tool may run a side's |
Source code in src/veridelta/mcp_server.py
__post_init__()
Resolve every root, so a path compares with them as the filesystem does.
Raises:
| Type | Description |
|---|---|
ConfigError
|
If no root is given, or |
Source code in src/veridelta/mcp_server.py
SuggestionReport
Bases: TypedDict
What suggest_rules returns: the suggestions, up to the cap.
Attributes:
| Name | Type | Description |
|---|---|---|
suggestions |
list[dict[str, Any]]
|
Each suggestion as |
total |
int
|
How many suggestions there are. |
truncated |
bool
|
Whether |
Source code in src/veridelta/mcp_server.py
build_server(settings)
Build the server and register its tools.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
settings
|
Settings
|
The roots every tool is held to. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
MCPServer |
MCPServer
|
The SDK's server, ready for |
Raises:
| Type | Description |
|---|---|
VerideltaError
|
If the |
Source code in src/veridelta/mcp_server.py
764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 | |
check_configuration(settings, path, *, schemas=False, allow_missing_env=False)
Check a configuration file for what would stop a run, as validate --json does.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
settings
|
Settings
|
The roots the server was started with. |
required |
path
|
str
|
The configuration file, under a root. |
required |
schemas
|
bool
|
Whether to also connect and check the rules against each side's columns, which reads no rows. |
False
|
allow_missing_env
|
bool
|
Whether an unset |
False
|
Returns:
| Name | Type | Description |
|---|---|---|
ValidationReport |
ValidationReport
|
The errors and warnings, with the resolved path. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If the path is outside the roots, or, with |
Source code in src/veridelta/mcp_server.py
check_data_paths(settings, source, target)
Refuse a side whose data on this machine lies outside the roots.
Every tool that opens a side calls this first, since a column name or an
error message can carry a file's text as a row does. The paths are
expanded and resolved as the readers do, so neither ~ nor a link leads
out. Data on another machine, such as an object store, a database server,
or a warehouse, is read as the command line reads it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
settings
|
Settings
|
The roots the server was started with. |
required |
source
|
SourceRef
|
The source configuration. |
required |
target
|
SourceRef
|
The target configuration. |
required |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If a file or folder a side reads is outside every root. |
Source code in src/veridelta/mcp_server.py
describe_side(settings, path, side)
List one side's columns and their types, as a run reads them before its first row.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
settings
|
Settings
|
The roots the server was started with. |
required |
path
|
str
|
The configuration file, under a root. |
required |
side
|
Literal['source', 'target']
|
The side to describe. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
SchemaReport |
SchemaReport
|
The side, and each of its columns mapped to its type. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If the file or its data is outside the roots, the file
does not load, or it reads the side through a |
ConnectorError
|
If the side cannot be reached or read. |
Source code in src/veridelta/mcp_server.py
environment_values(settings, path)
Return the value of each environment variable a configuration file references.
A tool masks each in its answer, since a path, a name, or an error built from the configuration carries the values it took. A file outside the roots, or one that cannot be read, gives none, and so does a variable that is unset or shorter than four characters.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
settings
|
Settings
|
The roots the server was started with. |
required |
path
|
str
|
The configuration file from the tool call. |
required |
Returns:
| Type | Description |
|---|---|
tuple[str, ...]
|
tuple[str, ...]: The values to mask. |
Source code in src/veridelta/mcp_server.py
propose_maps(settings, path, *, min_confidence=DEFAULT_MIN_CONFIDENCE, min_support=DEFAULT_MIN_SUPPORT, sample_fraction=1.0)
Propose value_map entries as veridelta crosswalk --json does, up to the server's cap.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
settings
|
Settings
|
The roots, the permission, and the cap. |
required |
path
|
str
|
The configuration file, under a root. |
required |
min_confidence
|
float
|
Share of a source value's rows that must agree on one target value. |
DEFAULT_MIN_CONFIDENCE
|
min_support
|
int
|
Agreeing rows an entry needs. |
DEFAULT_MIN_SUPPORT
|
sample_fraction
|
float
|
Share of source rows to read, chosen by primary key. |
1.0
|
Returns:
| Name | Type | Description |
|---|---|---|
ProposalReport |
ProposalReport
|
The proposals, how many there are, and whether some were left out. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If the server does not allow row values, the file or its
data is outside the roots, a side reads through a |
ConnectorError
|
If a source cannot be read. |
Source code in src/veridelta/mcp_server.py
read_rows(settings, path, kind, limit=DEFAULT_ROW_LIMIT)
Run the comparison and return the first rows of one kind, up to the server's cap.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
settings
|
Settings
|
The roots, the permission, and the cap. |
required |
path
|
str
|
The configuration file, under a root. |
required |
kind
|
Literal['added', 'removed', 'changed']
|
Which rows to return. |
required |
limit
|
int
|
The most rows the call asks for. The server's |
DEFAULT_ROW_LIMIT
|
Returns:
| Name | Type | Description |
|---|---|---|
DiscrepancyReport |
DiscrepancyReport
|
The rows, how many there are, and whether more were left out. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If the server does not allow row values, |
ConnectorError
|
If a source cannot be read. |
DataIntegrityError
|
If a primary key repeats on either side. |
Source code in src/veridelta/mcp_server.py
resolve_path(settings, path)
Resolve a path from a tool call, and refuse one outside the roots.
A relative path is read against the first root. Resolving follows symbolic
links and folds .., so neither leads out of a root.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
settings
|
Settings
|
The roots the server was started with. |
required |
path
|
str
|
The path from the tool call. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
Path |
Path
|
The resolved path, under one of the roots. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If the path resolves outside every root. |
Source code in src/veridelta/mcp_server.py
run_configuration(settings, path)
Compare the two datasets a configuration file names, as veridelta run --json does.
A run writes the rows that differ to output_path, so a configuration
whose output_path or data on this machine lies outside the roots is
refused before any row is read, as is a side's query unless the server
allows queries.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
settings
|
Settings
|
The roots the server was started with. |
required |
path
|
str
|
The configuration file, under a root. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
RunReport |
RunReport
|
The summary, with the verdict and the exit code. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If the file, its data, or its |
ConnectorError
|
If a source cannot be read. |
DataIntegrityError
|
If a primary key repeats on either side. |
Source code in src/veridelta/mcp_server.py
serve(settings)
Answer an agent's host over stdio until it disconnects.
A configuration's relative paths resolve against the working directory,
as on the command line, and the tools check them there. veridelta mcp
runs the server in its first root; a program that calls this from
another folder has its relative paths checked against that folder.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
settings
|
Settings
|
The roots every tool is held to. |
required |
Raises:
| Type | Description |
|---|---|
VerideltaError
|
If the |
Source code in src/veridelta/mcp_server.py
suggest(settings, path, *, max_share=DEFAULT_MAX_SHARE)
Suggest rules as veridelta suggest --json does, up to the server's cap.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
settings
|
Settings
|
The roots, the permission, and the cap. |
required |
path
|
str
|
The configuration file, under a root. |
required |
max_share
|
float
|
Largest gap a tolerance may explain, as a share of the larger of its two values. |
DEFAULT_MAX_SHARE
|
Returns:
| Name | Type | Description |
|---|---|---|
SuggestionReport |
SuggestionReport
|
The suggestions, how many there are, and whether some were left out. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If the server does not allow row values, the file or its
data is outside the roots, a side reads through a |
ConnectorError
|
If a source cannot be read. |
DataIntegrityError
|
If either dataset repeats a normalized primary key. |
Source code in src/veridelta/mcp_server.py
Reports
Standalone HTML reports and Markdown summaries, rendered from a DiffResult.
Standalone HTML reports and Markdown summaries for comparison results.
Renders a DiffResult into one self-contained file: no CDN reference, no
build step, no runtime dependency. Veridelta runs in CI, and CI runners are
often air-gapped, where a report that fetches a stylesheet from the internet
renders as unstyled text at exactly the moment someone needs to read it.
The Markdown summary is the short form CI posts to a job summary or a pull request comment: the verdict, the counts, and the columns that drifted. It lists changed values only when asked, and ends with the same counts as JSON in an HTML comment, which a reader never sees and a script can parse.
DEFAULT_MAX_ROWS = 1000
module-attribute
Rows embedded per table before truncation.
A diff of ten million rows would otherwise produce an HTML file nobody can open. The report states when it has truncated, so a reader never mistakes a capped table for the whole story.
render_html(result, *, max_rows=DEFAULT_MAX_ROWS)
Render a comparison result as a standalone HTML document.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
result
|
DiffResult
|
Completed comparison. |
required |
max_rows
|
int
|
Rows to embed per table before truncating. |
DEFAULT_MAX_ROWS
|
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
A complete HTML document with no external references. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If |
Source code in src/veridelta/report.py
198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 | |
render_markdown(result, *, max_rows=0)
Render a comparison result as a short Markdown summary.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
result
|
DiffResult
|
Completed comparison. |
required |
max_rows
|
int
|
Changed values to list, lowest keys first. 0, the default, lists none, since CI posts the summary where more people may read it than may read the data. |
0
|
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
The verdict, a table of counts, and the top drifting columns,
limited to the configured |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If |
Source code in src/veridelta/report.py
write_html(result, path, *, max_rows=DEFAULT_MAX_ROWS)
Write a standalone HTML report to disk.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
result
|
DiffResult
|
Completed comparison. |
required |
path
|
str | Path
|
Destination file. Parent directories are created. |
required |
max_rows
|
int
|
Rows to embed per table before truncating. |
DEFAULT_MAX_ROWS
|
Returns:
| Name | Type | Description |
|---|---|---|
Path |
Path
|
The file that was written. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If |
Examples:
>>> import polars as pl
>>> from veridelta.engine import DiffEngine
>>> from veridelta.models import DiffConfig
>>> source = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 20.0]})
>>> target = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 21.5]})
>>> result = DiffEngine(DiffConfig(primary_keys=["id"]), source, target).run()
>>> render_html(result).startswith("<!DOCTYPE html>")
True
>>> path = write_html(result, "reports/orders.html")
Source code in src/veridelta/report.py
write_markdown(result, path, *, max_rows=0)
Write the Markdown summary to disk.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
result
|
DiffResult
|
Completed comparison. |
required |
path
|
str | Path
|
Destination file. Parent directories are created. |
required |
max_rows
|
int
|
Changed values to list, as for |
0
|
Returns:
| Name | Type | Description |
|---|---|---|
Path |
Path
|
The file that was written. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If |
Examples:
>>> import polars as pl
>>> from veridelta.engine import DiffEngine
>>> from veridelta.models import DiffConfig
>>> source = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 20.0]})
>>> target = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 21.5]})
>>> result = DiffEngine(DiffConfig(primary_keys=["id"]), source, target).run()
>>> print(render_markdown(result).splitlines()[0])
### Veridelta: FAILED
>>> path = write_markdown(result, "summary.md")
Source code in src/veridelta/report.py
OpenTelemetry metrics
A run's counts, column drift, and verdict as OTLP/JSON metrics. See OpenTelemetry metrics.
OpenTelemetry metrics for comparison results.
Renders a DiffResult as one OTLP/JSON metrics export: the body an OTLP/HTTP
endpoint accepts at /v1/metrics, written on a single line so the
OpenTelemetry Collector's otlpjsonfile receiver can read the file as well.
No OpenTelemetry SDK is needed. The JSON follows the protobuf JSON mapping
that OTLP specifies, so 64-bit integers are written as strings.
Every value is a gauge: a snapshot of one run, which a backend graphs over time rather than adds up. The export carries counts, column names, and what was compared, never row values, connection URIs, credentials, or query text, since metrics usually end up in a third-party backend.
The standard OTEL_RESOURCE_ATTRIBUTES and OTEL_SERVICE_NAME variables add
resource attributes, as they do for an OpenTelemetry SDK.
send_otlp_metrics posts the same export to an OTLP/HTTP endpoint, which the
standard OTEL_EXPORTER_OTLP_* variables configure. Their headers often carry
an API key, so no header value reaches a log line or an error, and a redirect
is refused rather than followed to a second host.
SERVICE_NAME = 'veridelta'
module-attribute
service.name resource attribute and instrumentation scope name.
render_otlp_metrics(result, *, config_path=None, source=None, target=None, time_unix_nano=None)
Render a comparison as an OTLP/JSON metrics export.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
result
|
DiffResult
|
Completed comparison. |
required |
config_path
|
str | Path | None
|
Configuration file the run used,
recorded as |
None
|
source
|
SourceRef | None
|
Source configuration, recorded by type and by table or path. |
None
|
target
|
SourceRef | None
|
Target configuration, recorded likewise. |
None
|
time_unix_nano
|
int | None
|
Observation time, in nanoseconds since the epoch. Defaults to now. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
One line of JSON holding an |
Source code in src/veridelta/telemetry.py
send_otlp_metrics(result, *, config_path=None, source=None, target=None, time_unix_nano=None)
Send a comparison's OTLP/JSON metrics export to an OTLP/HTTP endpoint.
The standard variables configure the send. OTEL_EXPORTER_OTLP_METRICS_ENDPOINT
is the URL as written; otherwise /v1/metrics follows OTEL_EXPORTER_OTLP_ENDPOINT,
which defaults to http://localhost:4318. The HEADERS, TIMEOUT, and
PROTOCOL variables follow the same pattern, with the metrics variable
winning over the general one. The send is one attempt, with no retry.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
result
|
DiffResult
|
Completed comparison. |
required |
config_path
|
str | Path | None
|
As for |
None
|
source
|
SourceRef | None
|
As for |
None
|
target
|
SourceRef | None
|
As for |
None
|
time_unix_nano
|
int | None
|
As for |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
The endpoint the export reached, with no user part or query. |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If a variable holds an unusable value, or names a
protocol other than |
ConnectorError
|
If the endpoint cannot be reached, redirects, or answers with an HTTP error. |
Examples:
>>> import polars as pl
>>> from veridelta.engine import DiffEngine
>>> from veridelta.models import DiffConfig
>>> source = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 20.0]})
>>> target = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 21.5]})
>>> result = DiffEngine(DiffConfig(primary_keys=["id"]), source, target).run()
>>> send_otlp_metrics(result, config_path="veridelta.yaml")
'http://localhost:4318/v1/metrics'
Source code in src/veridelta/telemetry.py
450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 | |
write_otlp_metrics(result, path, *, config_path=None, source=None, target=None, time_unix_nano=None)
Write a comparison's OTLP/JSON metrics export to disk.
The file holds one line, ending in a newline, so it can be sent as is to an
OTLP/HTTP endpoint or read by the Collector's otlpjsonfile receiver.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
result
|
DiffResult
|
Completed comparison. |
required |
path
|
str | Path
|
Destination file. Parent directories are created. |
required |
config_path
|
str | Path | None
|
As for |
None
|
source
|
SourceRef | None
|
As for |
None
|
target
|
SourceRef | None
|
As for |
None
|
time_unix_nano
|
int | None
|
As for |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
Path |
Path
|
The file that was written. |
Examples:
>>> import polars as pl
>>> from veridelta.engine import DiffEngine
>>> from veridelta.models import DiffConfig
>>> source = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 20.0]})
>>> target = pl.LazyFrame({"id": [1, 2], "amount": [10.0, 21.5]})
>>> result = DiffEngine(DiffConfig(primary_keys=["id"]), source, target).run()
>>> import json
>>> export = json.loads(render_otlp_metrics(result, time_unix_nano=0))
>>> [
... metric["name"]
... for metric in export["resourceMetrics"][0]["scopeMetrics"][0]["metrics"]
... ]
['veridelta.dataset.rows', 'veridelta.diff.rows',
'veridelta.column.mismatched_rows', 'veridelta.diff.mismatch_ratio',
'veridelta.diff.match']
>>> path = write_otlp_metrics(result, "metrics.json")
Source code in src/veridelta/telemetry.py
Datasets
Sample datasets for the tutorials, downloaded once and cached.
Sample datasets for the tutorials and documentation examples.
load_nyc_taxi()
Load the NYC Taxi sample dataset.
The file is downloaded from the Veridelta repository once and cached under
~/.cache/veridelta/datasets. A download must match the sample's pinned
SHA-256, and stay under 1 MiB, before it is cached, and a cached copy that
no longer matches, such as a corrupt one, is downloaded again. A download
times out after 15 seconds.
Returns:
| Type | Description |
|---|---|
DataFrame
|
pl.DataFrame: The sample trips. |
Raises:
| Type | Description |
|---|---|
DatasetError
|
If the download fails, or is not the published sample. |