Changelog
Every release, newest first, with what changed and why. A release that needs a change from you also has a section in Upgrading.
v0.35.1 (2026-10-09)
Each run of the GitHub Action writes its reports to a folder of its own, so a job that
runs the action twice keeps both runs' reports. Before, the second run overwrote the
first's summary.json, summary.md, report.html, and otel-metrics.json, while the
first step's outputs still named those paths. Each artifact now holds only its own run's
files, and a run that ends before it writes a summary no longer copies the previous run's
summary into the job summary. The CI guide also shows the comment the action posts on a
pull request, and the rules page says what rules cannot do.
Fix
- give each GitHub Action run its own reports folder (#363)
v0.35.0 (2026-10-09)
The GitHub Action takes a baseline input: a baseline file, relative to
working-directory, whose drift the run accepts, so a pull request that changes data on
purpose fails only on drift the file does not list. Write the file once with
veridelta run --save-baseline accepted.json, commit it next to the configuration, and
pass baseline: accepted.json. The path reaches veridelta run --baseline as one
argument, whatever it holds.
Accepting drift in the CI
guide shows the step. The GitLab template does not take the input, since it stays as it
is unless a GitLab user asks.
Feat
- a baseline input on the GitHub Action (#360)
v0.34.0 (2026-10-09)
Baseline and AcceptedChange now import from veridelta, as the other models a user
builds do, so a baseline built in Python no longer needs from veridelta.models import
Baseline. veridelta.models still exports both. Accepting
drift gains a Python example:
build the file from a result with Baseline.of, write it, read it back, and pass it to
run_from_configs as baseline.
Feat
- import Baseline and AcceptedChange from veridelta (#358)
v0.33.6 (2026-10-09)
No behavior changed. engine.py held 3,421 lines, most of them helpers around one class.
It now holds DiffEngine and the names it exports, in 1,102 lines. The helpers live in
nine private modules beside it, each with one job: rule resolution, suggest, local
matching, reading, warehouse routing, results, pushdown, config checks, and value maps.
A decision record
says why. Every name veridelta.engine exports is the same, and the test suites compile
the same SQL and make the same engine calls, in the same order, as 0.33.5 did. Three
attributes left veridelta.engine, since the code that reads them moved: the fastexcel
and rapidfuzz_distance probes, which the API page showed, and logger. The logger keeps
its name, so logging.getLogger("veridelta.engine") returns the same logger, and
--verbose prints the same lines.
Refactor
- move crosswalk's proposals out of engine.py into _value_maps.py (#355)
- move the config checks out of engine.py into _checks.py (#354)
- move pushdown planning and collecting into _pushdown.py (#353)
- check a pushdown's probes without going through DiffEngine (#352)
- move pushdown rule resolution out of engine.py into _pushdown.py (#351)
- move artifacts and baselines out of engine.py into _results.py (#350)
- move warehouse routing out of engine.py into _warehouses.py (#349)
- move the loaders out of engine.py into _reading.py (#348)
- move local matching out of engine.py into _matching.py (#347)
- move suggest's proposals out of engine.py into _suggest.py (#346)
- move rule resolution out of engine.py into _resolution.py (#345)
v0.33.5 (2026-10-09)
No behavior changed. PyPI shows the README of the newest release, so this release brings two README changes to the project page there. A Status section says where the project stands: it is in alpha, its tests run on generated data, nobody but its maintainer is known to have run it, and its pushdown SQL has not yet run in a live cloud warehouse. The version badge now comes from shields.io, which reads PyPI directly. The old one, from badge.fury.io, still showed 0.14.0. GitHub Releases also take their notes from this changelog now.
v0.33.4 (2026-10-08)
A pushdown run, its check under validate --schemas, and its row sample each compiled the
same statements, with the same long argument lists. They now share one compile step, so
the check compiles exactly the SQL a run executes. The statements, and the order they run
in, are the same. One ordering changes: a statement that fails to compile now raises its
ConfigError before the duplicate key queries run, as a local run fails on a rule before it
reads a key. Before, repeated keys were reported first.
Refactor
- compile the pushdown statements once (#337)
v0.33.3 (2026-10-08)
No behavior changed. build_parser added every command's arguments in one function of 270
lines. It now names each command, its parents, and its help in one place, and one function
per command adds that command's arguments. Every command's --help prints the same text as
before, byte for byte. The docs also gained a
troubleshooting guide, which
starts from what a failed or surprising run shows and gives the cause and the fix.
Refactor
- one function per command in build_parser (#335)
v0.33.2 (2026-10-08)
A primary key whose two sides held two types a join cannot pair failed inside the join,
with an error that named no side and called itself a bug. A CSV and a Parquet export that
type one key differently were enough: Int64 against text, Int64 against Float64, two
decimal precisions, or a date against a timestamp. Each key's two types are now checked
before rows are paired, in run, validate --schemas, suggest, and crosswalk. Such a
key raises ConfigError, naming it, both of its types, and the cast_to rule that
resolves it. A key of one type, two integer types, or two float types pairs as before.
Pushdown is unchanged, since a database converts one side of the key itself, and
Pushdown
lists that difference. The docs also gained a How-to guides section, a changelog page,
and tests that hold them to the command line, the GitHub Action, and the MCP server.
Fix
- name a primary key whose two sides cannot pair rows (#332)
v0.33.1 (2026-10-08)
No behavior changed. The Snowflake and Databricks connectors share one cursor session for
what they did the same way: the check before each statement, a fresh cursor's Arrow
fetch, and close(). Each keeps only how it signs in and which fetch its driver offers,
and the statements, log lines, and errors stay the same. This is also the first release
the new pipeline publishes (#324): it builds in one job, and publishes from another that
runs no code from the release, with an attestation for each file.
Refactor
- one cursor session for Snowflake and Databricks (#325)
v0.33.0 (2026-10-08)
Pushdown and a local run could reach different verdicts on a column whose two sides hold
different types. A local run casts the target to the source type, but a database
converted one side by its own rules. In DuckDB, the text "007" against the integer 7,
or ISO text against a timestamp, matched where a local run reported drift. A date against
a timestamp at noon went the other way, and legacy text against a timestamp failed with a
raw conversion error. Pushdown now refuses such a column before reading a row, and names
it with both of its types. Two numeric types still compare by value, and a cast_to that
gives both sides one type resolves the rest. The Upgrading
page shows how. A new seeded-drift suite found it, and now holds both engines to drift
whose effect is known.
BREAKING CHANGE
- A pushdown comparison of a column whose two sides hold different types, other than two
numeric types, raises
ConfigError. Addcast_toto the column's rule, or setstrict_types: true.
Fix
- refuse a pushdown column of two types a database would convert (#317)
v0.32.0 (2026-10-08)
veridelta.config no longer re-exports the seven source models. Import BigQueryConfig,
DatabaseConfig, DatabricksConfig, DeltaLakeConfig, DuckDBConfig, IcebergConfig,
and SnowflakeConfig from veridelta or veridelta.models, which export them too. A new
Upgrading page says what each breaking
release since 0.6.0 asks of you. This is the release's one breaking change, so under 0.x it
bumps the minor version.
BREAKING CHANGE
from veridelta.config import SnowflakeConfig, and the same import of the six other source models, raisesImportError. Import them fromverideltaorveridelta.models.
Refactor
- stop re-exporting the source models from veridelta.config (#315)
v0.31.2 (2026-10-08)
No behavior changed. ReaderConnector now holds what connect() opened, and gives every
reader the same lazyframe() and close(). The database, DuckDB, Delta Lake, and Iceberg
readers each wrote that pair themselves, and now keep only connect(). Each still names
its own kind of source when read before connect(), in the words it used before. A reader
of your own may still define both methods.
Refactor
- give ReaderConnector a concrete lazyframe and close (#312)
v0.31.1 (2026-10-08)
Every message for an optional extra that is not installed now reads the same way: what
needs the extra, which extra it is, and the command that installs it. Twelve places wrote
their own, in five wordings, and now share missing_extra in veridelta.exceptions. Each
still raises the error it raised before, so exit codes do not change.
Refactor
- write the message for a missing extra in one place (#310)
v0.31.0 (2026-10-08)
veridelta run legacy.csv modern.csv --key id compares two files with no configuration
file. It runs what a file holding only primary_keys and the two paths would run, with the
same summary and exit codes, and each file's format still follows its suffix. Repeat
--key for a key of several columns. Two files without --key, --key without two
files, or two files with -c exit 2. In Python, files_config in veridelta.config
builds the same configuration. Each tutorial also opens in Google Colab now, where its
first cell installs this release.
Feat
- compare two files with run and --key, no configuration file (#307)
v0.30.0 (2026-10-08)
veridelta mcp serves a sixth tool, suggest_rules. It returns what
veridelta suggest --json prints: rules that would explain each column's differences,
each with its evidence. Its max_share argument works as the command's --max-share
does. Each suggestion's examples hold primary keys from the data, so the tool answers only
on a server started with --allow-row-values, and --max-rows caps those keys. It reads
both sides locally, so a pair compared in place is refused, as the command refuses it.
Feat
- mcp: suggest rules through a suggest_rules tool (#305)
v0.29.0 (2026-10-08)
veridelta validate --schemas now warns about a rule whose column_names holds a name
that neither side has, after normalize_column_names when it is on. Such a rule forgives
nothing, so a misspelled name used to pass unnoticed. The warning names each such column.
It works for files, databases, lakehouse tables, and warehouse pairs alike, and the MCP
server's validate_config reports it too. A warning leaves the exit code at 0, as before.
Feat
- warn when a rule names a column neither side has (#303)
v0.28.0 (2026-10-08)
SQLPushdownCompiler's seven compile_ methods that read both sides no longer take
source_alias and target_alias. No caller passed them: the engine always used the
defaults. Every statement still names the two relations src and tgt, byte for byte, so
a comparison inside a warehouse runs the same SQL as before. This is the one breaking
change of the release, so under 0.x it bumps the minor version.
Refactor
- sql: remove the unused source_alias and target_alias parameters (#295)
BREAKING CHANGE
compile_column_predicate,compile_query,compile_changed_sample_query,compile_missing_query,compile_added_query,compile_column_mismatch_query, andcompile_value_map_queryno longer acceptsource_aliasortarget_alias. A call that passes either raisesTypeError. Drop the argument, since the aliases were alwayssrcandtgt.
v0.27.1 (2026-10-08)
Every extra that installs pyarrow now asks for 23.0.1 or later, where it allowed 14.0.1. pyarrow before 23.0.1 frees memory still in use while reading an Arrow IPC file (CVE-2026-25087). A fresh install already got 23.0.1, so this changes an install only where another tool pinned an older pyarrow. The Snowflake extra keeps its cap below 24 on Python 3.14.
Fix
- deps: raise the pyarrow floor past CVE-2026-25087 (#293)
v0.27.0 (2026-10-07)
veridelta run --save-baseline accepted.json writes the file run --baseline reads:
every row of drift the run finds, in key order, with each changed row's differing
columns. A change made on purpose is accepted in one step. With --baseline too, the new
file keeps what the old one accepted, and leaves out entries for rows that no longer
drift. In Python, Baseline.of(result) builds the same file. The run's verdict and exit
code stay as they are.
Feat
- write the drift a run finds as a baseline with run --save-baseline (#283)
v0.26.0 (2026-10-07)
veridelta run --baseline accepted.json accepts the drift the file lists and fails only
on drift it does not list. The file lists added and removed rows by primary key, and
changed rows with the columns whose drift is accepted on each: drift in any other column
of that row still counts. Accepted drift is left out of the counts, the verdict, the
artifacts, and the reports. The summary gains accepted_count, in run --json, in the
MCP server's run_comparison, and in the published run schema, and
veridelta schema baseline prints the file's JSON Schema. A pair compared where it is
stored refuses --baseline.
Feat
- accept the drift a baseline file lists with run --baseline (#280)
v0.25.0 (2026-10-07)
veridelta suggest now proposes datetime_format for a column with text on one side and
dates or timestamps on the other, when the text reads as those dates under one format.
The common formats are tried, with or without a time, and the one that matches the most
differing rows wins, so the dates themselves settle whether the day or the month comes
first. A timestamp with a zone is left out.
Feat
- suggest a date format for text that reads as the dates beside it (#278)
v0.24.0 (2026-10-07)
veridelta suggest now proposes null_values for a column of any type when a common
spelling of NULL stands on one side where the other side is NULL: an empty string, N/A,
NULL, None, -, and the like, in any case, or a number such as -999. A real value
beside NULL, such as a city, is a change, and false is never suggested. A suggested
sentinel joins the column's sentinels today, and a column with both rounding and a
sentinel gets one rule with both settings.
Feat
- suggest null sentinels where one side spells NULL (#276)
v0.23.0 (2026-10-07)
veridelta suggest now proposes rules for text columns too. A column gets
whitespace_mode: both when its differing values match once both ends are stripped,
case_insensitive: true when they match once lowercased, and both settings only when
some rows need both. Text that differs in anything else is a change, and no rule explains
it. A rule that would make a row that matches today differ is never suggested, such as
case folding ahead of a value_map whose keys are capitals.
Feat
- suggest trimming and case folding for text that differs only in them (#274)
v0.22.0 (2026-10-07)
veridelta suggest runs the comparison, then suggests rules that would explain the
differences it finds, each with its evidence. Its first kind of rule is a numeric
tolerance. A column gets one when every gap between its differing values is at most
--max-share, 1% by default, of the larger of the two: absolute when the gaps stay about
one size, as rounding leaves them, and relative when they grow with the values. The
tolerance is the first round value above the largest gap, and the comparison runs again
with it to count the rows it explains. The rules print as YAML on stdout, ready to paste
first in rules, and the evidence on stderr, or both as JSON under --json, with a
published schema. No model is called. In Python, DiffEngine.suggest_rules() returns
RuleSuggestion models.
The results page shows the HTML report in light and dark, and a link to the docs shows a
preview card. The AI agents page shows a client calling the MCP tools, and demo/agent/
holds a kit for recording an agent's session. The sdist leaves out the recordings and
screenshots, which made up about 40% of its size.
Feat
- suggest a numeric tolerance from the pairs that differ (#272)
v0.21.2 (2026-10-07)
The run report counts one in the singular: a column with one mismatch reads "1 mismatch", and a target with one more row reads "+1 row", where both took the plural.
Fix
- count one mismatch and one row in the singular (#265)
v0.21.1 (2026-10-07)
While a run compares, veridelta run says "Comparing..." on stderr, where it said
"Executing semantic diff...".
The AI agents page and the agent skill give an agent the loop that fixes drift its own
pipeline change causes. Tutorial 6 compares two model evaluation runs. The command line
page shows recordings of validate and crosswalk, each beside the text it prints.
Fix
- say "Comparing..." while a run compares (#263)
v0.21.0 (2026-10-07)
The Markdown summary that CI posts to a pull request ends with an HTML comment, which
GitHub and GitLab hide from readers. It names the published run schema and holds the
run's summary as JSON on one line, so an agent that reads the comment through the API can
parse the result without running the comparison again. The JSON is what
veridelta run --json prints, except that column_mismatches lists only the columns the
drift table shows, and is left out when report_top_columns_limit is 0. <, >, and &
are written as JSON escapes, so a column name cannot close the comment. The 60,000 byte
budget for changed values counts the comment.
Feat
- end the Markdown summary with its counts as JSON for agents (#256)
v0.20.0 (2026-10-07)
veridelta schema takes an optional name, run, validate, crosswalk, or error, and
prints the JSON Schema of what that command prints with --json, or of the error object
any command prints at exit code 3. With no name, it prints the configuration schema, as
before. Each schema is published under docs/schema/ at the URL its $id names, and a
test validates real output from each command against it. typing-extensions is now a
declared dependency; Pydantic already required it.
The README and the docs home name the MCP server, and the user guide, the product documents, and the contributor guide now say what the code does.
Feat
- publish a JSON Schema for what each command prints with --json (#254)
v0.19.20 (2026-10-07)
No behavior changed. The engine opens and closes a warehouse session in one place, and the
package logger holds the one NullHandler that seven modules each added. A program that
adds no handler still sees no records from Veridelta.
Refactor
- open a warehouse session in one place, and silence logging once (#248)
v0.19.19 (2026-10-07)
No behavior changed, and every SQL statement is byte for byte the same: the unit and
integration suites compile 2,316 statements, and each is the same before and after. The
allowlist check every configured name passes before SQL is one function, the value map
query names its aliases once, and veridelta.models.POSTGRES_SCHEMES names the Postgres
URI schemes once.
Refactor
- name each SQL alias, the allowlist gate, and the Postgres schemes once (#246)
v0.19.18 (2026-10-07)
No behavior changed. The verdict and the reported mismatch_ratio divide by the source
rows through one function, mismatch_ratio_of, and five places that normalized a column
name share normalize_column_name. Both are in veridelta.models.
Refactor
- compute the mismatch ratio and normalize a column name in one place each (#244)
v0.19.17 (2026-10-07)
The engine writes each check once: the timezone guard, the null equality of stage 9, the
probe for each optional extra, and the names of the two sides. Results are unchanged. One
message changed: a timezone rule on a naive timestamp now gives the same advice on a
local run and in pushdown, to store the column with a zone or parse it with an offset such
as %z.
Refactor
- write each engine check once (#242)
v0.19.16 (2026-10-07)
No behavior changed. validate --json and the MCP server's validate_config build their
report through one function, validation_report in veridelta.mcp_server, and print the
same JSON as before. Two argument parsers share their whole-number parse, with every
message unchanged.
Refactor
- build the validation report once, and parse whole numbers once (#240)
v0.19.15 (2026-10-07)
A schema_mode failure lists the columns that break the mode, sorted and quoted, where it
printed a Python set in an order that changed between runs. exact says which columns
only the source has and which only the target has, and a run from a configuration says
what each side read. The messages keep their opening words.
Fix
- list the columns a schema_mode refuses in a fixed order (#227)
v0.19.14 (2026-10-07)
The HTML report and the Markdown summary say a keys-only comparison ran as pushdown,
inside the database that stores both tables. They said warehouse pushdown, which was wrong
for two Postgres or DuckDB tables that set pushdown.
Fix
- say pushdown, not warehouse pushdown, in the reports (#225)
v0.19.13 (2026-10-07)
OTEL_EXPORTER_OTLP_TIMEOUT must be ASCII digits, so a value such as ² names the
variable rather than crashing the send. A send with OTEL_EXPORTER_OTLP_HEADERS set warns
when it goes over plain http:// to another machine, naming the endpoint and never the
headers, since they often carry an API key.
Fix
- refuse a non-ASCII OTLP timeout, and warn on headers over http (#223)
v0.19.12 (2026-10-07)
load_nyc_taxi caches a download only when it is under 1 MiB and its SHA-256 matches the
published sample's. A cached copy that no longer matches, a corrupt one included, is
downloaded again with a warning.
Fix
- check the sample dataset's SHA-256 and size before caching it (#221)
v0.19.11 (2026-10-07)
The GitHub Action checks its extras, version, and artifact-name inputs before it
installs anything. A value outside the pattern each input takes fails the step with exit
code 2, the command line's code for invalid arguments. Before, a workflow that passed text
from a pull request into an input could add a requirement, or a second line to the step's
outputs.
Fix
- check the Action's extras, version, and artifact name before use (#219)
v0.19.10 (2026-10-07)
The snowflake extra needs snowflake-connector-python 4.7.3 or later, past six
advisories in earlier releases, such as logged credentials and a skipped TLS hostname
check. The snowflake and databricks extras need pyarrow 14.0.1 or later, past a flaw
that runs code from a crafted IPC file. An install that already has an older driver now
upgrades it.
Fix
- deps: raise the floors of two drivers past their advisories (#217)
v0.19.9 (2026-10-07)
Errors and logs name a file or table without the query of its URL, which can hold a token
or the signature of a pre-signed link, and without a login in its user part. Before, only
telemetry did: a failed read of a pre-signed s3:// link printed its signature, and a
database URI kept a ?password= parameter in every log line. redacted_location in
veridelta.models is the one redactor for all of them.
Fix
- leave the query and login out of every location printed (#215)
v0.19.8 (2026-10-07)
A Databricks connection that fails masks the access token in its error as ***, and no
longer chains the driver's exception, whose traceback printed the message again. Snowflake
already did both. BigQuery's connection error is no longer chained either, so the three
warehouses report a failed connection alike.
Fix
- keep the Databricks token out of a connection error (#213)
v0.19.7 (2026-10-07)
A failure Polars reports as a panic now exits 3, as any other error does, and --json
prints it as one error object. Before, Python exited 1 with a traceback: the code
veridelta run gives for drift, so CI read a crash as drift. The MCP server fails the one
call with the panic's type and message, where it would have stopped.
Fix
- report a Polars panic as an error, exit 3, not as drift (#211)
v0.19.6 (2026-10-07)
A run whose differing rows hold a binary column, such as a MySQL BIT column, writes its
discrepancy files in every format. csv, json, and ndjson hold the bytes as
hexadecimal text, while parquet and arrow keep them. Before, Polars refused the column
in CSV and panicked on it in JSON, and the run stopped with exit code 1. A column that
nests binary values in a list or a struct fails with a ConfigError that names the
formats which hold it.
Fix
- write a binary column to every artifact format (#209)
v0.19.5 (2026-10-07)
veridelta mcp describes its tools as they behave. propose_value_maps now says that
sample_fraction is a share of source rows chosen by primary key, as --sample-fraction
does, not a share of joined rows. The instructions a host reads say that two tools return
row values only when the server allows them, and that the server reads files only from the
folders it was started with.
Fix
- describe the MCP tools as they behave (#207)
v0.19.4 (2026-10-07)
veridelta mcp masks every value a configuration takes from an environment variable, four
characters or longer, wherever it appears in a tool's answer. Before, path: ${SECRET}.csv
returned the value inside a read's error, and a column named after a variable returned it
in a schema. veridelta.config.referenced_variables names the variables a configuration's
text references.
Fix
- mask values taken from the environment in MCP answers (#205)
v0.19.3 (2026-10-07)
veridelta mcp runs a side's query only when it is started with the new
--allow-queries, since a query runs as written, with the configuration's credentials.
Before, a DuckDB query such as SELECT * FROM read_text(...) could read a file outside
the roots, and a MotherDuck query ran read-write with the server's token. Every DuckDB file
the server opens is now held to the roots: extensions stop loading on their own, files open
only under the roots, and the configuration locks before the first read. A MotherDuck
connection is not held to the roots. A DuckDB connection whose setup fails is now closed.
Fix
- gate MCP queries behind --allow-queries and hold DuckDB to the roots (#203)
v0.19.2 (2026-10-07)
The MCP server checks a configuration's relative data path, and its output_path, where
the reader and the artifact writer open it: the working directory. Before, it checked them
against the first --root folder, which agrees only because veridelta mcp runs there. A
program that called serve() or build_server() from another folder passed the check for
a file under the root, then read or wrote the one beside the program. The command line is
unchanged.
Fix
- check a configuration's relative paths where they are opened (#201)
v0.19.1 (2026-10-07)
Every veridelta mcp tool that opens a side now reads data on this machine only from under
the server's --root folders. Before, only read_discrepancies and propose_value_maps
checked, while run_comparison, describe_schema, and validate_config with schemas
opened a file anywhere. A column name can carry a file's text: reader options that skip
lines and split on an unused separator turn any line of any file into the header, which
describe_schema returned. A check without schemas opens no data and is unchanged.
The PyPI summary, the docs, and veridelta --help now describe Veridelta in one sentence,
and the pushdown guide says which services the generated SQL has run in.
Fix
- hold every MCP tool that opens a side to the server's folders (#199)
v0.19.0 (2026-10-07)
veridelta mcp gains its last two tools, which return values from the data.
read_discrepancies runs the comparison and returns the rows of one kind, added,
removed, or changed, with how many there are. propose_value_maps returns what
veridelta crosswalk --json prints. Both refuse unless the server is started with
--allow-row-values, and then return at most --max-rows rows or value map entries a
call, 50 by default. They read files on this machine only from under the server's
--root folders, with ~ expanded and links followed, and read data elsewhere, such as
an object store or a warehouse, as the command line does.
Feat
- read rows and propose value maps from the MCP server when allowed (#195)
v0.18.0 (2026-10-07)
veridelta mcp gains a third tool, describe_schema, which lists one side's columns and
their types and reads no rows. It returns the side and a map from each column's name, as
stored, to its type as Polars names it, such as Int64. A side read through a query is
refused, since only running the query would name its columns.
In Python, DiffEngine.read_schema(config) returns one side's pl.Schema without
returning a row. A file is read by its loader, a database or DuckDB table by a probe
that returns no rows, and a warehouse table, or a table with pushdown, by the probe a
pushdown run starts with, in a session of its own.
Feat
- describe a side's schema from the MCP server (#193)
v0.17.0 (2026-10-07)
veridelta mcp gains a second tool, run_comparison, which runs the comparison a
configuration file describes. It returns the summary veridelta run --json prints, with
three more fields: verdict, which is match or drift; exit_code, the 0 or 1 that
veridelta run exits with; and artifacts_written, which says whether the rows that
differ were written. A run writes those rows to the configuration's output_path, so the
tool refuses an output_path outside the server's --root folders before it reads a row.
A configuration that cannot run fails the call with its error's type and message, as
run --json reports a failure.
Feat
- run a comparison from the MCP server (#191)
v0.16.0 (2026-10-07)
veridelta mcp serves Veridelta's checks to an AI agent's host as Model Context
Protocol tools, over stdio. It needs the new mcp extra, which installs the official
MCP Python SDK and is part of all. Its first tool, validate_config, returns the
object veridelta validate --json prints. The person who starts the server names the
folders a tool may read with --root, the current directory by default, and a path
outside them fails the call. A tool returns findings, never a value from the
configuration file. The AI agents page shows how to register the server with Claude
Code or Cursor.
In Python, DiffEngine.check_config_file(path) loads a configuration file and returns
the findings veridelta validate reports, with a file that does not load as one error.
Feat
- add
veridelta mcp, a Model Context Protocol server with the toolvalidate_config
v0.15.0 (2026-10-07)
A file source reads its format from the path's suffix when the key is absent: .csv;
.parquet or .pq; .json; .ndjson or .jsonl; .arrow, .ipc, or .feather;
.avro; and .xlsx or .xls, in any case and before a ? query or # fragment. A
format you write always wins, and a suffix Veridelta does not know keeps the csv
default. The smallest configuration is now the two paths and the keys that pair their
rows, and the README, the docs home, and tutorial 2 lead with it, which is the file the
README's recording runs.
Two names deprecated in 0.14 are gone. veridelta.DataIngestor, which loaded and aligned
two sources for a DiffEngine that aligns them again, goes: call
DiffEngine.run_from_configs(diff, source, target). VerideltaConnector.fetch_schema(),
which nothing in the package called, goes with the code that served only it: read a
reader's schema with lazyframe().collect_schema(), and probe a warehouse table with
compile_schema_probe_query(table) through execute_pushdown(sql, query_type="schema").
The connector base is split in two. VerideltaConnector declares the lifecycle,
connect() and close(), and nothing else. ReaderConnector adds lazyframe(), which
the Delta Lake, Iceberg, database, and DuckDB connectors implement. PushdownSession adds
compiler and execute_pushdown(), which the Snowflake, Databricks, and BigQuery
connectors and the Postgres and DuckDB pushdown sessions implement. A reader no longer has
an execute_pushdown method. A subclass written against the old base now subclasses the
kind it is.
This release ships before the live warehouse suite's first run against Snowflake, Databricks, BigQuery, and MotherDuck, by the maintainer's decision. That run comes when the maintainer has the time, and a difference it finds ships as a 0.15 patch.
Feat
- infer a file's format from its extension
Refactor
- remove
DataIngestor, deprecated in 0.14.6 - remove
fetch_schema, deprecated in 0.14.7 - split the connector base into readers and pushdown sessions
BREAKING CHANGE
veridelta.DataIngestorno longer exists. CallDiffEngine.run_from_configs(diff, source, target), which loads, aligns, and compares both sides.VerideltaConnector.fetch_schema()no longer exists. Read a reader's schema withlazyframe().collect_schema(), and probe a warehouse table withcompile_schema_probe_query(table)throughexecute_pushdown(sql, query_type="schema").VerideltaConnectordeclaresconnect()andclose()only. A connector the local engine reads subclassesReaderConnectorand defineslazyframe(); a connector that runs compiled SQL subclassesPushdownSessionand definescompilerandexecute_pushdown(). A reader no longer has anexecute_pushdownmethod.
v0.14.10 (2026-10-06)
veridelta.datasets computes its Git ref in one private function, where two public names
nothing else used built it, and logs through the logger's own formatting rather than an
f-string. The notice that the NYC taxi sample is downloading moved from WARNING to INFO,
since a download a tutorial asks for is routine; a corrupted cache still logs at WARNING.
Nothing else changed.
Refactor
- tidy
datasets.py: one private ref constant, and the download notice logged at INFO without an f-string
v0.14.9 (2026-10-06)
The connector base now refuses SQL pushdown by default, with one message, so a reader
that is compared locally no longer implements execute_pushdown only to raise. The Delta
Lake, Iceberg, database, and DuckDB readers lose those overrides and their three copies of
the message. A warehouse connector or a pushdown session overrides the default, as before.
The refusal's wording changed: it now says the source is compared locally and names
connect(), where it said the method was warehouse-only.
Refactor
- refuse pushdown from the connector base, so no reader implements
execute_pushdownonly to raise
v0.14.8 (2026-10-06)
The connectors now share one secret scrubber. Three of them each masked credentials in
driver output their own way, and two carried an identical property that names what a
side reads. mask_secrets and read_subject in veridelta.connectors.base replace
them, and every message, log line, and exception keeps its wording. Nothing a user sees
changed.
Refactor
- share one secret scrubber and one read subject across the warehouse, database, and DuckDB connectors
v0.14.7 (2026-10-06)
Every connector's fetch_schema now warns that it goes in 0.15.0. The method has nine
implementations and no caller in the package: the engine reads a side's columns from the
frame it loads. The method, its implementations, and the code that serves only it stay
until 0.15.0, where the removal ships as a breaking change. Nothing else changed.
Refactor
- deprecate
fetch_schemaon every connector with a warning, ahead of its removal in 0.15.0
v0.14.6 (2026-10-06)
DataIngestor now warns on construction that it goes in 0.15.0. No code in the package
calls it, and its own docstring warned against the one thing it is for: feeding its
output to DiffEngine aligns both sides twice. DiffEngine.run_from_configs(diff,
source, target) loads and aligns the sides itself. The class and its export stay until
0.15.0, and nothing else changed.
Refactor
- deprecate
DataIngestorwith a warning on construction, ahead of its removal in 0.15.0
v0.14.5 (2026-10-06)
The package behaves as 0.14.4 did: no command, rule, connector, or report changed. One
helper, optional_module in veridelta.connectors.base, now probes every extra the
connectors and the engine read through. Each probe was a try block of its own, with two
coverage pragmas, eight in all, and the engine kept a private copy of the helper that the
connectors could not import. Every probe keeps its name, so a test that patches one is
unchanged.
Refactor
- share one optional-import helper across the connectors and the engine, and drop the eight coverage pragmas the old probes carried
v0.14.4 (2026-10-06)
-v now prints each file a local run opens, with its format and its width, on the
veridelta.engine logger. Only the connectors logged before, so a run over two files
printed nothing under --verbose, though the 0.14.0 changelog promised each read. A run
without the flag prints no record, as before.
Fix
- log each file a local run opens under
--verbose, with its format and column count
v0.14.3 (2026-10-06)
A configuration failure inside source or target now names the block. A bad format
in the target was reported as [file -> format]. Pydantic opens the location with the
type tag of the model it chose, which names neither side. The location now opens
with source or target, so the line reads [target -> format], and a block whose
type matches no model reads [target]. The comparison settings keep their locations,
such as [rules -> 0 -> case_insensitive].
Fix
- name
sourceortargetin the location of a configuration failure inside either block, in place of Pydantic'stypetag
v0.14.2 (2026-10-06)
A missing primary key now names what each side read. format defaults to csv, so a
.parquet path without it is scanned as text, and the run stopped with "Primary keys
missing in SOURCE", which named a symptom. The error now names the side, the file, and the
format it was read as, and says when format was not set. It lists the columns it found,
so a newcomer changes format rather than the key. A Delta Lake or Iceberg side is named
by its table URI, and a database or DuckDB side by its table or as a query, never by a
credential. validate --schemas reports the same message.
Fix
- name the side, the file, the format it was read as, and the columns found when a
primary key is missing, in
run,crosswalk, andvalidate --schemas
v0.14.1 (2026-10-06)
The package is the one 0.14.0 shipped: no command, rule, connector, or report changed.
This release proves the shorter release path. A release now writes the version into
every pin a user copies, from one list in pyproject.toml: the docs, the GitHub Action,
and the bug report form. A test holds those pins to the package's version between
releases, where 0.14.0 edited each by hand.
The repository now says whom Veridelta serves. product/USERS.md holds the personas
and use cases, and product/KEY_METRICS.md the north star with its drivers and
guardrails. product/FEATURES.md maps every feature and roadmap item to the use case it
serves, and each roadmap item names its use case or says that none asks for it yet. The
GitLab CI template is frozen as it stands, until a GitLab user asks for more. The
pushdown parity suite has still not run against the live warehouses. That run comes
before 0.15.0, and a difference it finds ships as a 0.14 patch.
Chore
- bump every version pin from one list in
pyproject.toml, and hold each to the package's version with a test - write the personas, use cases, key metrics, and feature map under
product/, with a test that ties them together - freeze the GitLab CI template, and name a use case on every roadmap item
v0.14.0 (2026-10-06)
run, validate, and crosswalk now exit with 3 when they cannot finish, such as on
a configuration error, an unreachable source, or a missing extra. 1 means drift, or a
file validate finds invalid, and nothing else. Under --json, a failure prints one
object, {"error": {"type": ..., "message": ...}}, on stdout, so a script reads one
stream whatever happens. The GitHub Action and the GitLab CI template still read 1 as
drift only when the summary says so, since either can pin an older release.
Snowflake signs in with a key pair, from private_key_path and an optional
private_key_passphrase, since Snowflake now requires strong sign-in for scripted
users. -v or --verbose prints each connection, read, and pushdown statement on
stderr, with their timings and never a credential. run --otel-send posts the run's
OpenTelemetry metrics to the OTLP/HTTP endpoint that the standard OTEL_EXPORTER_OTLP_*
variables name, so no separate curl step is needed. The GitHub Action and the GitLab
CI template gain an otel-send input. The GitLab template also gains the Action's
working-directory, python-version, and upload-artifact inputs.
Three failures that the command line called a bug now name their cause before the
comparison starts. One is a missing or unreadable file. Another is a Delta Lake or
Iceberg table, version, or snapshot that cannot be read. The third is a cast_to that
the column's type cannot take, which validate --schemas reports too. A SQL Server
table read now reads each DATETIMEOFFSET at the instant it holds, where ConnectorX
0.4.6 shifted it by its offset. A query still reads what ConnectorX returns, and the
docs show the workaround.
The HTML report reads without JavaScript: every row is in the page, and the script only
splits long tables into pages of 25 rows. Each pager names its table and announces its
page, and a wide table scrolls from the keyboard. Cells show Python's text for a value,
so a float reads 1.0 and a struct shows its fields. On the docs site, text reaches the
4.5:1 contrast ratio in both color schemes, links are underlined, and the search dialog
has a name. ACCESSIBILITY.md says what Veridelta aims for, how it is checked, and how
to report a barrier.
Reads are now tested against real systems: Delta Lake and Iceberg tables each test writes, and MySQL 8.4 and SQL Server 2022 servers in CI. A suite started by hand runs the pushdown parity cases inside Snowflake, Databricks, BigQuery, and MotherDuck. This release ships before that suite's first run against those services. That run comes before 0.15.0, and a difference it finds ships as a 0.14 patch.
For AI agents, the docs site publishes llms.txt and llms-full.txt, and a new page
says how an agent should drive the command line. skills/veridelta/SKILL.md carries the
same steps as an agent skill. Contributors' coding agents read one rules file,
AGENTS.md, and settled choices keep short records under decisions/. A new CI job
checks every docs page and a sample report with axe-core.
Feat
- sign in to Snowflake with a key pair, from
private_key_pathandprivate_key_passphrase - print each connection, read, and pushdown statement on stderr with
--verbose - send a run's OpenTelemetry metrics to an OTLP/HTTP endpoint with
run --otel-send, and theotel-sendinput of the GitHub Action and the GitLab CI template - give the GitLab CI template the Action's
working-directory,python-version, andupload-artifactinputs - exit with
3when a command cannot finish, and print the failure as one JSON object under--json
Fix
- report a missing or unreadable file before the comparison starts
- report an unreadable Delta Lake or Iceberg table, version, or snapshot before the comparison starts
- refuse a
cast_tothe column's type cannot take before reading a row, and report it invalidate --schemas - read each SQL Server
DATETIMEOFFSETin atableread at the instant it holds - render every row of the HTML report in the page, name each pager's buttons after its table, announce its page, and let the keyboard scroll a wide table
- show report cells as Python writes their values, such as
1.0for a float - reach the 4.5:1 contrast ratio on the docs site in both color schemes, underline its links, and name its search dialog
- correct where the docs and the code disagreed
Test
- read real Delta Lake and Iceberg tables
- read real MySQL and SQL Server tables in CI
- run the pushdown parity suite inside live warehouses, started by hand
- check every docs page and a sample HTML report with axe-core and a keyboard in CI
Chore
- publish
llms.txt,llms-full.txt, and an AI agents page, and add AI workflows to the roadmap - add an agent skill,
AGENTS.mdfor contributors' coding agents, and decision records - add an accessibility statement, an accessibility issue form, and a funding file for GitHub Sponsors
- run the end-to-end tests on macOS, give CI jobs a read-only token, and update vulnerable locked dependencies
BREAKING CHANGE
run,validate, andcrosswalkexit with3, not1, when they cannot finish. Under--json, a failure prints{"error": {"type": ..., "message": ...}}on stdout, where it printed nothing.
v0.13.0 (2026-10-05)
A duckdb source reads a table, or the result of a query, from a DuckDB file or a
MotherDuck database written as md:name. It is compared locally and pairs with any
source but a warehouse. It needs the duckdb extra. A file opens read-only, and every
session reads time in UTC. MotherDuck needs a token, from the motherduck_token field
or the MOTHERDUCK_TOKEN variable, and is tested only against a stand-in for its
driver.
Two tables in one DuckDB database that both set pushdown are compared inside DuckDB,
as two Postgres tables already can be. Only counts, keys, and an optional row sample
come back. DuckDB refuses max_levenshtein_distance, since its levenshtein counts
bytes rather than characters.
A database table can be read in parallel: set partition_on to an integer column and
partitions to the number of ranges. The column must hold no NULL, which ConnectorX
would leave out. Veridelta counts NULLs first and fails the read if it finds any.
--markdown-max-rows lists changed values in the Markdown summary, which the GitHub
Action and the GitLab CI template post on pull and merge requests. The default, 0,
lists none, since everyone who can read the comment sees the values. OpenTelemetry
exports now read OTEL_SERVICE_NAME and OTEL_RESOURCE_ATTRIBUTES, and Veridelta's own
attributes keep their values.
The rest is internal. Single-use helpers are inlined across the engine, the compiler,
and the CLI, and private docstrings are cut to one line. Public docstrings follow the
style rules, with doctested examples, and copied tests are parametrized. CI has one
check to require on main, a time limit on each job, and a workflow that re-runs jobs
GitHub never started.
Feat
- read DuckDB files and MotherDuck databases as a
duckdbsource - compare two DuckDB tables inside DuckDB when both set
pushdown - read a database
tablein parallel partitions withpartition_onandpartitions - list changed values in the Markdown summary with
--markdown-max-rows, and themarkdown-max-rowsinput of the GitHub Action and the GitLab CI template - read
OTEL_SERVICE_NAMEandOTEL_RESOURCE_ATTRIBUTESfor OpenTelemetry exports
Refactor
- inline single-use helpers and drop repeated code in the engine
- inline single-use helpers in the SQL compiler and the connector sessions
- define
--configonce in the CLI and inline its single-use helpers
Chore
- add one
CI Passedcheck formainto require, stop each CI job after 15 minutes, and re-run CI jobs that GitHub never started - drop stale tool configuration
- mark each test module once, parametrize copied tests, and drop exact duplicates
- cut private docstrings to one line, and write the public ones to the style rules with doctested examples
- start the release steps with a
cz bumpdry run
v0.12.1 (2026-10-05)
A Postgres table now keeps the declared precision and scale of each numeric column,
with or without pushdown. Before, every numeric arrived as Decimal(38, 10),
rounded to ten decimal places. A value with more than 18 digits before the decimal point
failed the read, and strict_types could not tell numeric(10, 2) from
numeric(12, 4). Veridelta now reads the declarations from the pg_attribute catalog
before the rows, so a Postgres-compatible server without that catalog needs a query. A
query, a numeric declared without a precision, and one with a precision above 38 or
a negative scale keep Decimal(38, 10). A stored NaN fails the read, which now names
its column.
Pushdown also accepts timezone on text that datetime_format parses with an offset,
as a local run does, and cast_to: Boolean works on BigQuery numbers that are not
integers. OpenTelemetry exports keep the container in abfss:// source names.
The documentation follows written style rules, which a CI check enforces. The configuration guide is split into pages for configuration, sources, rules, pushdown, results, and the command line. A link to a section that moved lands at the top of the configuration page.
Fix
- read each Postgres
numericat its declared precision and scale, with or withoutpushdown - accept
timezonein pushdown on text thatdatetime_formatparses with an offset - compare BigQuery numbers with zero for
cast_to: Boolean, since BigQuery casts only integers and text toBOOL - keep the container in the
abfss://source names of OpenTelemetry exports
Chore
- pin workflow actions to commits, and let Dependabot keep them and the lockfile current
- deploy the documentation one commit at a time
- add a code of conduct, issue forms, and a pull request template
- hold the documentation and tutorials to written style rules, checked in CI, and split the configuration guide into focused pages
v0.12.0 (2026-10-05)
Veridelta now requires Python 3.11 or later, since Python 3.10 reaches end of life this month; pip and uv keep installing 0.11.1 on 3.10. Pushdown reports can now show values, and any run can feed a dashboard through OpenTelemetry.
Set pushdown_sample_rows on a comparison that runs in a warehouse or inside Postgres,
and one more statement fetches up to that many changed rows, lowest keys first, with
both sides' values and each column's match flag, laid out like a local run's changed
rows. The HTML report shows them in place of the bare keys, DiffResult.changed_sample
holds them, and output_path writes them as changed_rows_sample. The counts do not
change. Values do leave the warehouse, so the default, 0, fetches nothing, and the
Markdown summary that CI posts, the --json summary, and the logs never carry them.
veridelta run --otel PATH writes the run's metrics as one line of OTLP JSON for an
OpenTelemetry Collector or any OTLP/HTTP endpoint, with no new dependency: rows on each
side, added, removed, and changed rows, mismatched rows for every compared column, the
mismatch ratio, and the verdict, each a gauge. They carry counts and column names,
never row values, connection URIs, credentials, or SQL. The GitHub Action and the
GitLab CI template write the file into their artifact, and the Action exposes its path
as the otel-metrics output. A fifth tutorial takes a database source through
veridelta validate and into the CI integrations.
Feat
- fetch up to
pushdown_sample_rowschanged rows with both sides' values in a pushdown run, shown in the HTML report and written aschanged_rows_sample - write OpenTelemetry metrics with
veridelta run --otel, and from the GitHub Action and the GitLab CI template
Refactor
- make
SQLDialectaStrEnum, sostr()of a dialect is its value
Chore
- finish a release that is still waiting for approval when another change merges to
main, and upload only the files PyPI does not have yet - require Python 3.11 or later, and test 3.11 through 3.14
BREAKING CHANGE
- Veridelta requires Python 3.11 or later
v0.11.1 (2026-10-05)
Two Postgres tables can now be compared inside Postgres, and warehouse pushdown trims whitespace and reads regex replacements the way a local run does. Those two settings compared differently in a warehouse and now give the local verdict there too, so a run can report fewer differences than it did on 0.11.0.
Set pushdown: true on two type: database sources that name tables on one Postgres
connection, and the comparison compiles to SQL and runs inside Postgres through
ConnectorX, so only counts and primary keys come back. Both sides must set it; a
database source without it is still read into memory and compared locally. Postgres
refuses datetime_format and max_levenshtein_distance with ConfigError, since it
has no date parse that returns NULL and its levenshtein needs an extension, and the
server must keep standard_conforming_strings on, as it is by default. CI runs the
parity suite against a live Postgres. ConnectorX reads every Postgres numeric as
Decimal(38, 10) whether or not pushdown is set; the configuration guide explains
what that changes.
whitespace_mode strips the same characters in a warehouse as in a local run: tabs,
line breaks, no-break spaces, and the rest of Unicode's whitespace, not only spaces.
Before, Snowflake, Databricks, and DuckDB pushdown counted values padded with those
characters as changed, and keys padded with them as added and removed rows.
A regex_replace replacement now means the same in every warehouse. Write group
references as Polars reads them, $1 or ${1}, with $0 for the whole match and
$$ for a dollar sign, and pushdown rewrites them for each warehouse: \1 on
Snowflake, BigQuery, DuckDB, and Postgres, and $1 on Databricks. A reference to a
group by name, which includes $1a, or to a group above 9 raises ConfigError on a
warehouse pair, and veridelta validate warns about it. A replacement written for
Snowflake as \1 is now plain text in the warehouse, as it always was in a local run.
Feat
- compare two Postgres tables inside Postgres when both
databasesources setpushdown: true
Fix
- strip tabs, line breaks, and Unicode whitespace in pushdown
whitespace_mode, as Polars does - rewrite
regex_replacegroup references for each warehouse, and refuse references to a group by name or above 9
Chore
- tag the version, publish it to PyPI after approval, and create the GitHub release
when a version bump merges to
main
v0.11.0 (2026-10-04)
Veridelta now reads operational databases, compares BigQuery tables in place, and checks a configuration before it runs. Read the behavior changes first: a few comparisons that matched on 0.10.0 because of arithmetic or parsing bugs now report drift that is really there.
A type: database source reads a table, or the result of a query, from Postgres,
MySQL or MariaDB, SQL Server, Oracle, Redshift, ClickHouse, or SQLite through
ConnectorX, installed with the new database extra. The read is eager and the
comparison runs locally, so a database pairs with a file, a lakehouse table, or
another database, and crosswalk reads it too. A query is sent verbatim, so connect
with a read-only role. A password field is percent-encoded into the URI and kept out
of printed configs and error messages.
BigQuery joins Snowflake and Databricks for warehouse pushdown through the new
bigquery extra. It authenticates with Application Default Credentials or a key
file, and maximum_bytes_billed caps what each statement may scan. The dialect
follows the GoogleSQL reference and is tested against a stand-in client; it has not
yet run against a live project. veridelta crosswalk now also runs inside a
warehouse when both tables share one connection, counting every candidate column in
one statement. There, only columns stored as text on both sides qualify.
veridelta validate checks a configuration without reading any rows: backend
pairing, missing extras, and regex patterns Polars rejects, plus each side's columns
with --schemas. --allow-missing-env lets it run in CI without secrets. A JSON
Schema gives editors completion and validation as you type, and veridelta schema
prints it. veridelta run --markdown writes a summary for pull requests, and a
composite GitHub Action and a GitLab CI template run a comparison and keep one summary
comment up to date. Pin either to v0.11.0 or later. Avro files are read with
format: avro, from local paths and eagerly.
Several parity fixes change results. Local integer tolerances measure the true
difference, where Int8 100 against -100 used to wrap into a match. %f in
datetime_format reads a fraction of a second, so .5 is 500 ms rather than 5 ns.
strict_types: true now applies in warehouse comparisons, which used to treat 10.0
and 10 as equal. DuckDB pushdown replaces every regex match, and pushdown tolerances
widen integer operands, so narrow and extreme values no longer overflow. A Hypothesis
property test compares local and pushdown verdicts on generated data in every CI run.
veridelta[all] now includes ConnectorX, which publishes no wheels for musllinux or
Windows on ARM, so install only the extras you need there.
Feat
- read a table or query from Postgres, MySQL, SQL Server, Oracle, Redshift,
ClickHouse, or SQLite with
type: database, through the newdatabaseextra - compare two BigQuery tables in place with
type: bigquery, through the newbigqueryextra, with an optionalmaximum_bytes_billedcap - propose
value_mapentries inside the warehouse when both tables share one connection - check a configuration without reading rows with
veridelta validate, adding each side's columns with--schemas - generate a JSON Schema for configuration files, print it with
veridelta schema, and publish it with the docs for editors - write a Markdown summary with
veridelta run --markdown - run a comparison in CI with a composite GitHub Action or a GitLab CI template, each keeping one summary comment on the pull or merge request
- read Avro files with
format: avro
Fix
- widen integer pairs before applying a local tolerance, so a difference no longer wraps into a match
- read
%findatetime_formatas a fraction of a second, without Polars'ChronoFormatWarning - replace every regex match in DuckDB pushdown
- widen integer operands in pushdown tolerances, so narrow and extreme values no longer overflow
- enforce
strict_typesin warehouse comparisons - reject crosswalk thresholds that are not plain numbers, and a
min_supportthat is not anint
Refactor
- route every warehouse pair through one registry, and plan a run from the schemas
alone, so
validatechecks exactly what a run would compile
Chore
- publish to PyPI only for
vX.Y.Ztags, so a floating action tag cannot publish - compare local and pushdown verdicts on generated data with Hypothesis in CI
BREAKING CHANGE
%findatetime_formatreads.5as 500 ms instead of 5 ns- local integer tolerances use the true difference, so pairs that wrapped into a match now mismatch
strict_types: truefails warehouse columns whose types differ after normalization- crosswalk thresholds must be
intorfloatvalues other thanbool, andmin_supportmust be anint
v0.10.0 (2026-09-26)
Warehouse pushdown now reaches the same verdict as a local run in every case the differential harness covers, and several local comparisons that reported a match on values that differ are fixed. Read the behavior changes before upgrading: a run that passed on 0.9.1 can fail on 0.10.0 because the data really does differ.
Snowflake pushdown failed on every real run before this release. The driver returns
None for a zero-row result, and every run starts with a zero-row schema probe, so
the snowflake extra now requires snowflake-connector-python 3.7.0. String literals
are escaped per dialect: a backslash in a regex or sentinel reaches Snowflake and
Databricks intact, an apostrophe survives on Databricks, and a value ending in a
backslash can no longer close its literal and run as part of the statement. Primary
keys go through the same normalization in a warehouse join as in a local run, and a
key that repeats after normalization raises DataIntegrityError in the warehouse too,
so added, removed, and changed counts can change.
Numeric comparisons no longer cast the target to the source's type. With
strict_types off, an integer 10 and a float 10.7 now differ, as they already did
in a warehouse, and a Float32 0.1 no longer equals a Float64 0.1; add a
tolerance or a cast_to to forgive precision gaps. Under a tolerance, NaN matches only
NaN and an infinity only itself, on both paths, and unsigned differences no longer
wrap.
Strings in the source and target blocks expand ${NAME} and ${NAME:-default}
from the environment, so credentials can stay out of the file. A literal ${ must now
be written $${. Printed configs leave out password, access_token, and
storage_options, and validation errors no longer quote the input they reject;
model_dump() still returns every field.
Two additions: max_levenshtein_distance and min_jaro_winkler_similarity forgive
typos in text columns through the new fuzzy extra, and veridelta crosswalk proposes
value_map entries from how the source and target values line up, with the evidence
for each. Jaro-Winkler runs locally only; pushdown refuses it before any comparison
query runs.
The tutorials are now one numbered path of four notebooks, 01_core_concepts through
04_html_reports, so the old notebook URLs return 404. The README and the site index
were rewritten around them.
Feat
- expand
${NAME}and${NAME:-default}insourceandtargetstrings, with$${for a literal${, and name the variable and its location when one is unset - forgive text typos within
max_levenshtein_distanceormin_jaro_winkler_similarity, scored locally through the newfuzzyextra. Pushdown compiles the edit distance toEDITDISTANCEorlevenshteinand refuses Jaro-Winkler before any comparison query runs - propose
value_mapentries withveridelta crosswalkandDiffEngine.propose_value_maps(), printing paste-ready rules on stdout and the evidence for each entry on stderr - normalize primary keys through stages 1-7 in every warehouse join, accept renamed
keys, and reject keys that repeat after normalization with the same
DataIntegrityErrora local run raises - give connectors
close()and context-manager support, log connections and statements without SQL or credentials, and tell a missing lakehouse extra apart from a failed scan
Fix
- pass
force_return_table=Trueto Snowflake's Arrow fetch and requiresnowflake-connector-python>=3.7.0, so a zero-row result no longer fails the run - escape string literals per dialect, so backslashes and apostrophes survive and a trailing backslash cannot escape its literal
- compare mixed numeric types in their common type instead of casting the target to the source's type
- stop NaN and infinity from matching any value under a tolerance, and measure unsigned differences without wrapping
- apply tolerances in pushdown only to columns a local run compares as numbers, and resolve one rule per column by its post-rename name on both paths
- count a row with one NULL side in pushdown's
changed_count, and select no rows when every compared column is ignored - count rows with null keys in the row totals, as the warehouse's
COUNT(*)does - drop pattern-ignored columns from the target side too, so
schema_mode: exactno longer fails on them - reject an empty
primary_keys, an infinite tolerance, the same table on both sides of a pushdown, headers that collide after normalization, and a negative--html-max-rows - keep the HTML report readable when a value is NaN or infinite, and show integers beyond 2^53 exactly
- leave
password,access_token, andstorage_options, including a file source's nested one, out of printed configs, and keep inputs out of validation errors
Refactor
- share one rule folder and one rename and drop matcher between the two engines, and
split
DiffEngine.runinto named stages - raise
DatasetError, aVerideltaError, fromload_nyc_taxiinstead ofRuntimeError
Chore
- run CI on stacked pull requests and execute the tutorial notebooks against their recorded output
- enable Ruff's McCabe complexity check at 10 and tighten the Cursor rules
BREAKING CHANGE
- a literal
${insidesourceortargetmust be written$${ - mixed numeric types compare by value, so integer and float columns that differ only in the fraction now mismatch locally
load_nyc_taxiraisesDatasetError, whichexcept RuntimeErrorno longer catches- the tutorial notebooks were renamed, and their old URLs return 404
v0.9.1 (2026-09-11)
CI now pins JavaScript actions that declare Node 24, so GitHub-hosted runners stop forcing the deprecated Node 20 runtime onto checkout, setup-uv, and codecov.
Chore
- bump
actions/checkoutto v5,astral-sh/setup-uvto v7, andcodecov/codecov-actionto v6 so every workflow step we pin runs on Node 24
v0.9.0 (2026-09-11)
Internal. The public API does not change. This is the 1.0 candidate: the breaking-change budget for the 0.7–0.9 line was spent in 0.7.0, 0.8.0 was additive, and 0.9.0 is compiler and toolchain work.
Warehouse SQL now projects stages 1–7 once per column through a pair of normalization CTEs. The join and the mismatch tally read those values, so adding a transform no longer copies the expression tree into every later predicate. Emitted SQL stays roughly linear in the number of compared columns; the DuckDB harness is the same check it was in 0.6.0.
engine.py, models.py, sentinels.py, and connectors/sql.py are gated
at 100% branch coverage. cast_to is resolved through a TypedDict and an
exhaustive lookup table. make lint runs strict pyright. Global mypy
ignore_missing_imports is gone; missing stubs stay on per-module overrides
for optional drivers and test-only imports.
Refactor
- project warehouse transforms through
_src_normalized/_tgt_normalizedCTEs so each column is normalized once, verified by the existing DuckDB harness and a SQL-length regression test - convert
_get_effective_ruleto anEffectiveRuleTypedDict socast_tostays aCastTargetand_CAST_TARGETSis exhaustive
Test
- fail
make testand the core CI job unless the four core modules stay at 100% branch coverage, and drop the deadraise NotImplementedErrorcoverage exclude
Chore
- add strict pyright to
make lintand the CI lint job - drop the global mypy
ignore_missing_importsin favor of the existing per-module overrides
v0.8.0 (2026-09-11)
Reporting and CLI polish. Purely additive: new flags default off, and the exit codes stay 0 for a match and 1 for drift.
--html PATH writes one self-contained file. Styles and a 25-row pager are
inlined, so an air-gapped runner can open the report without fetching anything.
Embedded rows are capped (--html-max-rows, default 1000) and the truncation is
stated. Pushdown reports are labeled as primary-keys-only. Markup in column
names is escaped, and < inside the embedded JSON is rewritten so a value
cannot close the script block.
Progress chatter moved to stderr so veridelta run --json | jq does not have
to strip it. --quiet silences that chatter. --version / -V print the
package version.
docs/configuration.md is now bound to DiffConfig, DiffRule, and
SourceConfig by a test, so a new field cannot ship undocumented. A Getting
Started notebook covers local diffing and the DiffResult accessors added in
0.7.0.
Feat
- write a standalone HTML report from the Python API or
--html, with an inlined pager and a row cap that is visible in the page - add
--jsonso CI can readDiffSummaryfrom stdout,--quietto suppress progress, and--version/-V - route every progress line to stderr, including the artifact path, so a piped
--jsonrun stays valid JSON - add a Getting Started notebook and bind
docs/configuration.mdto the three config models so missing fields fail the test suite
Fix
- explain configuration errors as a problem with the file, not the data, and point unexpected failures at the issue tracker with the exception class
v0.7.0 (2026-09-10)
Two breaking changes, both mechanical.
DiffEngine.run() and DiffEngine.run_from_configs() now return a DiffResult rather
than a DiffSummary. Existing code reaches the old value through .summary:
SourceType and output_format are now closed sets. SourceType accepts csv,
parquet, json, ndjson, arrow, and excel; output_format accepts everything
but excel. The removed names (fixed_width, netcdf, shapefile, geopackage,
sql, avro, xml, delta) never had an implementation and already raised
ConfigError at run time, so a config using one was already broken; it now fails when
the config loads instead. Delta Lake is unaffected and still reached through the
delta_lake source type.
Feat
- return
DiffResultfrom the engine, carrying the added, removed, and changed frames alongside the summary. The rows were already materialized, counted, and discarded, so reaching them previously meant configuringoutput_pathand reading files back - add
DiffResult.get_mismatches(column)to narrow the changed set to one column with the source and target values side by side, andDiffResult.to_pandas()for notebooks - record the compared column set on
DiffResult, so a mistyped column name is rejected on the pushdown path too, where the primary-key frames cannot reveal which columns were compared - read JSON, NDJSON, Arrow IPC, and Excel. NDJSON and Arrow scan lazily; JSON and Excel are read whole, because Polars has no lazy reader for either
- write discrepancy artifacts as JSON, NDJSON, or Arrow IPC in addition to CSV and Parquet
- add an
excelextra, reporting a missing install as a hint rather than anImportErrorraised from inside Polars - export the exception hierarchy, the warehouse and lakehouse configs, and the public
type aliases from the package root. The docs already instructed users to catch
VerideltaErrorand to construct aSnowflakeConfig, neither of which was reachable without importing from a submodule
Fix
- ship the
py.typedmarker. The package advertised theTyping :: Typedclassifier without it, so PEP 561 had every downstream mypy and pyright treat Veridelta as unannotated no matter how complete its annotations were - bind
SourceTypeandArtifactFormatto the registries that implement them, with tests. A name could previously be advertised in the config schema with nothing behind it
BREAKING CHANGE
DiffEngine.run()andDiffEngine.run_from_configs()returnDiffResultSourceTypeandoutput_formatare constrained to implemented formats
v0.6.0 (2026-09-10)
Warehouse pushdown now implements all nine transform stages. pad_zeros,
datetime_format, timezone, and cast_to previously raised ConnectorError on the
warehouse path, so any rule using them could only run locally.
cast_to is now a closed set of Int64, Float64, String, Boolean, Date, and
Datetime. Any other value is rejected at load time rather than silently skipping the
cast, so a config with a typo that previously reported a clean diff will now fail
validation. Configs already using valid Polars type names are unaffected.
Feat
- compile
pad_zerosas a sign-aware, non-truncating expression rather than anLPAD, which pads in front of a minus sign (0-12where Python'szfillgives-012) and discards characters past the target width - compile
cast_tothrough a per-dialect keyword table, truncating toward zero on a float-to-integer cast to match Polars where Snowflake and DuckDB round - compile
datetime_formatby translating one directive at a time against a per-dialect table, with literal runs restricted to a separator allowlist and wrapped in the dialect's quoting. An untranslatable directive raisesConfigErrorrather than passing through, since an unrecognized directive parses nothing and returns NULL for every row - enforce the
timezoneprecondition against the probed warehouse schema, so a naive or non-temporal column fails the same way it does locally. The conversion itself emits no SQL: Polars rewrites only a column's timezone label and every downstream cast still reads the UTC instant, while warehouses have no per-column label to rewrite - add a differential test harness that runs both engines over the same frames and compares the results, executing real compiler output through DuckDB
Fix
- skip the text transform stages on non-text columns during pushdown, matching the local
engine. A currency string compared against a native float previously asked the
warehouse to run
REGEXP_REPLACEover a number
BREAKING CHANGE
cast_tois constrained to aLiteralof supported Polars type names. It previously accepted any string and resolved it withgetattr, so an unrecognized name left the column uncast and the diff green. The field also reaches SQL asCAST(x AS <type>), where a type name cannot be quoted or bound as a parameter
v0.5.1 (2026-09-10)
Unsupported formats now raise ConfigError instead of NotImplementedError. Code
catching NotImplementedError around LoaderFactory or a diff run should catch
ConfigError, or VerideltaError for all framework failures.
Fix
- raise
ConfigErrorrather thanNotImplementedErrorfor a source format with no loader or anoutput_formatwith no writer, so the CLI reports "Configuration Error" instead of "Unexpected System Error" for an ordinary misconfiguration - validate the artifact format before writing rather than inside the write loop. The
check previously ran after the empty-frame guard, so a bad
output_formatwas caught only when there was drift to write and a clean comparison passed silently - name the supported formats in both messages, derived from the registry that backs the behavior so the message cannot drift from what actually works
v0.5.0 (2026-09-09)
Null sentinels are no longer strings only. null_values and default_null_values
accept mixed scalars, and quoting now carries meaning: -999 nulls out -999 in a
numeric column and is ignored on a text column, while "-999" behaves the other way
around. A list that previously read ["N/A", "-999"] still works unchanged; add the
unquoted form if the value also appears as a number.
Feat
- widen
null_valuesanddefault_null_valuesfromlist[str]to a mixed union ofstr | int | float | bool, withstrict=Trueso Pydantic preserves the exact type each sentinel was written as - filter sentinels against each column's dtype before use, in both the local engine and the warehouse compiler, so one global list can span a mixed schema. Text sentinels reach string, categorical, and enum columns; numbers reach any numeric column including decimals; booleans reach boolean columns only
- carry probed dtypes out of the warehouse schema probe into both pushdown query builders, filtering each side independently since the two relations can disagree
- raise
ConfigErrorwhen an explicit per-columnnull_valuesrule holds no sentinel its column's type can match, locally and against the probed warehouse schema. A globaldefault_null_valuesstill skips silently, since spanning a mixed schema is its purpose - reject
.nanand.infsentinels at load time, as NaN never compares equal to itself and infinity has no portable SQL literal
Refactor
- emit one
CASE WHEN col IN (...) THEN NULL ELSE col ENDper side instead of nestedNULLIFcalls, rendering numbers and booleans unquoted so they cannot break a numeric column's cast. Strings still route through_literalfor apostrophe escaping
v0.4.0 (2026-09-09)
Transforms now apply uniformly to every column, primary keys included, before the
joins. A case_insensitive, whitespace_mode, or value_map rule on a key column
therefore changes how rows are matched, not only how they are compared. Key
uniqueness is asserted after normalization, so a rule that collapses two keys into
one raises DataIntegrityError instead of exploding the join.
Feat
- consolidate every column transform into one normalization pass per dataset, executed
before the joins, and make the
DiffRuledocstring the canonical transform order that both the local engine and the SQL compiler follow - implement
pad_zeros, which stringifies first so a numeric123matches a text"00123" - implement
datetime_format, which parses text into timestamps so the column is compared as a timestamp rather than as text - implement
timezone, which converts timezone-aware data only and raisesConfigErrorfor naive timestamps rather than assuming an origin zone and shifting every value - compare every shared column on the warehouse path, with global
default_*settings folded in, instead of only columns carrying an explicit rule - populate
column_mismatcheson warehouse pushdown from a per-columnSUM(CASE ...)tally, usingCOALESCE(pred, FALSE)to match the local engine under SQL three-valued logic - write warehouse pushdown artifacts through the shared exporter under
_pks_onlyfilenames, since those queries project primary keys only - reject a
pad_zeroswidth supplied as a string or float instead of coercing it
Fix
- stop applying
regex_replacetwice, once in therun()pre-pass and again during comparison, which corrupted any non-idempotent pattern - stop a global
default_null_valuesfrom failing every run containing numeric columns, by gating text transforms on string dtypes - reject an unknown
timezonename as aConfigErrornaming the column, rather than surfacing a raw Polars error
v0.3.0 (2026-09-08)
Feat
- warehouse and lakehouse connectors with SQL pushdown (#5)
- add Snowflake and Databricks warehouse extras with Arrow SQL pushdown
- add Delta Lake and Iceberg lakehouse scans, including Iceberg
snapshot_idtime travel - route YAML
source/targetthrough a discriminatedSourceRefunion andDiffEngine.run_from_configs - compile warehouse anti-joins and row counts so pushdown reports added, removed, and changed counts against exact source and target totals
- enforce
schema_modeand primary-key existence on warehouse relations via zero-row column probes - allowlist dotted warehouse table identifiers and reject string-coerced numeric config fields
- report artifact persistence through
DiffSummary.artifacts_written - add
make docsandmake docs-serve, and build the documentation strictly in CI
Fix
- stop the CLI from announcing discrepancy artifacts that warehouse pushdown never writes
- keep the Snowflake extra inside the driver's supported pyarrow range on Python 3.14
- update Python classifiers and correct homepage URL in pyproject.toml
v0.2.0 (2026-04-30)
Feat
- core: migrate to Polars Lazy computation graphs and implement enterprise CI/CD (#4)
- implement core semantic diffing engine and GitOps configuration (v0.1.0-alpha) (#3)
- add dynamic output_format for diff artifacts with strict fallback
- add RELAXED_ORDER schema validation mode
- implement core diff engine, schema validation, and documentation (#2)
Fix
- update dataset URL to use versioning from package metadata
- satisfy CI requirements and trigger docs deployment
- add missing tests directory and modernize uv config