2. YAML and CLI¶
A configuration file declares a comparison in YAML: its sources, primary keys, and rules. veridelta run reads the file, compares the two sides, and exits 0, 1, or 3, which makes a comparison one step in a CI job.
This tutorial writes two Parquet files and a configuration, runs the comparison from the command line and from Python, and reads its artifacts. The core concepts tutorial covers the Python models, and Command line lists every flag.
Open this tutorial in Google Colab. Its first cell installs Veridelta there.
# A kernel without Veridelta, such as Colab's, installs the release this tutorial shows.
import importlib.util
import subprocess
import sys
if importlib.util.find_spec("veridelta") is None:
subprocess.run(
[sys.executable, "-m", "pip", "install", "--quiet", "veridelta==0.35.1"],
check=True,
)
1. Source and target files¶
Three accounts appear on both sides, and one appears only in the target. The target abbreviates status codes, and one balance drifted by 99 cents:
import polars as pl
pl.DataFrame(
{
"user_id": [1, 2, 3],
"status": ["Active", "Pending", "Closed"],
"balance": ["$100.50", "$50.00", "$0.00"],
}
).write_parquet("source.parquet")
pl.DataFrame(
{
"user_id": [1, 2, 3, 4],
"status": ["ACT", "PND", "CLS", "ACT"],
"balance": [100.50, 50.99, 0.00, 12.00],
}
).write_parquet("target.parquet")
print("wrote source.parquet and target.parquet")
# Output:
# wrote source.parquet and target.parquet
2. Write the configuration¶
The file names both sources, the primary key, and two rules. Each path's suffix says what format it is. output_path is the directory for artifacts: the added, removed, and changed rows, each written only when it has rows:
%%writefile veridelta.yaml
source:
path: "source.parquet"
target:
path: "target.parquet"
primary_keys: ["user_id"]
output_path: "./diff_results"
output_format: "parquet"
rules:
- column_names: ["status"]
value_map:
Active: ACT
Pending: PND
Closed: CLS
- column_names: ["balance"]
regex_replace:
"\\$": ""
cast_to: Float64
# Output:
# Writing veridelta.yaml
3. Run the comparison¶
Exit code 0 means the comparison fell within threshold. Exit code 1 means drift, and exit code 3 a run that could not finish. --json prints the DiffSummary on stdout, and --quiet keeps progress off stderr, so veridelta run --json --quiet | jq reads clean JSON:
!veridelta run -c veridelta.yaml --quiet
# Output:
#
# Veridelta Execution Summary
# ===========================
# Status: FAILED
# Match Rate: 33.33%
# Source Rows: 3
# Target Rows: 4
# Volume Shift: +1 row
#
# Row-Level Discrepancies:
# ---------------------------
# Added: 1
# Removed: 0
# Changed: 1
# Total Issues: 2
#
# Top Column-Level Drifts:
# ---------------------------
# - balance: 1 mismatch
!veridelta run -c veridelta.yaml --json --quiet; echo "exit=$?"
# Output:
# {
# "total_rows_source": 3,
# "total_rows_target": 4,
# "added_count": 1,
# "removed_count": 0,
# "changed_count": 1,
# "column_mismatches": {
# "balance": 1
# },
# "is_match": false,
# "accepted_count": 0,
# "total_mismatches": 2,
# "mismatch_ratio": 0.6666666666666666,
# "match_rate_percentage": 33.33,
# "is_perfect_match": false,
# "volume_shift": 1,
# "report_summary": "Veridelta Execution Summary\n===========================\nStatus: FAILED\nMatch Rate: 33.33%\nSource Rows: 3\nTarget Rows: 4\nVolume Shift: +1 row\n\nRow-Level Discrepancies:\n---------------------------\nAdded: 1\nRemoved: 0\nChanged: 1\nTotal Issues: 2\n\nTop Column-Level Drifts:\n---------------------------\n- balance: 1 mismatch\n"
# }
# exit=1
4. The same file from Python¶
load_config returns the parsed models. DiffEngine.run_from_configs loads both sources and compares them, with no LazyFrame built by hand:
from veridelta import DiffEngine, load_config
diff_cfg, src_cfg, tgt_cfg = load_config("veridelta.yaml")
summary = DiffEngine.run_from_configs(diff_cfg, src_cfg, tgt_cfg).summary
print(f"is_match={summary.is_match} changed={summary.changed_count} added={summary.added_count}")
# Output:
# is_match=False changed=1 added=1
5. Environment variables¶
One reviewed configuration can serve every environment, with each environment saying where its data lives. Any string inside source or target can read an environment variable: ${NAME}, or ${NAME:-default} for a fallback when the variable is unset or empty. A warehouse password stays out of the file the same way. Rules and root settings are read as written; see Environment variables.
From Python, model_copy(update={"path": ...}) overrides a field of a loaded configuration. This file reads its data directory from VERIDELTA_DATA_ROOT:
%%writefile veridelta.env.yaml
source:
path: "${VERIDELTA_DATA_ROOT:-.}/source.parquet"
target:
path: "${VERIDELTA_DATA_ROOT:-.}/target.parquet"
primary_keys: ["user_id"]
rules:
- column_names: ["status"]
value_map:
Active: ACT
Pending: PND
Closed: CLS
- column_names: ["balance"]
regex_replace:
"\\$": ""
cast_to: Float64
# Output:
# Writing veridelta.env.yaml
import os
import shutil
from pathlib import Path
os.environ.pop("VERIDELTA_DATA_ROOT", None)
_, src_cfg, _ = load_config("veridelta.env.yaml")
print(f"unset: {src_cfg.path}")
Path("data").mkdir(exist_ok=True)
for name in ("source.parquet", "target.parquet"):
shutil.copy(name, Path("data") / name)
os.environ["VERIDELTA_DATA_ROOT"] = "data"
diff_cfg, src_cfg, tgt_cfg = load_config("veridelta.env.yaml")
summary = DiffEngine.run_from_configs(diff_cfg, src_cfg, tgt_cfg).summary
print(f"set: {src_cfg.path}")
print(f"is_match={summary.is_match}")
# Output:
# unset: ./source.parquet
# set: data/source.parquet
# is_match=False
6. Artifacts¶
Each frame with rows is written under output_path. This run added and changed rows, but removed none:
from pathlib import Path
print(sorted(path.name for path in Path("diff_results").iterdir()))
# Output:
# ['added_rows.parquet', 'changed_rows.parquet']
The advanced rules tutorial reconciles a real dataset, and the HTML reports tutorial turns a result into a report to share. The cell below removes the files this tutorial wrote:
import shutil
from pathlib import Path
for name in ("source.parquet", "target.parquet", "veridelta.yaml", "veridelta.env.yaml"):
Path(name).unlink(missing_ok=True)
shutil.rmtree("diff_results", ignore_errors=True)
shutil.rmtree("data", ignore_errors=True)
print("cleaned")
# Output:
# cleaned