4. HTML reports¶
An HTML report presents a DiffResult as one file: the verdict, the row counts, the column-level drift, and the differing rows, in pages. It has no external references, so it opens offline from a CI artifact, an email attachment, or a shared drive.
This tutorial builds a comparison worth reporting, writes the report with veridelta.report, and walks through how to read it. The command line writes the same document with --html; see Running a comparison.
Open this tutorial in Google Colab. Its first cell installs Veridelta there.
# A kernel without Veridelta, such as Colab's, installs the release this tutorial shows.
import importlib.util
import subprocess
import sys
if importlib.util.find_spec("veridelta") is None:
subprocess.run(
[sys.executable, "-m", "pip", "install", "--quiet", "veridelta==0.35.1"],
check=True,
)
1. A comparison with something to report¶
Two billing exports disagree in every way a report distinguishes. One invoice exists only in the target, one only in the source, and two shared invoices changed in different columns. A tolerance rule forgives a one-cent rounding difference, so the report shows a rule at work:
import polars as pl
from veridelta import DiffConfig, DiffEngine, DiffRule
from veridelta.report import render_html, write_html
# Source: last night's billing export
source = pl.DataFrame(
{
"invoice_id": ["INV-001", "INV-002", "INV-003", "INV-004", "INV-005"],
"amount": [100.00, 250.50, 75.25, 310.00, 42.00],
"status": ["paid", "open", "paid", "void", "open"],
}
)
# Target: the replacement pipeline. INV-005 is gone, INV-006 is new, INV-002 moved
# by a cent, INV-003 by a dollar, and INV-004 changed its status casing.
target = pl.DataFrame(
{
"invoice_id": ["INV-001", "INV-002", "INV-003", "INV-004", "INV-006"],
"amount": [100.00, 250.51, 76.25, 310.00, 18.00],
"status": ["paid", "open", "paid", "VOID", "open"],
}
)
config = DiffConfig(
primary_keys=["invoice_id"],
rules=[DiffRule(column_names=["amount"], absolute_tolerance=0.01)],
)
# The engine consumes LazyFrames, so in-memory DataFrames are wrapped with .lazy()
result = DiffEngine(config, source.lazy(), target.lazy()).run()
print(result.summary.report_summary)
# Output:
# Veridelta Execution Summary
# ===========================
# Status: FAILED
# Match Rate: 20.0%
# Source Rows: 5
# Target Rows: 5
# Volume Shift: +0 rows
#
# Row-Level Discrepancies:
# ---------------------------
# Added: 1
# Removed: 1
# Changed: 2
# Total Issues: 4
#
# Top Column-Level Drifts:
# ---------------------------
# - amount: 1 mismatch
# - status: 1 mismatch
2. Write the report¶
write_html takes the result and a destination, creates any missing parent directories, and returns the path it wrote. Styles and the small paging script are embedded, so the file is complete on its own:
report_path = write_html(result, "reports/nightly.html")
print(report_path)
print(f"{report_path.stat().st_size / 1024:.0f} KiB, no external references")
# Output:
# reports/nightly.html
# 6 KiB, no external references
3. Read the report¶
Open reports/nightly.html in a browser. From top to bottom, it shows:
- Verdict badge:
PASSEDorFAILED, the sameis_matchdecision the command line turns into an exit code, judged againstthreshold. The time the report was generated sits beside it. - Six cards: the match rate, the source and target row counts, and the added, removed, and changed counts. They are the
DiffSummaryfields, so they agree withreport_summaryabove. - Column-level drift: one row per compared column that mismatched at least once, ranked by mismatch count. Columns that matched everywhere are left out.
- Changed rows, Added rows, and Removed rows: the differing rows, in pages. Added and removed rows are complete records. A changed row holds both values side by side, as
<column>_sourceand<column>_target, plus a<column>_is_matchflag per compared column. Afalseflag names the column that failed.
The changed rows table shows result.changed, whose flags are:
flags = result.changed.select("invoice_id", "amount_is_match", "status_is_match")
print(flags.sort("invoice_id"))
# Output:
# shape: (2, 3)
# ┌────────────┬─────────────────┬─────────────────┐
# │ invoice_id ┆ amount_is_match ┆ status_is_match │
# │ --- ┆ --- ┆ --- │
# │ str ┆ bool ┆ bool │
# ╞════════════╪═════════════════╪═════════════════╡
# │ INV-003 ┆ false ┆ true │
# │ INV-004 ┆ true ┆ false │
# └────────────┴─────────────────┴─────────────────┘
INV-002 is absent: its one-cent drift is within the absolute_tolerance, so it counts as a match. INV-003 failed on amount and INV-004 on status, as the two rows of the drift table say.
4. Capping large tables¶
A comparison of ten million rows should not produce a ten-million-row HTML file. Each table holds at most max_rows rows, 1000 by default, and says when it is truncated, so a reader never mistakes the visible rows for all of them. render_html returns the document as a string, which suits a check like this one, or attaching a report in CI without writing a file:
capped = render_html(result, max_rows=1)
note_start = capped.index("Showing the first")
print(capped[note_start : capped.index("</p>", note_start)])
# Output:
# Showing the first 1 of 2 rows. Export artifacts with <code>output_path</code> for the complete set.
Only the changed rows table was truncated: the added and removed tables held one row each, within the cap. When a report is truncated, set output_path in the configuration to export the complete added, removed, and changed frames alongside it.
A report from a warehouse comparison is marked as keys only, because pushdown reads back counts and keys, not values. Its tables list keys, unless pushdown_sample_rows fetches some values; see Row samples. From the command line, veridelta run -c veridelta.yaml --html report.html --html-max-rows 1000 writes the same document.
The validate and CI tutorial checks a configuration before it runs, then holds a pull request to the verdict. The cell below removes the report this tutorial wrote:
import shutil
# Remove the report directory (housekeeping)
shutil.rmtree("reports", ignore_errors=True)