6. Model evaluation runs¶
A team that changes how its model is served, such as a new inference server, runs its evaluation suite again. The question is whether any answer changed beyond noise. Each run is a table with one row per example, so Veridelta compares the two on the example ID. No model is called. The rules you declare say which differences are noise, and every other difference is reported.
This tutorial builds two runs of 200 examples in the notebook, so it reads no data and needs no network. max_levenshtein_distance needs the fuzzy extra: pip install "veridelta[fuzzy]".
Open this tutorial in Google Colab. Its first cell installs Veridelta there.
# A kernel without Veridelta, such as Colab's, installs the release this tutorial shows.
import importlib.util
import subprocess
import sys
if importlib.util.find_spec("veridelta") is None:
subprocess.run(
[sys.executable, "-m", "pip", "install", "--quiet", "veridelta[fuzzy]==0.35.1"],
check=True,
)
1. Two runs¶
The baseline is the evaluation before the change, and the candidate is the same 200 examples after it. Each example has an output, a grader's score from 0 to 1, and its latency_ms. The candidate differs in four ways the team expects, and in one it does not:
- every latency is different;
- most scores move by up to 0.004, the grader's own noise;
- one answer in ten drops its final period;
- a few answers come back in capitals, padded with spaces;
- examples 41, 97, and 158 now refuse to answer, and their scores fall.
import polars as pl
from veridelta import DiffConfig, DiffEngine, DiffRule
from veridelta.models import DiffResult
ANSWERS = [
"Paris is the capital of France.",
"Water boils at 100 degrees Celsius at sea level.",
"The Pacific is the largest ocean on Earth.",
"Light from the Sun reaches Earth in about eight minutes.",
"A triangle's angles add up to 180 degrees.",
"Mount Everest is the highest mountain above sea level.",
"The human heart has four chambers.",
"Photosynthesis turns light into chemical energy.",
]
ids = range(1, 201)
baseline = pl.DataFrame(
{
"example_id": list(ids),
"output": [ANSWERS[i % len(ANSWERS)] for i in ids],
"score": [round(0.70 + (i * 37 % 30) / 100, 3) for i in ids],
"latency_ms": [120 + i * 7 % 90 for i in ids],
}
)
def candidate_output(i: int, output: str) -> str:
"""Return what the candidate answered for example `i`."""
if i in (41, 97, 158):
return "I am not able to answer that question."
if i % 25 == 7:
return " " + output.upper() + " "
if i % 10 == 3:
return output.rstrip(".")
return output
candidate = pl.DataFrame(
{
"example_id": list(ids),
"output": [
candidate_output(i, out) for i, out in zip(ids, baseline["output"], strict=True)
],
"score": [
0.12 if i in (41, 97, 158) else round(score + ((i * 13) % 9 - 4) / 1000, 3)
for i, score in zip(ids, baseline["score"], strict=True)
],
"latency_ms": [
latency - 30 + i % 11 for i, latency in zip(ids, baseline["latency_ms"], strict=True)
],
}
)
print(candidate.filter(pl.col("example_id").is_in([3, 7, 41])))
# Output:
# shape: (3, 4)
# ┌────────────┬─────────────────────────────────┬───────┬────────────┐
# │ example_id ┆ output ┆ score ┆ latency_ms │
# │ --- ┆ --- ┆ --- ┆ --- │
# │ i64 ┆ str ┆ f64 ┆ i64 │
# ╞════════════╪═════════════════════════════════╪═══════╪════════════╡
# │ 3 ┆ Light from the Sun reaches Ear… ┆ 0.909 ┆ 114 │
# │ 7 ┆ PHOTOSYNTHESIS TURNS LIGHT I… ┆ 0.887 ┆ 146 │
# │ 41 ┆ I am not able to answer that q… ┆ 0.12 ┆ 115 │
# └────────────┴─────────────────────────────────┴───────┴────────────┘
2. Without rules¶
Compared value by value, every example differs, since every latency does:
def verdict(label: str, result: DiffResult) -> None:
"""Print the verdict, the changed count, and the drift per column."""
summary = result.summary
drift = ", ".join(f"{column}={count}" for column, count in summary.column_mismatches.items())
status = "PASSED" if summary.is_match else "FAILED"
print(f"{label}: {status} changed={summary.changed_count} drift=[{drift}]")
plain = DiffEngine(DiffConfig(primary_keys=["example_id"]), baseline.lazy(), candidate.lazy()).run()
verdict("1 no rules", plain)
# Output:
# 1 no rules: FAILED changed=200 drift=[output=31, score=177, latency_ms=200]
3. Rules for the noise¶
Three rules say what the team counts as noise. Latency is not part of an answer, so ignore leaves it out. A score may move by the grader's noise, so absolute_tolerance allows 0.005. An answer may differ in case, in the spaces around it, and in one character, such as a dropped period, which case_insensitive, whitespace_mode, and max_levenshtein_distance allow. Three examples remain:
rules = [
DiffRule(column_names=["latency_ms"], ignore=True),
DiffRule(column_names=["score"], absolute_tolerance=0.005),
DiffRule(
column_names=["output"],
case_insensitive=True,
whitespace_mode="both",
max_levenshtein_distance=1,
),
]
config = DiffConfig(primary_keys=["example_id"], rules=rules)
result = DiffEngine(config, baseline.lazy(), candidate.lazy()).run()
verdict("2 noise rules", result)
# Output:
# 2 noise rules: FAILED changed=3 drift=[output=3, score=3]
4. The regressions¶
get_mismatches lists the examples behind a column's count, with both values side by side. Text is shown as the rules leave it, in lower case and trimmed:
print(result.get_mismatches("score"))
print(result.get_mismatches("output").select("example_id", "output_target").row(0))
# Output:
# shape: (3, 3)
# ┌────────────┬──────────────┬──────────────┐
# │ example_id ┆ score_source ┆ score_target │
# │ --- ┆ --- ┆ --- │
# │ i64 ┆ f64 ┆ f64 │
# ╞════════════╪══════════════╪══════════════╡
# │ 41 ┆ 0.87 ┆ 0.12 │
# │ 97 ┆ 0.89 ┆ 0.12 │
# │ 158 ┆ 0.96 ┆ 0.12 │
# └────────────┴──────────────┴──────────────┘
# (41, 'i am not able to answer that question.')
5. The verdict in CI¶
The run fails, and veridelta run exits 1 on it, so CI holds the change until someone reads the three examples. A threshold of 0.02 would pass this run, since 3 of 200 examples is 1.5%, and the regressions with it. The rules forgive only the differences they name, so what they leave is the finding.
In CI, the two runs are files, and veridelta run reads the same comparison from a configuration file. This cell writes both runs to Parquet, writes the configuration beside them, and runs the file:
import yaml
from veridelta import load_config
baseline.write_parquet("baseline.parquet")
candidate.write_parquet("candidate.parquet")
settings = config.model_dump(mode="json", exclude_none=True, exclude_defaults=True)
with open("veridelta.yaml", "w", encoding="utf-8") as file:
yaml.safe_dump(
{
"source": {"path": "baseline.parquet"},
"target": {"path": "candidate.parquet"},
**settings,
},
file,
sort_keys=False,
)
with open("veridelta.yaml", encoding="utf-8") as file:
print(file.read())
verdict("3 from the file", DiffEngine.run_from_configs(*load_config("veridelta.yaml")))
# Output:
# source:
# path: baseline.parquet
# target:
# path: candidate.parquet
# primary_keys:
# - example_id
# rules:
# - column_names:
# - latency_ms
# ignore: true
# - column_names:
# - score
# absolute_tolerance: 0.005
# - column_names:
# - output
# max_levenshtein_distance: 1
# case_insensitive: true
# whitespace_mode: both
#
# 3 from the file: FAILED changed=3 drift=[output=3, score=3]
CI integrations runs this file on every pull request, and Rules describes each rule it uses.