The token-accounting conformance suite has moved to lizhuojunx86/token-accounting-conformance, with its history. It was a different subject with a different audience, and burying it in a repository about look-ahead bias helped nobody looking for either. The old paths still resolve here, as stubs pointing there.
A routing audit for your team, fixed price. The routing-audit method in this repo (
traceguard.routing_audit, written up in the dev.to series) runs on any Claude Code trace store. Send me 30 days of your team's usage records (model, tokens, timestamps, agent and session ids; I'll send a one-linejqfilter that drops prompt and answer text before anything leaves your machine) and within ten working days you get one number and a five-page report: what share of your spend ran on a tier your own routing policy would not have chosen, what that cost at list price, and which components caused it. On my own 26,131 traces the number was 22.6% and $1,248.13, all of it on subagents and none on the main thread. US$1,500, flat. Write to info@zhuojun.li with the subject "routing audit".
Point-in-time correct LLM instrumentation — the time-integrity layer for LLM pipelines.
TraceGuard makes it structurally impossible for a run over historical data to use a model, prompt, or feature that did not exist yet. Tracing, version pinning, and look-ahead-bias invariants for research pipelines that have to be reproducible in time, not just in code.
When you run LLMs over historical data — backtesting a trading signal, replaying a research pipeline, re-scoring an archive — a normal observability stack will happily let your "2023 backtest" call a model released in 2025, rendered through a prompt you rewrote last week. The numbers come out great and mean nothing.
TraceGuard is a small Python SDK that makes that class of mistake structurally hard:
- Model registry with two timestamps —
released_at(when the model existed in the world) andavailable_to_us_at(when your system could first call it).select_model(..., strict=True)refuses anachronistic choices;stricthas no default, so every call site states its intent. - Git-tracked prompt registry — prompts are versioned YAML files;
history is
git log, and the template hash is pinned into every trace. - Reproducible input hashing — one canonical
normalize_input/input_hashimplementation (sorted keys, fixed float precision, normalized whitespace) so identical inputs hash identically across runs and machines. - Four look-ahead invariants as callable validators — call them in
pytest/CI; violations raise, nothing is silently logged-and-forgotten.
Invariants 1 and 3 are pure functions; 2 and 4 necessarily read the
registry/store and take an optional
engine. - Lightweight tracing — a
@tracer.tracedecorator, atracer.span()context manager, andwrap_anthropic/wrap_openaiclient wrappers that record every LLM/embedding/ML call (input hash, model, prompt version, output, latency, tokens, cost) into SQLite/SQLAlchemy.
"Look-ahead bias" in an LLM pipeline is really two distinct failure modes, and they need different tools. Conflating them is how teams fix one and ship the other.
| (1) Training contamination | (2) Harness / pipeline leakage | |
|---|---|---|
| What | The model was pre-trained on the future it is predicting — it recalls rather than reasons | Your code uses a model, prompt, or feature that did not exist at the simulated time |
| Lives in | The model weights | Your pipeline / orchestration code |
| Symptom | Suspiciously good on pre-cutoff data, decays after | A backtest that looks great and means nothing |
| Tooling | Membership-inference (MIN-K%), performance decay across regimes, claim-level temporal checks | Model/prompt registries, canonical input hashing, look-ahead invariants |
| TraceGuard today | Groundwork — opt-in traceguard[contamination] (interfaces + baselines) |
Primary focus — structurally refused at the registry/validator layer |
TraceGuard's mature surface targets (2): leakage that rides in through harness code — a "2023 backtest" calling a 2025 model, a prompt you rewrote last week, a vendor "actual" that was silently revised. Detection for (1) is younger and lives behind an optional extra; see docs/POSITIONING.md.
The wedge audience is people for whom a wrong-by-one-timestamp result is a correctness failure, not a cosmetic one:
- Quant / AI-for-finance researchers backtesting LLM-derived signals, where a single anachronistic model or revised "actual" inflates a Sharpe ratio.
- LLM-eval researchers measuring contamination and temporal generalization, who need provenance on which model/prompt produced which score, as of when.
- Teams replaying extraction pipelines over document archives who must answer "could this result have been produced at that point in time?"
TraceGuard is not a dashboard (Langfuse, Phoenix, LangSmith), not a proxy/gateway (Helicone), and not a general-purpose eval harness (Braintrust). Those answer "what happened and how much did it cost?". TraceGuard answers a different, lower-level question: "could this have happened at the time you're simulating?"
It is the time-integrity layer that sits underneath those tools — and it aims
to interoperate, not compete. SQLite is the default local store; an
OpenTelemetry / OpenInference exporter (traceguard[otel]) lets the same
time-correct traces flow up into Langfuse, Phoenix, or any OTLP backend
unchanged. Use your dashboard for observability; use TraceGuard to guarantee the
timeline underneath it. Step-by-step:
docs/integrations/otel-langfuse-phoenix.md.
The same goes for OpenAI-compatible gateways. TraceGuard instruments whatever
client you hand it, so a gateway is just a base_url; presets for OrcaRouter and
OpenRouter ship in traceguard.gateways, in alphabetical order and with no
provider recommended over another. Read
docs/integrations/gateways.md before pointing
one at historical data: a routing alias like orcarouter/auto or
openrouter/auto records the alias rather than the model that actually served
the call, which defeats look-ahead invariant 2 silently. Pin a concrete model
id for anything you intend to reproduce.
One boundary worth stating here rather than 200 lines down. TraceGuard makes a record correct in time; it does not make that record something a third party can check without trusting you. That is tg-attest — RFC 3161 timestamps over Merkle epoch roots, aimed at EU AI Act Article 12 rather than at backtests. TraceGuard answers "could this have happened then?"; tg-attest answers "can you prove this record has not been edited since?". Separate repo, separate package, no code dependency in either direction. Detail: Evidence layer.
TraceGuard exists because of one number. Polling a commercial fundamentals feed
against a live trading strategy's own logs over four months of 2026, 41.4% of
epsActual values differed between the value the vendor served first and the
value it serves now, and 15.3% differed enough to flip a long-entry decision.
Those are the two numbers quoted in
tg-attest and in the case study, so
the record set behind them is published and the arithmetic runs in one command
with no dependencies and no vendor account:
$ python analysis/eps_revision.pyIt prints N, the capture window, both rates with 95% Wilson intervals, the magnitude distribution, and a nine-point sweep showing what the flip rate would have been at other decision thresholds. It also re-verifies every row against its digest pair and exits non-zero if anything disagrees.
Three things worth knowing before quoting the number:
- The capture cannot be re-run. A first-seen vendor value is gone once it is overwritten; there is no vintage endpoint. What ships is the record captured at the time plus the code that turns it into the statistic, not a script that pretends it can go back and collect it again.
- The raw values are not published, because the vendor's terms forbid redistributing data "contained in or derived from" the service. Each record carries a keyed digest of each value, whether they differ, the direction and a coarse magnitude bucket. That is enough to recompute 41.4% and not enough to reconstruct anything the vendor sells.
- A second, better-designed capture gives 18.6% and 4.6%, over a broader universe and a much shorter post-print horizon. It is published in the same directory. Neither number is hidden behind the other.
Method, decision-flip definition, and a twelve-item limitations section: docs/eps-revision-methodology.md · narrative: docs/case-studies/fmp-revision.md (中文)
A Claude Code subagent whose definition omits model: runs whatever the main
thread is running. Usually that is what you want. It stops being what you want
when a frontier model is doing work nobody would have chosen a frontier model
for — at that point no routing decision is being made, and the cost lands
anyway.
python -m traceguard.routing_audit.agent_lint
# or, with nothing installed at all:
python agent_lint.py ~/.claude/agents ./.claude/agentsReads the YAML frontmatter of .claude/agents/**/*.md and nothing else — two
keys, name and model. No transcripts, no network, no dependencies, ~35 ms
on a 15-agent tree. Exits 1 if anything is unpinned, so it works as a CI gate.
model: absent and model: inherit are reported separately. Both run the
parent's model; only one of them is a decision somebody made.
That answers whether you have the exposure. What those agents actually ran, and what it cost, needs a trace store —
The opt-in traceguard.routing_audit extension applies the same discipline to
your own agent history: it ingests Claude Code session transcripts
(~/.claude/projects/**/*.jsonl) into an append-only, message.id-keyed
trace store, so usage history stops changing underneath you. Claude Code
rewrites session files in place on resume/compact; anything that recomputes
totals from live files inherits that drift.
From that log it scores each (unit, component) routing decision against a
policy file you write: the tier the policy expected, the model that actually
ran, and a verdict. The verdict has three shapes rather than two — compliant,
deviation, and unresolved, the last split into no_rule (nothing in the
policy matched) and unknown_model (the model sits outside every tier). A
default that quietly resolves those cases fabricates verdicts in both
directions: on this corpus it scored one component compliant and another
deviant, each against a rule nobody had written.
Current corpus, 789 decisions: 627 compliant, 160 deviation, 2
unresolved:no_rule, 0 unresolved:unknown_model. The two unresolved classes
are counted apart because they decay differently. unknown_model is an
operational gap that clears itself once the tier table catches up; no_rule
stays until somebody writes a rule or decides the case belongs outside policy.
The summary carries two coverage counts next to them: decisions no rule reached, and rules no decision reached. Those catch different failures, uncovered behavior on one side, dead or shadowed policy on the other. Brian Jin, whose comment prompted the three-state verdict, put it this way: "That makes unresolved much more than an error bucket. It becomes an observable property of the policy surface itself."
Having a stable reference log turned out to be enough to audit other tools. Run against splitrail, a Rust usage tracker for agentic CLIs, it has produced three upstream fixes:
| finding | upstream |
|---|---|
| Totals drift because resume/compact rewrites live JSONL | #200 → SQLite history store in 3.6.0 |
Subagent transcripts (<session>/subagents/**) sit below a depth-2 discovery cap — on this corpus, 54% of live messages and ~1/3 of the spend never entered any total |
#207 → maintainer-authored #209 in 3.6.1 |
| Partial streaming snapshots summed per fingerprint, inflating input +62% / cache_read +69% on subagent transcripts (this corpus); inflation ratios predicted from the corpus before the fix branch existed, then matched on every field | #220 → #222, corpus-verified and merged the same day |
On the subset splitrail scans, 3.6.0 and the audit log agree token-exact — 18,548,947 output tokens on both sides across ~13.5k messages. Two independent implementations, different languages, different dedup strategies, same number to the digit; that exactness is what turned the remaining gap into a nameable defect instead of a rounding argument. The end-to-end regression fixture from that work was merged upstream.
The same reference log now also emits the cross-tool
usage-drift-log
record each scheduled run (--usage-report-history) — a six-field per-run
spec published by clauderank after the drift pattern reproduced at
leaderboard scale (viberank#83);
routing_audit is its second independent implementation.
viberank became the third, and adopting it there forced six revisions —
per-month scoping, absence is not deletion, and the multi-agent source gap
among them. Which tool implements what, each revision and the measurement
that forced it, and what is still open:
docs/specs/usage-drift-log.md.
Method and reproduction protocol: splitrail-validation/ ·
write-up: An append-only audit log caught two accounting bugs in a 216-star usage tracker
The same protocol pointed at
claude-code-templates
(30k★) found the opposite sign — a 2.36× over-count, fix submitted as
PR #754:
cct-dedup-check/ ·
write-up: The vendor documents this bug. A 30k-star repo shipped it anyway.
What the series adds up to is a catalog: eleven invariants anything counting
tokens from ~/.claude/projects has to hold, each one measured in a shipped
tracker before it was written down — CONFORMANCE.md.
Maintainers can run the checks in CI with a drop-in workflow:
ci/.
pip install traceguardRequires Python 3.11+. Core dependencies: SQLAlchemy 2, Pydantic 2, PyYAML.
The Anthropic and OpenAI wrappers are extras:
pip install "traceguard[anthropic]" / pip install "traceguard[openai]".
Anchoring the audit chain to OpenTimestamps needs
pip install "traceguard[anchors]"; everything else in traceguard.audit
works without it. traceguard.sources needs no extra and is
experimental — it is outside the frozen public surface and outside the
contract-guard CI job, so its API can still change in a minor.
To track the development version instead of PyPI releases:
# pyproject.toml
[project]
dependencies = [
"traceguard @ git+https://github.com/lizhuojunx86/traceguard.git@main#subdirectory=packages/traceguard",
]Everything below is synthetic and runnable — see examples/quickstart for the full script.
from datetime import datetime, timezone
from traceguard.registry.models import register_model, select_model
from traceguard.store.models import make_engine
engine = make_engine("sqlite:///:memory:")
UTC = timezone.utc
register_model("demo-llm-2024", model_family="internal-ml",
capability_class="general-llm",
released_at=datetime(2024, 1, 10, tzinfo=UTC),
available_to_us_at=datetime(2024, 2, 1, tzinfo=UTC),
engine=engine)
register_model("demo-llm-2026", model_family="internal-ml",
capability_class="general-llm",
released_at=datetime(2026, 1, 5, tzinfo=UTC),
available_to_us_at=datetime(2026, 1, 15, tzinfo=UTC),
engine=engine)
# Backtesting as of mid-2025: the 2026 model must be invisible.
backtest_date = datetime(2025, 6, 30, tzinfo=UTC)
model_id = select_model("general-llm", available_at=backtest_date,
strict=True, engine=engine)
# -> "demo-llm-2024"; at a 2023 date it raises NoEligibleModelErrorTrace a call with version pinning:
from traceguard.registry.prompts import load_prompt
from traceguard.sdk.tracer import Tracer
prompt = load_prompt("demo/extractor/v1", prompts_root="prompts")
tracer = Tracer(engine)
with tracer.span("myproject", "extractor", "llm_complete",
correlation_id="doc-001", feature_as_of=backtest_date) as span:
span.record_input({"text": prompt.render(text="...")})
span.record_model_prompt(model_id=model_id,
prompt_template_id=prompt.prompt_template_id,
prompt_template_hash=prompt.prompt_template_hash)
# ... call the model ...
span.record_output(parsed={"entities": []}, parse_status="success")
span.record_perf(latency_ms=42, tokens_in=120, tokens_out=18)Enforce the invariants in CI:
from traceguard.validators.lookahead import (
validate_feature_as_of, validate_model_timing, InvariantViolation,
)
# Invariant 2: a 2025 feature may not be computed by a 2026 model.
validate_model_timing("demo-llm-2026", backtest_date, strict=True, engine=engine)
# -> raises InvariantViolation: [invariant 2] model 'demo-llm-2026'
# available_to_us_at=2026-01-15 is after feature_as_of=2025-06-30| # | Invariant | Validator |
|---|---|---|
| 1 | A derived feature's feature_as_of ≤ the earliest timestamp of all its inputs |
validate_feature_as_of |
| 2 | The model used must satisfy available_to_us_at ≤ feature_as_of (strict), or carry an explicit anachronism flag (loose) |
validate_model_timing |
| 3 | Any time-versioned reference data (prompt templates, alias tables, lookup dictionaries) must satisfy valid_from ≤ feature_as_of |
validate_reference_timing |
| 4 | A locked replay set is immutable | assert_replay_set_locked |
The full interface contract — table schemas, SDK signatures, semantics, and SemVer rules — lives in docs/SPEC.md (English) and TRACEGUARD_SPEC.md (Chinese original, authoritative).
The four invariants cover the model, the prompt, and feature ordering. They do
not cover the data the pipeline fetched. A backtest can use the right model,
the right prompt and the right feature_as_of, be handed a vendor value that
was rewritten into existence weeks later, and pass all four while being wrong.
That is measured, not hypothetical: 41.4% of vendor epsActual values
differ between first sight and today, and 15.3% flip a binary entry
decision. A second capture put the same two figures at 18.6% and 4.6%.
analysis/eps_revision.py recomputes both offline
from data committed to this repo.
traceguard.sources (opt-in, off the frozen import surface) records one
source_snapshot per retrieval — content_hash, retrieved_at, the source's
claimed published_at, and a verdict from invariant 3:
| verdict | meaning |
|---|---|
verified |
published_at ≤ feature_as_of — the content demonstrably already existed |
anachronistic |
it did not exist yet |
unverifiable |
the source states no published_at, so existence can be neither shown nor ruled out |
unchecked |
the call site passed no feature_as_of; nothing was compared |
strict=True refuses the last two; strict=False records them. strict is
keyword-only with no default, so every call site states its intent — the same
discipline select_model uses. python -m traceguard.sources --db URL drift
then reports which sources changed content between retrievals, as a rate with
its n and a Wilson 95% interval.
No retrieved content is stored — digests and metadata only. And the limits
are stated rather than glossed: a snapshot proves which bytes the host handed
over and how their claimed publication time relates to feature_as_of. It does
not prove the host actually fetched them from source_uri (TraceGuard never
made the request), nor that published_at is true — that is what the source
says about itself. Details: docs/sources.md.
Time-correct traces are only worth as much as the guarantee that they were not
edited afterwards. The opt-in traceguard.audit submodule (1.1.0; contract-stable
since SPEC v1.1, off the frozen import surface) adds that guarantee to the traces table:
- Append-only guard at the ORM layer — blocks accidental UPDATE/DELETE
against
traces;cost_usd-only updates pass, since repricing is the one legal mutation the spec allows. - Row hash chain — every inserted trace is chained as
sha256(prev_hash || canonical(entry)), with the algorithm frozen by golden tests.verify_chain()recomputes the chain in two passes and reports BREAK / WARN / GAP findings. - Exportable anchor —
export_anchor()emits the chain head for storage outside the database, which is what turns tamper-evident into something an auditor can actually check. - CLI —
python -m traceguard.audit enable|disable|verify|anchor|reconcile;anchor --sink file:…|git-note:…|webhook:… [--every SECONDS]stores the head outside the DB on a cadence, andreconcilechecks self-reported token volume against the provider's usage report (capture_mismatch). - Per-request existence check (L1.5) — totals can cancel: an
under-reported call and an over-reported one net out, and the provider
usage API reports tokens but no call counts.
reconcile --source requests-json:PATHinstead joinstraces.provider_response_idagainst arequest-ledger/v1document from an out-of-band source, one call at a time. A call present on only one side iscapture_unmatched(WARN) carrying a direction —out_of_band_only(a call the capture layer did not see) orself_reported_only(a record the provider side does not vouch for). It confirms that matched calls exist on both sides; it cannot vouch for their content, which stays out of scope.
Boundaries are stated rather than glossed: this is tamper-evident, not tamper-proof. Core SQL, raw drivers, and bulk APIs bypass the ORM guard, and without an external anchor a full-chain rewrite is undetectable. Details: docs/audit.md.
Where this stops. export_anchor() emits the chain head; what you do with it
afterwards is out of scope here, and that last sentence is the reason it matters.
tg-attest is the sibling project that
closes the gap — RFC 3161 timestamps over Merkle epoch roots, selective disclosure
of a single record without exposing the rest of the batch, and a disclosure bundle
an auditor checks with openssl ts and a CA certificate they fetch themselves.
Same point-in-time problem, aimed at EU AI Act Article 12 rather than at backtests.
Separate repo, separate package, no code dependency in either direction.
TraceGuard's harness-leakage invariants are the engineering counterpart to a
growing body of work on temporal validity and contamination in LLMs. The
contamination groundwork (extra traceguard[contamination]) draws on:
- A Test of Lookahead Bias in LLM Forecasts — Gao, Jiang & Yan, arXiv 2512.23847
- Look-Ahead-Bench: a Standardized Benchmark of Look-ahead Bias in Point-in-Time LLMs for Finance — Benhenda, arXiv 2601.13770
- All Leaks Count, Some Count More: Interpretable Temporal Contamination Detection in LLM Backtesting (TimeSPEC / Shapley-DCLR) — Zhang, Chen & Stadie, arXiv 2602.17234
- MIN-K% PROB — Detecting Pretraining Data from Large Language Models, Shi et al., arXiv 2310.16789
See docs/POSITIONING.md for how these map onto the two kinds of look-ahead.
This repo hosts two Python packages:
| Package | Path | Status |
|---|---|---|
traceguard — the SDK described above |
packages/traceguard/ | Active development; all new features land here (public API frozen under SemVer since 1.0.0) |
pipeline-guardian (import name guardian) — checkpoint validation for multi-agent pipelines: structural checks, LLM-as-Judge, retry/abort actions, dashboard |
repo root (guardian/) |
Frozen: bugfixes only; its 4-symbol public API stays stable for existing integrators |
Pipeline Guardian's full documentation is in docs/pipeline-guardian.md. The two packages share no imports and release independently.
analysis/ belongs to neither. It holds the published evidence for
the epsActual revision claim — the disclosure dataset, the script that
recomputes the two headline rates from it, and the builder that produced it from
the private captures.
# SDK
cd packages/traceguard
uv sync --extra openai # keep the extra: a bare `uv sync` uninstalls it
uv run pytest # 1053 tests (3 skip without the contamination-hf extra)
# Pipeline Guardian (legacy)
uv sync && uv run pytest # 293 tests, from repo rootRoadmap: TRACEGUARD_ROADMAP.md. Phase 0 was accepted in
June 2026, and 1.0.0 froze the contract: every SPEC MUST implemented and
enforced (invariants 1–4), a curated 29-symbol public surface held by a required
contract-guard CI job, and fail-open instrumentation that never breaks or masks
the host call. The public API is stable under SemVer from 1.0.0 onward.
Everything since has been additive and opt-in, off the frozen surface:
OpenTelemetry export and live dual-write (traceguard[otel]),
training-contamination detection incl. MIN-K%++ (traceguard.contamination),
loop evidence-gating (traceguard.loop), wrap_openai / wrap_anthropic,
the 1.1.0 audit evidence layer (traceguard.audit), and the 1.2.0 read-only
prompt-cache efficiency audit (traceguard.routing_audit.cache_audit). Full
history: CHANGELOG.
1.3.0 rebuilt that cache audit's keep-alive answer after three separate
recomputations overturned it: the give-up cap is now solved by a sweep instead
of hand-picked, expiry cost is a bracket instead of one bound, mid-gap model
switches are measured instead of assumed, and the reportable conclusion is a
band (RECOMMENDED CAP: 9h..12h) rather than a single argmax. --benchmark
pins the window every quoted number comes from.
1.4.0 lets that audit leave the laptop. --emit-share writes an aggregate
JSON summary and --show-share prints the exact bytes first, so the sender
reads all of it before deciding; a corpus.fingerprint identifies which traces
a window actually loaded, because a closed window bounds timestamps and not the
corpus; entries are immutable and refuse to be overwritten; and
generated_at / settling_days record when the pull happened, not only what
it pulled. benchmark/ collects the results. See the
package README.
1.5.0 makes the audit layer part of the contract and gives every trace
an owner. SPEC v1.1 adds agent_id / session_id to traces — nullable,
indexed, added to an existing database on open — so several executors can be
correlated after the fact; both columns sit outside the algo v1 hash envelope,
and docs/audit.md says so rather than implying otherwise.
traceguard.audit is no longer experimental: its surface, finding kinds and
boundary statements are frozen by the same CI job that guards the public
import surface. On top of it, anchor sinks (file: / git-note: /
webhook:, --every for a cadence) shrink the exposure window that boundary
statement 1 names, and reconcile compares self-reported token volume with
the provider's usage report per model and UTC bucket; a capture_mismatch
says which side has more and what that usually means. The prompt was the
2026-08-26 METR/Redwood investigation of the OpenAI–Hugging Face incident:
roughly 7% of the evaluated transcripts had spoofed tool calls, and agents
researched how to spoof, edit or delete their own transcripts. The chain
answers whether a stored record was changed afterwards; reconcile is the
first honest step on whether it was true.
1.6.0 (SPEC v1.2) works on the other two halves of that question: which
call produced a number, and whether anyone else can check it. traces gains
provider_response_id (nullable, outside the algo v1 hash envelope), and
reconcile_requests joins it against an out-of-band request-ledger/v1
document one call at a time — totals can cancel, because an under-report and an
over-report net out and the provider's usage API gives no call counts at all;
existence cannot. export_bundle writes an evidence-bundle/v1 document that a
recipient verifies with no database, no network and no traceguard installed; it
reads VERIFIED only when an anchor covers the entries the bundle carries, and
INTERNALLY CONSISTENT otherwise, because a rewrite that re-chains the segment is
internally consistent too. The chain can now be anchored to OpenTimestamps
(ots:), whose still-pending proofs are reported as a calendar server's promise
rather than as attested time. New and experimental: traceguard.sources
records a digest and the timing of each retrieval — never the content — grades
invariant 3 against feature_as_of, and reports which sources rewrote what they
had already served with an n and a Wilson interval instead of a bare percentage.
Licensed under the Apache License 2.0.