Consolidate the ported experiment code onto shared helpers (0.2.1) - #2
Merged
Merged
Conversation
…(0.2.1)
Behaviour-preserving cleanup of the code shipped in 0.2.0, verified against the
paper's own run artifacts where they exist locally (cross-lens summaries and
worked-example tables, error tails, paired matched-KL records, DE bank:
byte-identical apart from added provenance).
Structure
- research/cross_lens.py, research/wsd.py, research/row_geometry.py hold the
helpers that were copied across scripts (jlens import, model/lens loading,
dump parsing, Wilson CIs, agreement rule, bundle IO, cluster bootstrap,
spherical k-means, bare-first token resolver); scripts route through
qwen_readout.load_sae / encode_topk, factorizers.topk_mask,
data.center_normalize_rows, utils.load_causal_lm / write_json,
run_io.run_provenance; one script loader in tests/conftest.py.
- Unused AmbiStory / anchors paths removed from the WSD scripts.
Paper-level choices made explicit (defaults reproduce the paper)
- --centering {live,trained} on the cross-lens, causal, stability and WSD
scripts (data.centering_mean); on Qwen3.5-2B the two means differ by ~0.04%
of a centred row norm.
- --agreement-rule {half,strict} and --null-population {all,cross} on the
cross-lens aggregators; --knn-exclude-self on loo_core_recovery.
Fixed
- feature_group_matching: null pool no longer depends on PYTHONHASHSEED;
below-null tail uses the unrounded Jaccard.
- run_readout_baseline_comparisons: per-row-support methods aligned before
A-B subtraction (coverage/compactness columns only); k-means fitted once per
(n_clusters, seed); determinism claim corrected.
- fit_jlens: OOM fallback reaches dim_batch 1; shard/merge resume guarded by
metadata sidecars. fit_ridge_lens: --holdout >= 1, holdout-only diagnostic,
guarded resume.
- run_wsd_feature_alignment: left truncation in cloze mode; --analyze-only no
longer overwrites run_config.json. run_causal_contribution_validation:
self-test's realized-change check compares against the dense LM head.
- nearest_rows_baseline: empty tables no longer crash.
- analyze_error_tails bootstrap vectorised (identical draws, 3x faster).
Docs: THIRD_PARTY.md lists jlens (Apache-2.0, git source), pandas,
scikit-learn, CoarseWSD-20 and the C4 lens corpora; config headers and
REPRODUCE rows corrected; version 0.2.1.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
There was a problem hiding this comment.
🔵 Needs a closer look
It refactors and rethreads a large set of research scripts/helpers with behavioral and operational implications, so it needs careful human validation beyond the targeted review findings.
Pull request overview
This PR consolidates previously duplicated “ported experiment” logic into shared sparse_readout_prism.research helpers, makes paper-default conventions explicit via CLI flags, and adds standardized provenance metadata across outputs to support reproducibility for the 0.2.1 release.
Changes:
- Centralizes shared experiment logic (centering policies, k-means, token resolution, WSD bundle/stat utilities) into new/expanded
research/*helpers and routes scripts through them. - Updates multiple scripts’ CLIs/output schemas to expose paper-default behavior explicitly (e.g.,
--centering,--agreement-rule,--null-population,--knn-exclude-self,--device) and to record provenance. - Expands/adjusts tests to pin byte-/bit-identical behavior where required and to validate new provenance blocks; bumps project version to
0.2.1and updates docs/licenses.
File summaries
| File | Description |
|---|---|
| uv.lock | Bumps editable package version to 0.2.1. |
| tests/test_paired_matched_kl_bootstrap.py | Updates smoke test for new JSON top-level {comparisons, provenance} output. |
| tests/test_lexical_control_directions.py | Switches to shared load_script helper for script import. |
| tests/test_error_tails_and_nearest_rows.py | Adds regression tests pinning vectorized bootstrap gathers and CSV byte output behavior. |
| tests/test_causal_contribution_validation.py | Updates/extends self-test assertions and shared token resolver usage. |
| tests/test_baseline_and_dla.py | Adds coverage for consolidated _finish, k-means memoization, and margin_from_rows support alignment. |
| tests/conftest.py | Adds shared load_script helper for importing scripts/ entry points in tests. |
| src/sparse_readout_prism/research/wsd.py | New shared CoarseWSD-20 bundle reader/validators, split rules, nulls, and bootstrap utilities. |
| src/sparse_readout_prism/research/seed_stability.py | Refactors stability helpers to shared centering/top-k/token-resolution utilities and updates checkpoint loading. |
| src/sparse_readout_prism/research/row_geometry.py | New shared spherical k-means and bare-first single-token resolver. |
| src/sparse_readout_prism/research/init.py | Updates module overview to include new research helpers. |
| src/sparse_readout_prism/data.py | Adds centering_mean policy helper and token_mask_from_tokenizer. |
| scripts/run/run_readout_baseline_comparisons.py | Consolidates baseline method accounting via _finish, adds feature-space alignment for margins, and memoizes k-means centroids. |
| scripts/run/run_qwen_profanity_suppression_eval.py | Clarifies PCA rank-4 description wording. |
| scripts/run/run_cross_lens_readouts.py | Routes through shared cross-lens/qwen_readout helpers; adds centering/device/k handling + manifest provenance. |
| scripts/run/fit_ridge_lens.py | Adds resume metadata sidecar checks, holdout validation, device flag, and provenance in report. |
| scripts/run/fit_jlens.py | Adds safer OOM fallback schedule, shard sidecars + merge validation, device flag, and provenance metadata. |
| scripts/README.md | Documents cross-lens --device and --centering flags. |
| scripts/figures/compute_cross_lens_shared_feature.py | Uses shared dump parsing/manifest writing and adds provenance to JSON output. |
| scripts/eval/run_causal_contribution_validation.py | Adds explicit centering policy + shared token resolver; updates self-test to dense-head measurement; provenance includes centering. |
| scripts/eval/paired_matched_kl_bootstrap.py | Changes output to include provenance alongside comparison records. |
| scripts/eval/loo_core_recovery.py | Adds explicit paper conventions (--knn-exclude-self, centering), shared k-means, and provenance in output. |
| scripts/eval/feature_group_matching.py | Fixes hash-seed dependence in null pool enumeration, aligns summaries, adds centering + provenance. |
| scripts/eval/cross_seed_stability.py | Refactors to shared helpers, adds centering + provenance. |
| scripts/eval/analyze_error_tails.py | Vectorizes clustered bootstrap gathers while pinning output identity; documents change. |
| scripts/eval/aggregate_cross_lens_en_zh.py | Moves aggregation logic into shared cross-lens helpers; adds --agreement-rule + provenance. |
| scripts/eval/aggregate_cross_lens_en_de.py | Moves aggregation logic into shared cross-lens helpers; adds --agreement-rule, --null-population + provenance. |
| scripts/data/extract_model_readout.py | Uses shared token_mask_from_tokenizer. |
| scripts/data/build_cross_lens_de_bank.py | Uses shared fold_text, adjusts null token handling, adds provenance to exclusion report. |
| scripts/analyze/nearest_rows_baseline.py | Writes CSVs via shared write_csv with fixed field lists; handles empty outputs. |
| scripts/analyze/cross_lens_three_lens_prompt.py | Adds manifest provenance when writing CSVs via shared helper. |
| scripts/analyze/cross_lens_antonym_layers.py | Routes through shared cross-lens/qwen_readout helpers; adds centering/device/k handling + manifest provenance. |
| scripts/analyze/analyze_wsd_sense_groups.py | Uses shared WSD bundle helpers; fixes HF cache resolution; adds centering + provenance plumbing. |
| README.md | Documents 0.2.1 consolidation, flags, and scope changes. |
| pyproject.toml | Bumps version to 0.2.1; updates dependency comments. |
| docs/THIRD_PARTY.md | Expands dependency/license notes; documents jlens sourcing and data notes. |
| docs/DATA.md | Updates data layout description (removes data/wsd/ mention). |
| data/wsd/wsd_sense_anchors.json | Removes unused optional WSD anchor file. |
| data/cross_lens/README.md | Clarifies bank schema details and control/null semantics. |
| configs/sweeps/README.md | Clarifies recipe-matching semantics and adds notes about omitted inert keys. |
| configs/sweeps/qwen35_2b_seedvar_base.yaml | Adds log/checkpoint cadence keys (documented as cadence-only). |
| configs/sweeps/qwen35_0p8b_paper_topk_32x_k256_s2.yaml | Clarifies recipe matching and adds cadence-only keys. |
| configs/sweeps/qwen35_0p8b_paper_topk_32x_k256_s1.yaml | Clarifies recipe matching and adds cadence-only keys. |
| configs/sweeps/lowk_qwen35_2b_16x_base.yaml | Adds cadence-only keys. |
| configs/sweeps/lowk_qwen35_0p8b_16x_base.yaml | Adds cadence-only keys. |
| CITATION.cff | Bumps version to 0.2.1 (release date retained). |
Review details
- Files reviewed: 52/53 changed files
- Comments generated: 1
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Comment on lines
789
to
+793
| cur = ( | ||
| d.original_logit, | ||
| d.reconstructed_logit, | ||
| d.residual_term, | ||
| d.feature_contributions, | ||
| _full_contributions(d), |
hematteo
added a commit
that referenced
this pull request
Aug 31, 2026
tests/test_wsd_sense_groups.py pinned SHA-256 hashes of full-precision outputs (and a few float literals) captured on Apple Silicon; on the ubuntu CI runners BLAS accumulation differs at ~1e-7 relative, so three tests failed after PR #2 merged (the merge gate was mis-wired and let a red build through). The pins are now fixture values in tests/fixtures/wsd_pins.json compared with tolerances (rel 1e-5, atol 1e-6; ints, ids and orderings exact), old-vs-new comparisons use allclose, and the file passes under perturbed fixtures and different PYTHONHASHSEED values. No script code changes. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Behaviour-preserving consolidation of the experiment code shipped in 0.2.0, following a code review of that release. Shared helpers replace the copies that had accumulated across scripts; the paper-level choices those scripts made implicitly are now explicit flags whose defaults reproduce the paper; the small bugs the review found are fixed; every result file records provenance.
Structure.
research/cross_lens.py,research/wsd.py,research/row_geometry.py; scripts route throughqwen_readout.load_sae/encode_topk,factorizers.topk_mask,data.center_normalize_rows/centering_mean,utils.load_causal_lm/write_json,run_io.run_provenance; one script loader intests/conftest.py; unused AmbiStory/anchors paths removed (WSD run script 1,236 → 792 lines).Flags (defaults = paper).
--centering {live,trained}on the cross-lens, causal, stability and WSD scripts;--agreement-rule {half,strict}and--null-population {all,cross}on the cross-lens aggregators;--knn-exclude-selfon LOO core recovery;--deviceon the GPU cross-lens scripts;--kdefaults to the checkpoint's k.Fixed.
PYTHONHASHSEED-dependent null pool in feature-group matching; per-row-support alignment inmargin_from_rows(coverage/compactness columns only); k-means fitted once per(n_clusters, seed); jlens OOM fallback reachingdim_batch1; resume metadata sidecars forfit_jlens/fit_ridge_lens;--holdoutvalidation; left truncation for WSD cloze prompts;--analyze-onlyno longer overwritesrun_config.json; the causal self-test's realized-change check now compares against the dense LM head; empty nearest-rows tables; error-tails bootstrap vectorised with identical draws (3× faster).Verification against the paper's own artifacts (byte-identical unless noted)
--analyze-onlymetrics identical on the Qwen3.5-2B and R1-Llama-8B paper bundles; scoring path pinned bit-for-bit on synthetic checkpoints.decompose_rowfield and scalar margin bit-exact across all eight method families; stability scripts: set-derived quantities identical, 19 float32 cosine stats move ≤ 2e-7 (shared reader's unit-norm guard).ruffclean (89 files), 242 tests pass, synthetic smoke pipeline passes.Documented, not changed
The paper's centering (live full-vocabulary mean; measured ~0.04 % of a centred row norm from the training mean on Qwen3.5-2B), the ≥-half agreement rule (a strict majority gives 72/80), the ridge λ grid (both paper fits selected the grid's top value at every layer), k-means baseline leakage and the LOO kNN self-exclusion asymmetry are stated in the docstrings and
REPRODUCE.mdand left at the paper's behaviour by default.🤖 Generated with Claude Code