Skip to content

Consolidate the ported experiment code onto shared helpers (0.2.1) - #2

Merged
hematteo merged 1 commit into
mainfrom
cleanup/consolidate-ported-code
Aug 31, 2026
Merged

hematteo merged 1 commit into
mainfrom
cleanup/consolidate-ported-code

Conversation

@hematteo

Copy link
Copy Markdown
Owner

Summary

Behaviour-preserving consolidation of the experiment code shipped in 0.2.0, following a code review of that release. Shared helpers replace the copies that had accumulated across scripts; the paper-level choices those scripts made implicitly are now explicit flags whose defaults reproduce the paper; the small bugs the review found are fixed; every result file records provenance.

Structure. research/cross_lens.py, research/wsd.py, research/row_geometry.py; scripts route through qwen_readout.load_sae / encode_topk, factorizers.topk_mask, data.center_normalize_rows / centering_mean, utils.load_causal_lm / write_json, run_io.run_provenance; one script loader in tests/conftest.py; unused AmbiStory/anchors paths removed (WSD run script 1,236 → 792 lines).

Flags (defaults = paper). --centering {live,trained} on the cross-lens, causal, stability and WSD scripts; --agreement-rule {half,strict} and --null-population {all,cross} on the cross-lens aggregators; --knn-exclude-self on LOO core recovery; --device on the GPU cross-lens scripts; --k defaults to the checkpoint's k.

Fixed. PYTHONHASHSEED-dependent null pool in feature-group matching; per-row-support alignment in margin_from_rows (coverage/compactness columns only); k-means fitted once per (n_clusters, seed); jlens OOM fallback reaching dim_batch 1; resume metadata sidecars for fit_jlens / fit_ridge_lens; --holdout validation; left truncation for WSD cloze prompts; --analyze-only no longer overwrites run_config.json; the causal self-test's realized-change check now compares against the dense LM head; empty nearest-rows tables; error-tails bootstrap vectorised with identical draws (3× faster).

Verification against the paper's own artifacts (byte-identical unless noted)

  • Cross-lens: EN–ZH summary key-for-key (77/80, floors 0/159 · 0/240), both antonym-layer tables, three-lens tables, shared-feature metrics, DE bank re-built byte-identical.
  • Error tails: nine CSVs identical on the six paper run dirs (10,000 resamples). Paired matched-KL: 96 records identical.
  • WSD: sense-group sweep (g = 1/2/4/8), classifier framing and --analyze-only metrics identical on the Qwen3.5-2B and R1-Llama-8B paper bundles; scoring path pinned bit-for-bit on synthetic checkpoints.
  • Baseline runner: every decompose_row field and scalar margin bit-exact across all eight method families; stability scripts: set-derived quantities identical, 19 float32 cosine stats move ≤ 2e-7 (shared reader's unit-norm guard).
  • ruff clean (89 files), 242 tests pass, synthetic smoke pipeline passes.

Documented, not changed

The paper's centering (live full-vocabulary mean; measured ~0.04 % of a centred row norm from the training mean on Qwen3.5-2B), the ≥-half agreement rule (a strict majority gives 72/80), the ridge λ grid (both paper fits selected the grid's top value at every layer), k-means baseline leakage and the LOO kNN self-exclusion asymmetry are stated in the docstrings and REPRODUCE.md and left at the paper's behaviour by default.

🤖 Generated with Claude Code

…(0.2.1)

Behaviour-preserving cleanup of the code shipped in 0.2.0, verified against the
paper's own run artifacts where they exist locally (cross-lens summaries and
worked-example tables, error tails, paired matched-KL records, DE bank:
byte-identical apart from added provenance).

Structure
- research/cross_lens.py, research/wsd.py, research/row_geometry.py hold the
  helpers that were copied across scripts (jlens import, model/lens loading,
  dump parsing, Wilson CIs, agreement rule, bundle IO, cluster bootstrap,
  spherical k-means, bare-first token resolver); scripts route through
  qwen_readout.load_sae / encode_topk, factorizers.topk_mask,
  data.center_normalize_rows, utils.load_causal_lm / write_json,
  run_io.run_provenance; one script loader in tests/conftest.py.
- Unused AmbiStory / anchors paths removed from the WSD scripts.

Paper-level choices made explicit (defaults reproduce the paper)
- --centering {live,trained} on the cross-lens, causal, stability and WSD
  scripts (data.centering_mean); on Qwen3.5-2B the two means differ by ~0.04%
  of a centred row norm.
- --agreement-rule {half,strict} and --null-population {all,cross} on the
  cross-lens aggregators; --knn-exclude-self on loo_core_recovery.

Fixed
- feature_group_matching: null pool no longer depends on PYTHONHASHSEED;
  below-null tail uses the unrounded Jaccard.
- run_readout_baseline_comparisons: per-row-support methods aligned before
  A-B subtraction (coverage/compactness columns only); k-means fitted once per
  (n_clusters, seed); determinism claim corrected.
- fit_jlens: OOM fallback reaches dim_batch 1; shard/merge resume guarded by
  metadata sidecars. fit_ridge_lens: --holdout >= 1, holdout-only diagnostic,
  guarded resume.
- run_wsd_feature_alignment: left truncation in cloze mode; --analyze-only no
  longer overwrites run_config.json. run_causal_contribution_validation:
  self-test's realized-change check compares against the dense LM head.
- nearest_rows_baseline: empty tables no longer crash.
- analyze_error_tails bootstrap vectorised (identical draws, 3x faster).

Docs: THIRD_PARTY.md lists jlens (Apache-2.0, git source), pandas,
scikit-learn, CoarseWSD-20 and the C4 lens corpora; config headers and
REPRODUCE rows corrected; version 0.2.1.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Copilot AI lite review requested due to automatic review settings August 31, 2026 15:44
@hematteo
hematteo merged commit 73e3074 into main Aug 31, 2026
3 of 5 checks passed
@hematteo
hematteo deleted the cleanup/consolidate-ported-code branch August 31, 2026 15:45

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

It refactors and rethreads a large set of research scripts/helpers with behavioral and operational implications, so it needs careful human validation beyond the targeted review findings.

Pull request overview

This PR consolidates previously duplicated “ported experiment” logic into shared sparse_readout_prism.research helpers, makes paper-default conventions explicit via CLI flags, and adds standardized provenance metadata across outputs to support reproducibility for the 0.2.1 release.

Changes:

  • Centralizes shared experiment logic (centering policies, k-means, token resolution, WSD bundle/stat utilities) into new/expanded research/* helpers and routes scripts through them.
  • Updates multiple scripts’ CLIs/output schemas to expose paper-default behavior explicitly (e.g., --centering, --agreement-rule, --null-population, --knn-exclude-self, --device) and to record provenance.
  • Expands/adjusts tests to pin byte-/bit-identical behavior where required and to validate new provenance blocks; bumps project version to 0.2.1 and updates docs/licenses.
File summaries
File Description
uv.lock Bumps editable package version to 0.2.1.
tests/test_paired_matched_kl_bootstrap.py Updates smoke test for new JSON top-level {comparisons, provenance} output.
tests/test_lexical_control_directions.py Switches to shared load_script helper for script import.
tests/test_error_tails_and_nearest_rows.py Adds regression tests pinning vectorized bootstrap gathers and CSV byte output behavior.
tests/test_causal_contribution_validation.py Updates/extends self-test assertions and shared token resolver usage.
tests/test_baseline_and_dla.py Adds coverage for consolidated _finish, k-means memoization, and margin_from_rows support alignment.
tests/conftest.py Adds shared load_script helper for importing scripts/ entry points in tests.
src/sparse_readout_prism/research/wsd.py New shared CoarseWSD-20 bundle reader/validators, split rules, nulls, and bootstrap utilities.
src/sparse_readout_prism/research/seed_stability.py Refactors stability helpers to shared centering/top-k/token-resolution utilities and updates checkpoint loading.
src/sparse_readout_prism/research/row_geometry.py New shared spherical k-means and bare-first single-token resolver.
src/sparse_readout_prism/research/init.py Updates module overview to include new research helpers.
src/sparse_readout_prism/data.py Adds centering_mean policy helper and token_mask_from_tokenizer.
scripts/run/run_readout_baseline_comparisons.py Consolidates baseline method accounting via _finish, adds feature-space alignment for margins, and memoizes k-means centroids.
scripts/run/run_qwen_profanity_suppression_eval.py Clarifies PCA rank-4 description wording.
scripts/run/run_cross_lens_readouts.py Routes through shared cross-lens/qwen_readout helpers; adds centering/device/k handling + manifest provenance.
scripts/run/fit_ridge_lens.py Adds resume metadata sidecar checks, holdout validation, device flag, and provenance in report.
scripts/run/fit_jlens.py Adds safer OOM fallback schedule, shard sidecars + merge validation, device flag, and provenance metadata.
scripts/README.md Documents cross-lens --device and --centering flags.
scripts/figures/compute_cross_lens_shared_feature.py Uses shared dump parsing/manifest writing and adds provenance to JSON output.
scripts/eval/run_causal_contribution_validation.py Adds explicit centering policy + shared token resolver; updates self-test to dense-head measurement; provenance includes centering.
scripts/eval/paired_matched_kl_bootstrap.py Changes output to include provenance alongside comparison records.
scripts/eval/loo_core_recovery.py Adds explicit paper conventions (--knn-exclude-self, centering), shared k-means, and provenance in output.
scripts/eval/feature_group_matching.py Fixes hash-seed dependence in null pool enumeration, aligns summaries, adds centering + provenance.
scripts/eval/cross_seed_stability.py Refactors to shared helpers, adds centering + provenance.
scripts/eval/analyze_error_tails.py Vectorizes clustered bootstrap gathers while pinning output identity; documents change.
scripts/eval/aggregate_cross_lens_en_zh.py Moves aggregation logic into shared cross-lens helpers; adds --agreement-rule + provenance.
scripts/eval/aggregate_cross_lens_en_de.py Moves aggregation logic into shared cross-lens helpers; adds --agreement-rule, --null-population + provenance.
scripts/data/extract_model_readout.py Uses shared token_mask_from_tokenizer.
scripts/data/build_cross_lens_de_bank.py Uses shared fold_text, adjusts null token handling, adds provenance to exclusion report.
scripts/analyze/nearest_rows_baseline.py Writes CSVs via shared write_csv with fixed field lists; handles empty outputs.
scripts/analyze/cross_lens_three_lens_prompt.py Adds manifest provenance when writing CSVs via shared helper.
scripts/analyze/cross_lens_antonym_layers.py Routes through shared cross-lens/qwen_readout helpers; adds centering/device/k handling + manifest provenance.
scripts/analyze/analyze_wsd_sense_groups.py Uses shared WSD bundle helpers; fixes HF cache resolution; adds centering + provenance plumbing.
README.md Documents 0.2.1 consolidation, flags, and scope changes.
pyproject.toml Bumps version to 0.2.1; updates dependency comments.
docs/THIRD_PARTY.md Expands dependency/license notes; documents jlens sourcing and data notes.
docs/DATA.md Updates data layout description (removes data/wsd/ mention).
data/wsd/wsd_sense_anchors.json Removes unused optional WSD anchor file.
data/cross_lens/README.md Clarifies bank schema details and control/null semantics.
configs/sweeps/README.md Clarifies recipe-matching semantics and adds notes about omitted inert keys.
configs/sweeps/qwen35_2b_seedvar_base.yaml Adds log/checkpoint cadence keys (documented as cadence-only).
configs/sweeps/qwen35_0p8b_paper_topk_32x_k256_s2.yaml Clarifies recipe matching and adds cadence-only keys.
configs/sweeps/qwen35_0p8b_paper_topk_32x_k256_s1.yaml Clarifies recipe matching and adds cadence-only keys.
configs/sweeps/lowk_qwen35_2b_16x_base.yaml Adds cadence-only keys.
configs/sweeps/lowk_qwen35_0p8b_16x_base.yaml Adds cadence-only keys.
CITATION.cff Bumps version to 0.2.1 (release date retained).
Review details
  • Files reviewed: 52/53 changed files
  • Comments generated: 1
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines 789 to +793
cur = (
d.original_logit,
d.reconstructed_logit,
d.residual_term,
d.feature_contributions,
_full_contributions(d),
hematteo added a commit that referenced this pull request Aug 31, 2026
tests/test_wsd_sense_groups.py pinned SHA-256 hashes of full-precision
outputs (and a few float literals) captured on Apple Silicon; on the
ubuntu CI runners BLAS accumulation differs at ~1e-7 relative, so three
tests failed after PR #2 merged (the merge gate was mis-wired and let a
red build through). The pins are now fixture values in
tests/fixtures/wsd_pins.json compared with tolerances (rel 1e-5,
atol 1e-6; ints, ids and orderings exact), old-vs-new comparisons use
allclose, and the file passes under perturbed fixtures and different
PYTHONHASHSEED values. No script code changes.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants