Official code for the paper: "Simple LLM Baselines are Competitive for Model Diffing"
-
Updated
Feb 13, 2026 - Python
Official code for the paper: "Simple LLM Baselines are Competitive for Model Diffing"
ActDiff: diagnose, repair, and prevent persistent domain priors in narrowly finetuned vision-language models.
Code and artifacts for The Convergence Gap: when instruction-tuned models settle on next-token predictions.
Supplementary code & results for "Variant-specific crosscoder features are seed-stable but not detectably task-causal in a GRPO-LoRA math setting" (ICML 2026 Mech Interp Workshop, Spotlight)
Mechanistic study of contextual-integrity post-training in Qwen2.5-7B, testing whether improved privacy behavior comes from new mechanisms or better use of machinery already present in the base model.
Behavioral auditing of finetuned language models using blinded policy recovery, model diffing, and causal activation interventions.
Sealed, preregistered benchmark for black-box model-diffing agents: five LoRA finetunes of Qwen3.5-9B (one null, three planted behaviours, one dropped backdoor), audited blind by Neel Nanda's diffing-agent recipe and four cheaper conditions. The recipe fails by not asking, and the auditor itself is a failure mode. MATS 12 application.
CDD Diffing Technique [Contrastive Decoding Diffing]
Code and artifacts for Same Targets, Different Computation: how post-training divides work across model layers.
Research workspace for model diffing between pretrained and post-trained language models.
To associate your repository with the model-diffing topic, visit your repo's landing page and select "manage topics."