Empirical Research Framework for Superalignment: Can a weaker, imperfect supervisor reliably align a vastly superior intelligence without forcing the stronger model to inherit its supervisor's errors and hallucinations?
As AI models surpass human-level capabilities in complex reasoning, software engineering, and scientific proof discovery, human feedback (RLHF) transitions from an authoritative ground truth to a weak supervisory signal. Standard fine-tuning on weak supervision forces high-capacity student models to mimic supervisor mistakes.
W2S-Eval implements a confidence-gated supervisory alignment framework. By dynamically masking out low-confidence supervisor logits during fine-tuning, the strong student model uses its internal pre-trained representations to surpass its teacher—achieving measurable Weak-to-Strong Generalization.
Standard Kullback–Leibler (KL) divergence forces the student distribution
In W2S-Eval, we introduce a Confidence Gate
The resulting Confidence-Gated Alignment Loss is defined as:
Where:
-
$\tau \in [0, 1]$ represents the strictness threshold (default$\tau = 0.6$ ). - When
$P_{\text{weak}}$ is uncertain,$\mathbb{I}_{\tau}(x) = 0$ , preventing the strong student from learning noisy or hallucinated outputs.
To quantify how effectively a strong student model recovers performance beyond its weak supervisor's ceiling, we compute the Performance Gap Recovered (PGR):
| Benchmark Metric | Symbol | Description |
|---|---|---|
| Weak Supervisor Accuracy | Accuracy of the low-capacity teacher alone. | |
| Ground-Truth Ceiling | Accuracy of the strong student when trained directly on gold labels. | |
| Weak-to-Strong Accuracy | Accuracy of the strong student trained only on weak supervision via W2S-Eval. |
-
$\text{PGR} \le 0%$ : Student blindly imitated teacher errors. -
$\text{PGR} > 0%$ : Successful generalization—the student elicited latent capabilities beyond teacher intelligence.
w2s_eval/
│
├── w2s/ # Core Library
│ ├── __init__.py # Package initializers
│ ├── loss.py # Confidence-gated alignment loss
│ ├── supervisor.py # Synthetic model definitions (16-dim vs 512-dim)
│ └── trainer.py # PyTorch execution engine
│
├── tests/ # Unit Tests
│ └── test_loss.py # Numerical stability & gradient pass tests
│
├── run_experiment.py # Synthetic tensor pipeline execution
├── run_hf_experiment.py # Real Hugging Face model evaluation pass
├── pytest.ini # Pytest workspace configuration
├── README.md # Documentation
└── LICENSE # Open-source license
# Clone repository
git clone https://github.com/your-username/w2s-eval.git
cd w2s-eval
# Create and activate virtual environment
python -m venv venv
.\venv\Scripts\Activate.ps1
# Install core dependencies
pip install torch transformers datasets accelerate pytest
python -m pytest
python run_experiment.py
python run_hf_experiment.py
Below are benchmark results collected on a sentiment classification task comparing a weak supervisor (DistilBERT) against an aligned strong student (BERT-Base):
| Model Architecture | Training Signal | Accuracy | PGR Score |
|---|---|---|---|
| DistilBERT (Weak Supervisor) | Gold Labels | ||
| BERT-Base (Strong Student) | Naive Weak Labels ( |
||
| BERT-Base (W2S-Eval Aligned) | Gated Weak Labels ( |
||
| BERT-Base (Upper Bound) | Gold Labels |
If you build upon this project for research, please cite:
@software{w2s_eval_2026,
author = {Rahul},
title = {W2S-Eval: Confidence-Gated Weak-to-Strong Supervision Protocol for AI Alignment},
year = {2026},
publisher = {GitHub},
journal = {GitHub repository},
url = {https://github.com/rahuldesai101/w2s_eval}
}