Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Formal verification of refusal-direction ablation in Qwen2.5-1.5B-Instruct using auto_LiRPA's CROWN bounds.

## Files

- model\_io.py — model + tokenizer loading

- refusal.py — refusal direction extraction (Arditi et al. method)

- verification.py — CROWN wrappers, MLP ablation, projection bounds

- data\_io.py — dataset loaders (JailbreakBench, Alpaca)

- run.py — gate-by-gate experiment runner

- data/ — saved JSON results and singular value plot

About

Proving robustness of LLM refusal mechanisms using formal neural network verification

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages