Does clinical co-occurrence strength predict when a medical VLM's internal representation of a finding is decodable but not causally used? Interpretability study on CheXagent-2-3b using activation probing and patching.
deep-learning medical-imaging radiology causal-inference interpretability probing chest-xray hallucination mechanistic-interpretability vision-language-model activation-patch explainable-algorithms
-
Updated
Aug 20, 2026 - Python