Interpretability and efficient reasoning in language models. Previously an ML Engineering Intern at Shopify, building and evaluating support-triage models.
Computer Science, University of Waterloo · website · linkedin
| The Truth Was Never Gone: Perfect Aliasing in Compliant-Context Truth Probes Sole author |
In the studied binary tasks, compliant truth and action labels coincide, leaving the probes unidentified. Complementary evaluation labels give AUROC(action) = 1 − AUROC(truth). Mixed-context fitting reaches 1.000 versus 0.006 conventional AUROC on Gemma-9B, averaged over three seeds. code · arXiv |
| The Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoning First author · COLM 2026, Efficient Reasoning workshop · Spotlight |
A causal halt direction at layer 18, moved into the weights. Hook-free per-problem self-halt: −24% thinking at held accuracy, five unseen benchmarks, 24 problems, no RL. code · arXiv |
| Deep Learning Model for Invasive Ductal Carcinoma Detection Sole author · IEEE CCECE 2025 |
Deep model for invasive ductal carcinoma detection. repo · IEEE Xplore |
| Human Action Detection using FMCW mmWave Radar First author · CVIS 2024 |
Action recognition off raw radar returns. JCVIS |
| Off-Axis Drift: Internalizing a Halt Direction Needs More Than Its Scalar Projection First author · with Tinuade Adeleke |
Evaluated activation-target training for hook-free early stopping across 1.5B–14B reasoning models, measuring compression, accuracy costs, and cross-domain transfer. Completed a 376-job development comparison of 46 candidate recipes, selecting full-vector and on-axis reconstruction pairs at three shortening targets; held-out validation is pending. |
| RL training dynamics | Dense-checkpoint probing across three historical GRPO seeds shows a gold-free confidence monitor failing to flag a length-penalty reward hack: it holds at 0.71–0.85 while held-out accuracy halves to 0.38–0.48. A four-arm factorial testing sensitivity to GRPO normalization is implemented and pre-registered, not yet run. |
| LOB-Engine | A small C++20 matching-engine prototype exploring price-time priority, SPSC queues, and pooled order storage. |
| RAG Tradeoffs | Benchmarks retrieval accuracy and latency across context lengths, chunk sizes, and 10+ LLMs. |
| Firefighter Robot | Autonomous maze-solving flame extinguisher. Set course records on two mazes, with recorded demonstrations. |
| Mr. Nutz | Poker robotics combining card perception, poker logic, and integrated hardware. |

