Hey, I'm Yusef. My main focus is building Scopehaven: permission and outcome testing for AI apps and agents. I'm also a first-year student at the University of Toronto, intending Math + CS.
- 🔍 Scopehaven: checking who can access data or use a tool, comparing the agent's reply with observable outcomes, and rerunning checks after a repair. Current evidence comes from local tests and internal pilots with synthetic data. The product is early stage, and the code is private. See the approach and request a first test.
- 🧪 Agent evaluation: my eval lab studies tool failures, partial execution, and missing evidence. I'm interested in evaluations that distinguish an intended action from a verified result.
- 🧩 Open source: contributing focused fixes to evaluation and research tools, with regression tests and links to the upstream review.
- Make Scopehaven useful to external teams: reproduce a concrete permission failure, verify a repair, and preserve the check for future releases.
- Build evaluations that catch forbidden actions while checking that legitimate use still works, and report uncertainty when the evidence is incomplete.
- Strengthen my independent Python, algorithms, debugging, and mathematical foundations alongside my coursework.
Each card links to my contribution. Green labels indicate merged PRs; the PyTorch PR is still open as of October 8, 2026.
My merged work includes Apache Arrow slicing, Agent Lightning timeout-report retries, Inspect Scout postponed-annotation support, and W&B RAI Toolkit prediction-error accounting. The evidence page lists the verified contributions and the date checked.
- Aesthetics AI — my iOS workout-planning and tracking app. App Store · How it's built
- Tiraz — my privacy-first digital wardrobe app. App Store · How it's built
- CallReclaim — a private missed-call recovery MVP. Sample-data demo · How it's built
The public CallReclaim Agent Desk is a separate synthetic demo with no messaging backend. ShiftProof is also a local prototype using synthetic data.
Languages: Python · TypeScript · JavaScript · Swift · SQL
Apps: React · React Native · Expo · Next.js · SwiftUI · watchOS
Experiments: PyTorch · NumPy · Inspect · pytest
Infrastructure: PostgreSQL · SQLite · Supabase · Docker · GitHub Actions
The eval lab includes a frozen 624-trial local-model study. Invalid-output rates differed between conditions, so I withheld improvement claims and kept the descriptive results and missing-output analysis.
The Tiraz study uses garment annotations, not images. Its three seeded models achieved 58.5–59.0% top-1 against a 52.7% baseline on 962 held-out groups; removing context exposed a large drop in prediction-set coverage.
Methods, results, and limitations · PyTorch local validation supplement
If you're building an AI app or agent and want to discuss permission testing, evaluation failures, or a focused open-source contribution, you can reach me below.
Typing intro and compact grids inspired by DenverCoder1.