ML/research engineering — verified environments, RLVR, agentic systems
Pinned Loading
-
Weak-Data-vs-RLVR
Weak-Data-vs-RLVR PublicControlled comparison of boosted SFT vs. GRPO at matched compute. i.e., which training loop escapes the problems a base model never solves. Reproduces Amin et al. (NeurIPS 2025) as the baseline. In…
Python
-
Chronic-Care-Agentic-AI-Companion
Chronic-Care-Agentic-AI-Companion PublicForked from Haris320/tc-pilot
Won 1st Place Overall in NYC Datadog Hackathon. Helping chronically ill patients decode their medical records & find matched clinical trials. Sponsored by Google DeepMind, Datadog, & ClickHouse. Al…
Python
-
-
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.



