A selection of robotics learning materials that I’ve personally read, used, and found worthwhile.
面向通用机器人策略(generalist policy)的 VLA 基础模型团队。研究列表:https://www.pi.website/blog
- π0, A Vision-Language-Action Flow Model for General Robot Control, 2024.10. [📄 Paper] [📝 Blog] [💻 Code] [🤗 Model]
- FAST, FAST: Efficient Action Tokenization for Vision-Language-Action Models, 2025.01. [📄 Paper] [📝 Blog] [🤗 Model]
- openpi, Open Sourcing π0, 2025.02. [📝 Blog] [💻 Code]
- Hi Robot, Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models, 2025.02. [📄 Paper] [📝 Blog]
- π0.5, π0.5: a Vision-Language-Action Model with Open-World Generalization, 2025.04. [📄 Paper] [📝 Blog]
- π0 + KI, Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better, 2025.05. [📄 Paper] [📝 Blog]
- RTC, Real-Time Execution of Action Chunking Flow Policies, 2025.06. [📄 Paper] [📝 Blog]
- π*0.6, π*0.6: a VLA that Learns from Experience, 2025.11. [📝 Blog]
- Human-to-Robot, Emergence of Human to Robot Transfer in VLAs, 2025.12. [📝 Blog]
- Memory VLA, VLAs with Long and Short-Term Memory, 2026.03. [📝 Blog]
- RLT, Precise Manipulation with Efficient Online RL, 2026.03. [📝 Blog]
- π0.7, π0.7: a Steerable Model with Emergent Capabilities, 2026.04. [📝 Blog]
openpi 仓库(Apache-2.0)包含 π0(flow-based VLA)、π0-FAST(自回归 VLA)、π0.5,基座 checkpoint 基于 10k+ 小时机器人数据预训练,提供 ALOHA / DROID / LIBERO 微调 checkpoint,支持 JAX 与 PyTorch。
面向通用人形机器人(generalist humanoid)的开放 VLA 基础模型。
- GR00T N1, GR00T N1: An Open Foundation Model for Generalist Humanoid Robots, 2025.03. [📄 Paper] [🌍 Website] [💻 Code] [🤗 Model]
- 历史分支:N1.5 | N1.6 | 模型合集 🤗 | 数据 Physical AI
- 架构:视觉-语言基础模型 + Diffusion Transformer 动作头;代码 Apache-2.0,权重 NVIDIA Open Model License。
- Cosmos-Reason1, Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning, 2025.03. [📄 Paper] [🌍 Website] [💻 Code]
从 Robotics Transformer 到 Gemini Robotics 的通用机器人基础模型序列。研究主页:https://deepmind.google/models/gemini-robotics/
- RT-1, RT-1: Robotics Transformer for Real-World Control at Scale, 2022.12. [📄 Paper] [🌍 Website] [💻 Code]
- PaLM-E, PaLM-E: An Embodied Multimodal Language Model, 2023.03, ICML 2023. [📄 Paper] [🌍 Website]
- RoboCat, RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation, 2023.06, TMLR. [📄 Paper] [📝 Blog]
- RT-2, RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, 2023.07, CoRL 2023. [📄 Paper] [🌍 Website] [📝 Blog]
- Q-Transformer, Q-Transformer: Scalable Offline RL via Autoregressive Q-Functions, 2023.09, CoRL 2023. [📄 Paper]
- Open X-Embodiment / RT-X, Open X-Embodiment: Robotic Learning Datasets and RT-X Models, 2023.10, ICRA 2024. [📄 Paper] [🌍 Website] [💻 Code]
- RoboVQA, RoboVQA: Multimodal Long-Horizon Reasoning for Robotics, 2023.11. [📄 Paper] [🌍 Website]
- RT-Trajectory, RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches, 2023.11. [📄 Paper] [🌍 Website]
- SARA-RT, SARA-RT: Scaling up Robotics Transformers with Self-Adaptive Robust Attention, 2023.12. [📄 Paper]
- AutoRT, AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents, 2024.01. [📄 Paper] [🌍 Website]
- ALOHA Unleashed, ALOHA Unleashed: A Simple Recipe for Robot Dexterity, 2024.10, CoRL 2024. [📄 Paper] [🌍 Website]
- Gemini Robotics, Gemini Robotics: Bringing AI into the Physical World, 2025.03. [📄 Paper] [🌍 Website] [📝 Blog]
- Gemini Robotics On-Device, 首个可微调的端侧 VLA, 2025.06. [📝 Blog]
- Gemini Robotics 1.5, Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer, 2025.10. [📄 Paper] [📝 Blog]
世界模型 Genie 系列(用于具身训练环境生成):Genie Generative Interactive Environments, 2024.02, ICML 2024 Best Paper [📄 Paper] | Genie 2, 2024.12 | Genie 3, 2025.08。
面向通用人形机器人的 System 1 / System 2 VLA 模型 Helix。技术更新以博客为主:https://www.figure.ai/news
- Helix, Helix: A Vision-Language-Action Model for Generalist Humanoid Control, 2025.02. [📝 Blog]
- Helix Logistics, Accelerating Real-World Logistics, 2025.02. [📝 Blog]
- RL Walking, Natural Humanoid Walk Using Reinforcement Learning, 2025.03. [📝 Blog]
- Scaling Helix, A New State of the Art in Humanoid Logistics, 2025.06. [📝 Blog]
- Project Go-Big, Internet-Scale Humanoid Pretraining and Direct Human-to-Robot Transfer, 2025.09. [📝 Blog]
- Helix 02, Introducing Helix 02: Full-Body Autonomy, 2026.01. [📝 Blog]
Helix 采用双系统架构:7B VLM 骨干(System 2,7-9 Hz)+ 80M 交叉注意力 Transformer(System 1,200 Hz),端到端训练,覆盖 35-DoF 上半身动作空间。
Tesla 人形机器人 Optimus 沿用与 FSD 一致的纯视觉端到端神经网络技术栈,信息主要来自 AI Day 演示与产品页,暂无正式论文。
- Optimus, Tesla 通用人形机器人, AI Day 2021.08 首次发布,2022.09 展示原型。[🌍 Website]
面向通用机器人的具身基础模型团队,随物理交互数据规模扩展。博客:https://generalistai.com/blog
- GEN-0, Embodied Foundation Models That Scale with Physical Interaction, 2025.11. [📝 Blog]
- GEN-1, Scaling Embodied Foundation Models to Mastery, 2026.04. [📝 Blog]
- Beyond World Models, Going Beyond World Models & VLAs, 2026.04. [📝 Blog]
Qwen 团队的多模态与具身相关工作。组织主页:https://github.com/QwenLM
- Qwen-RobotManip, 机器人操作(Manipulation)官方仓库. [💻 Code]
- Qwen-RobotNav, 机器人导航(Navigation)官方仓库. [💻 Code]
- Qwen-AgentWorld, Language World Models for General Agents(语言世界模型,Apache-2.0). [💻 Code]
- Qwen3-VL, 系列最强视觉-语言模型:3D grounding / 空间推理 / 具身 AI,可作为 VLA 视觉-语言底座. [💻 Code] [🤗 Model]
- Qwen2.5-VL, Qwen2.5-VL Technical Report, 2025.02. [📄 Paper] [📝 Blog] [💻 Code] [🤗 Model]
ByteDance Research / Seed 的视频生成预训练 + 机器人操作工作。
- RoboFlamingo, Vision-Language Foundation Models as Effective Robot Imitators, 2023.11, ICLR 2024 Spotlight. [📄 Paper] [💻 Code]
- GR-1, Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation, 2023.12, ICLR 2024. [📄 Paper]
- GR-2, GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation, 2024.10. [📄 Paper]
与 OpenDriveLab 合作的大规模操作平台与世界模型。组织主页:https://github.com/AgibotTech
- GO-1 / AgiBot World, AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems, 2025.03, IROS 2025. [📄 Paper] [🌍 Website] [💻 Code] [📊 Dataset]
- Genie Envisioner, Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation, 2025.08. [📄 Paper] [💻 Code]
- WholebodyVLA, Towards Unified Latent VLA for Whole-body Loco-manipulation Control, ICLR 2026. [💻 Code]
灵初智能(PsiBot)的灵巧操作 VLA 工作。组织主页:https://github.com/Psi-Robot
- DexGraspVLA, DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous Grasping, 2025.02, AAAI 2026 Oral. [📄 Paper] [💻 Code]
- Awesome-VLA-Papers, VLA 论文合集(含 Action Tokenization 综述). [💻 Code]
Psi R0 / R0.5 / R1 为公司产品级模型(快慢双脑 + Chain-of-Action-Thought),以官网发布为准:https://www.psibot.ai。
端到端具身基础模型 WALL 系列。官网:https://x2robot.com
宇树开源的机器人学习 / 强化学习与 VLA 代码库(Go2 / H1 / G1)。组织主页:https://github.com/unitreerobotics
- unitree_rl_gym, Isaac Gym 强化学习示例(Go2/H1/G1 运动控制). [💻 Code]
- unitree_rl_lab, 基于 IsaacLab 的强化学习实现. [💻 Code]
- unitree_mujoco, MuJoCo 仿真与 sim-to-real(C++/Python). [💻 Code]
- xr_teleoperate, XR 设备遥操作与数据采集. [💻 Code]
- unitree_lerobot, 端到端具身智能:LeRobot 策略训练/推理与部署. [💻 Code]
- unifolm-vla, 面向人形操作的视觉-语言-动作大模型. [💻 Code]
- unifolm-world-model-action, 开源世界模型-动作架构. [💻 Code]
- RT-2, RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, 2023.07, CoRL 2023. [📄 Paper] [🌍 Website]
- Open X-Embodiment, Open X-Embodiment: Robotic Learning Datasets and RT-X Models, 2023.10, ICRA 2024. [📄 Paper] [🌍 Website] [💻 Code]
- Octo, Octo: An Open-Source Generalist Robot Policy, 2024.05, RSS 2024. [📄 Paper] [🌍 Website] [💻 Code] [🤗 Model]
- OpenVLA, OpenVLA: An Open-Source Vision-Language-Action Model, 2024.06, CoRL 2024. [📄 Paper] [🌍 Website] [💻 Code] [🤗 Model]
- RDT-1B, RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation, 2024.10, ICLR 2025. [📄 Paper] [🌍 Website] [💻 Code] [🤗 Model]
- RoboBrain, RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete, 2025.02, CVPR 2025. [📄 Paper] [🌍 Website] [💻 Code] [📊 Dataset]
- GO-1 / AgiBot World, AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems, 2025.03. [📄 Paper] [🌍 Website] [💻 Code] [📊 Dataset]
- LeRobot, 端到端机器人学习库(集成 SmolVLA / ACT / Diffusion Policy / π0 移植). [💻 Code] [🤗 Model]
- SmolVLA, SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics, 2025.06. [📄 Paper] [💻 Code]
- ACT / ALOHA, Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware, 2023.04, RSS 2023. [📄 Paper] [🌍 Website] [💻 Code]
- Mobile ALOHA, Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation, 2024.01. [📄 Paper] [🌍 Website] [💻 Code]
- Diffusion Policy, Diffusion Policy: Visuomotor Policy Learning via Action Diffusion, 2023.03, RSS 2023. [📄 Paper] [🌍 Website] [💻 Code]
- 3D Diffusion Policy (DP3), 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations, 2024.03, RSS 2024. [📄 Paper] [💻 Code]
- Open-TeleVision, Open-TeleVision: Teleoperation with Immersive Active Visual Feedback, 2024.07, CoRL 2024. [📄 Paper] [🌍 Website] [💻 Code]
- Prismatic VLMs, Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models, 2024.02, ICML 2024. [📄 Paper] [💻 Code]
- legged_gym, Learning to Walk in Minutes Using Massively Parallel Deep RL, 2021.09, CoRL 2021. [📄 Paper] [💻 Code]
- HumanPlus, HumanPlus: Humanoid Shadowing and Imitation from Humans, 2024.06, CoRL 2024. [📄 Paper] [🌍 Website] [💻 Code]
- H2O / OmniH2O, H2O: Learning Human-to-Humanoid Real-Time Whole-Body Teleoperation, 2024.03, IROS 2024. [📄 Paper] [🌍 Website] [💻 Code]
- Genesis, 通用物理 / 生成式机器人仿真平台, 2024.12. [💻 Code] [📝 Blog]
- Isaac Lab, NVIDIA 基于 Isaac Sim 的机器人学习框架. [💻 Code] [📝 Blog]
- ManiSkill3, ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AI, 2024.10. [📄 Paper] [💻 Code]
- robosuite, robosuite: A Modular Simulation Framework and Benchmark for Robot Learning. [📄 Paper] [🌍 Website] [💻 Code]
- LIBERO, LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning, 2023.06, NeurIPS 2023. [📄 Paper] [🌍 Website] [💻 Code]
- RLBench, RLBench: The Robot Learning Benchmark & Learning Environment, 2019.09, RA-L 2020. [📄 Paper] [💻 Code]
- RoboCasa, RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots, 2024.06, RSS 2024. [📄 Paper] [🌍 Website] [💻 Code]
- CALVIN, CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks, 2021.12, RA-L 2022. [📄 Paper] [💻 Code]