Replies: 2 comments 1 reply
|
Your work is highly valuable. However, Blogrxiv primarily accepts blog posts rather than project homepages, so it may not be suitable for inclusion. That said, I strongly recommend you rewrite your content into a blog post and submit it to Blogrxiv. |
0 replies
|
Hi hqKing,
Thank you very much for your thoughtful feedback and recommendation. I understand that BlogrXiv primarily features blog-style articles rather than project homepages.
I will rewrite the current content into a self-contained blog post that more clearly presents the motivation, methodology, key results, and takeaways of DriveMA, and then submit it to BlogrXiv for consideration.
Thanks again for your time and encouragement.
Best regards,
Weicheng Zheng
…-----原始邮件-----
发件人:hqKing ***@***.***>
发送时间:2026-07-26 15:10:56 (星期日)
收件人: OpenEnvision/BlogrXiv ***@***.***>
抄送: "Weicheng Zheng" ***@***.***>, Author ***@***.***>
主题: Re: [OpenEnvision/BlogrXiv] [Recommendation] DriveMA: Driving Vision-Language-Action Models with Verifiable Meta-Actions (Discussion #12)
Your work is highly valuable. However, Blogrxiv primarily accepts blog posts rather than project homepages, so it may not be suitable for inclusion. That said, I strongly recommend you rewrite your content into a blog post and submit it to Blogrxiv.
—
Reply to this email directly, view it on GitHub, or unsubscribe.
You are receiving this because you authored the thread.Message ID: ***@***.***>
|
1 reply
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Blog title
DriveMA: Driving Vision-Language-Action Models with Verifiable Meta-Actions
Canonical URL
https://tsinghua-mars-lab.github.io/DriveMA/
Are you the author?
Yes, I wrote this post
Author / Lab / Organization
Weicheng Zheng et al. — MARS Lab, IIIS, Tsinghua University
Source name
No response
Suggested BlogrXiv category
LLM & MLLM
Suggested tags
Autonomous Driving, VLA, Language-Action Alignment, RL
Short summary
DriveMA addresses the language–action gap in driving Vision-Language-Action models by introducing compact, verifiable meta-actions between visual observations and trajectory generation. It combines trajectory-grounded annotation, action-centric pretraining, and turn-level credit assignment reinforcement learning to make generated trajectories faithfully execute the model’s stated driving intent. DriveMA achieves state-of-the-art planning performance on WOD-E2E and competitive closed-loop results on NAVSIM.
Why is this technically valuable?
DriveMA makes intermediate language decisions verifiable by mapping predicted trajectories back into the meta-action space. Its turn-level RL assigns decision and trajectory rewards to their corresponding tokens, improving language–action consistency and enabling data-efficient, state-of-the-art driving performance.
Related paper / code, optional
https://github.com/Tsinghua-MARS-Lab/DriveMA
Submission confirmation
All reactions