Speaker 1: Jiafei Duan
Jiafei Duan is an incoming Assistant Professor at the National University of Singapore. He completed his Ph.D. in Computer Science & Engineering at the University of Washington, advised by Dieter Fox and Ranjay Krishna. His research centers on robotics foundation models, scalable data collection and generation, and grounding vision-language models in robotic reasoning. He has published at venues including ICLR, ICML, and RSS, and received the Ubiquitous Robots 2023 Best Paper Award.
Speaker 2: Haoquan Fang
Haoquan Fang is an incoming CS PhD student at Stanford University, where he will be advised by Fei-Fei Li at the Stanford Vision and Learning Lab. He also works with the PRIOR and Robotics teams at the Allen Institute for AI.
Abstract
Vision-language-action (VLA) models aim to serve as generalist robot controllers. In this talk, we present MolmoAct and MolmoAct2, a family of open robotics foundation models that introduce interpretable action reasoning pipelines, open datasets, embodied reasoning backbones, and adaptive-depth reasoning for faster closed-loop control. We discuss why open robotics foundation models matter for reproducible progress and practical real-world deployment.
Video
Coming soon. Stay tuned. :-)