PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models
Published 15 Sept 2026arXiv:2609.14973
Updated 12 h ago · first seen 14 Sept 2026
paper_01M2HN8YAGY873NNH1W0YP4YD3
Abstract
We present PhysBrain 1.5, a unified model for understanding physical environments, generating actions, and predicting future states. Motivated by the physical loop of observation, interaction, and environmental change, we bring these capabilities into a common learning framework. Starting from a general vision--language model, we encode language responses, end-effector motion, and dense visual targets as discrete sequences and jointly optimize them with autoregressive next-token prediction. Pre-training draws its embodied supervision entirely from human interaction videos, using task-centered episodes to pair semantic and spatial context with recovered motion and subsequent observations. We then adapt the model through supervised fine-tuning on a mixture of human demonstrations, robot trajectories, and simulated experience. Across 28 embodied understanding benchmarks, our 8B model achieves an average score of 72.5, setting a new open-source state of the art and performing on par with leading proprietary models such as GPT-6-Astra and Gemini 3.6 Flash. It achieves the best open-source results on 14 benchmarks while retaining general multimodal capabilities. Beyond these understanding evaluations, qualitative examples show the model's ability to produce end-effector trajectories and predict future scenes through spatially aligned RGB, depth, and robot-mask outputs.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 3
PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models: published at changed from 2026-09-14T00:00:00+00:00 to 2026-09-15T04:00:00+00:00
Published14 Sept 2026→15 Sept 2026arxivPhysBrain 1.5: From Vision-Language Models to Physical Foundation Models: authors changed from ["Changti Wu", "Chaoyi Ruan", "Chenliu Hao", "Cong Huang"… to ["Changti Wu", "Chaoyi Ruan", "Chenliu Hao", "Cong Huang"…
AuthorsChangti Wu, Chaoyi Ruan, Chenliu Hao, Cong Huang, DeepCybo Team, Haibao Liu, Haipeng Cao, Hang Yuan, Hanwen Zhang, Hao Wu, Haochen Liu, Haoyang Ge, Hong Li, Jiyan He, Kai Chen, Kai Hu, Kailin Deng, Peize Li, Peng Ren, Qiuzhi Liu, Qiyuan Su, Ruimeng Zhang, Ruoqi Yang, Shengcai Liu, Shijie Lian, Shuo Ren, Tao Luo, Tuopusen Huang, Xiaopeng Lin, Xiaotong Fu, Xueyin Xu, Xuguo He, Yakun Hou, Yao Zhang, Yibo Zhang, Yichao Du, Yining Wang, Youning Chen, Yu Bin, Yu Huang, Yukun Shi, Yun Lin, Yunlong Guo, Yuxiang Zhang, Yuxuan Tian, Zhaolong Shen, Zhaoyang Yang, Zhaoyang Zeng, Zheng Chang, Zhiqiang Liu→Changti Wu, Chaoyi Ruan, Chenliu Hao, Cong Huang, DeepCybo Team, Haibao Liu, Haipeng Cao, Hang Yuan, Hanwen Zhang, Hao Wu, Haochen Liu, Haoyang Ge, Hong Li, Jiyan He, Kai Chen, Kai Hu, Kailin Deng, Peize Li, Peng Ren, Qiuzhi Liu, Qiyuan Su, Ruimeng Zhang, Ruoqi Yang, Shengcai Liu, Shijie Lian, Shuo Ren, Tao Luo, Tuopusen Huang, Xiaopeng Lin, Xiaotong Fu, Xueyin Xu, Xuguo He, Yakun Hou, Yao Zhang, Yibo Zhang, Yichao Du, Yining Wang, Youning Chen, Yu Bin, Yu Huang, Yukun Shi, Yun Lin, Yunlong Guo, Yuxiang Zhang, Yuxuan Tian, Zhaolong Shen, Zhaoyang Yang, Zhaoyang Zeng, Zheng Chang, Zhiqiang Liu, Zhirui Zhang, Zishen Zhuang, Ziyi Zhang, Zubin ZhengarxivNew paper: PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models
huggingface
Sources
Sources 2
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.