On August 27, 2026, Guangzhou, Xiaopeng Group held a Physics AI Sharing and Second Generation VLA New Version Experience Day event with the theme of "TIME Time". Xiaopeng's second generation VLA big model received its first major upgrade, and the new XOS 6.3.0 version will be globally released on Xiaopeng G9L. The core of this model upgrade is to enable AI to truly understand time: moving from static 3D spatial understanding of the physical world in the past to dynamic 4D spatiotemporal understanding.
Xiaopeng Group continues to validate Scaling Law in the field of autonomous driving, bringing significant improvements to experience, safety, and generalization. The upcoming release of the all-new XOS 6.3.0 major version has increased the parameter count of the end side model by 3.5 times, opening up an order of magnitude gap with common small models in the industry. The ultra long temporal context memory ability enables AI drivers to truly achieve "accurate reading, thinking ability, and fast response", promoting the transition of Xiaopeng Physics AI base model from spatial understanding to spatiotemporal understanding at the bottom level. The new version also introduces the whole vehicle brain Master Agent, achieving the integration of VLA+VLM cockpit, and downgrading some Robotaxi L4 level experiences to mass-produced models, bringing a new experience to the vast number of users.
From understanding space to understanding spacetime, the second generation of VLA models begins to possess a sense of time
【2026年8月27日,广州】今日,小鹏集团举办以“TIME 时间”为主题的物理AI分享暨第二代VLA全新版本体验日活动,小鹏第二代VLA大模型迎来首次重大升级,全新XOS 6.3.0版本将于小鹏G9L全球首发。此次模型升级的核心,是让AI开始真正理解时间:从过去对物理世界的静态3D空间理解,进一步走向动态4D时空理解。
小鹏集团持续在自动驾驶领域验证Scaling Law,给体验、安全和泛化性带来巨大的提升。即将发布的全新XOS 6.3.0大版本,端侧模型参数量扩大了3.5倍,与业内常见小模型拉开了数量级差距,超长时序的上下文记忆能力,让AI司机真正做到了“看得准,会思考,反应快”,推动小鹏物理AI基座模型实现从空间理解到时空理解的底层能力跃迁。全新版本还引入了整车大脑Master Agent,实现VLA+VLM的驾舱融合,并将部分Robotaxi的L4级体验下放到量产车型,为广大用户带来全新体验。
从理解空间到理解时空,第二代VLA大模型开始拥有“时间感”
The real challenge of physical AI is to understand the world and act in a restricted real world. The real world is not static images, but a constantly changing and evolving process. To understand the world like humans, models must first understand what "time" is.
Traditional models can recognize vehicles, pedestrians, roads, and traffic signals, but to truly make complex decisions in the physical world, simply "seeing" is far from enough. Vehicles need to know what may have happened in the past, present, and future. Therefore, the core of this second-generation VLA upgrade is to establish the model's understanding of the time dimension for the first time. The Xiaopeng physics world base model incorporates "time" into the model for the first time, allowing AI to understand the dimensions of the world and transition from "3D space" to "4D spacetime".
Previously, models saw more current information, but the second-generation VLA introduced the Infini VLA long temporal architecture. Remembering the first 30 seconds of the world, driving judgments began to have longer context and more effective information entered the model. Continuous road events are no longer fragmented into independent images, and models can connect previous actions, locations, and scenes to form a more complete temporal understanding.
物理AI真正的难题,是理解世界,并在充满限制的真实世界里行动。真实世界并不是一张张静止的图片,而是一个不断变化、连续演进的过程。模型想要像人类一样理解世界,必须首先理解什么是“时间”。
传统的模型可以识别车辆、行人、道路和交通信号,但要真正完成复杂的物理世界决策,仅仅“看见”还远远不够,车辆需要知道过去、现在、未来可能发生什么。因此,此次第二代VLA升级的核心,是让模型第一次建立起对时间维度的理解。小鹏物理世界基座模型,首次把“时间”纳入模型,AI理解世界的维度,从“3D空间”跃迁至“4D时空”。
以前模型看到的更多是当下的信息,第二代VLA全新引入Infini-VLA长时序架构,记得前30秒的世界,驾驶判断开始拥有更长的上下文,更多有效信息进入模型。连续发生的道路事件不再被割裂成一个个独立画面,模型可以将此前发生的动作、位置和场景串联起来,形成更加完整的时序理解。
The road conditions are constantly changing, and the model also needs to think while watching. If the decision-making speed cannot keep up with the real changes, it cannot form a truly reliable physical world intelligence. So, the second-generation VLA synchronously improves the efficiency of model inference by adopting Streaming Inference, moving from discrete inference to continuous inference, and increasing end-to-end response speed by 300%. In the face of constantly changing road conditions, the traditional model is a "look calculate output look again" mode, while the second-generation VLA can achieve parallel processing of "look, think, and act at the same time". When encountering unexpected situations such as sudden braking of the preceding vehicle, pedestrian turning back, and forced congestion of the adjacent vehicle, the vehicle's response is as skillful and agile as an experienced driver.
On real roads, the preceding vehicle may suddenly stop, pedestrians may temporarily change direction, and the adjacent vehicle may suddenly become congested. Predicting more possibilities in the future is necessary to choose a better course of action. The model not only needs to recognize past and current image information, but also needs to make predictions based on the position, speed, motion trend, and road environment of the target object. Xiaopeng X-Foresight predicts that the world model will board the car for the first time and can predict what will happen within 6 seconds, deducing the possible behaviors of surrounding traffic participants in the future.
Remembering the past, understanding the present, and predicting the future, Xiaopeng's physical world base model has come to understand time. At the same time, the new version of the model introduces a MoT hybrid architecture, which dynamically allocates model capabilities for different tasks, reduces task interference between different scenarios such as urban roads, parks, and parking, and assigns different problems to different "experts", allowing the model to switch more stably between different driving tasks.
道路情况是时刻在变化的,模型也需要边看边想,如果决策速度跟不上现实变化,也无法形成真正可靠的物理世界智能。所以,第二代VLA同步提升了模型推理效率,采用Streaming Inference(流式自回归推理),从离散推理走向连续推理,端到端响应速度提升300%。面对不断变化的道路状况,传统的模型是“看 → 算 → 输出 → 再看”的模式,而第二代VLA能够做到“边看、边想、边行动”的并行处理。遇到前车突然刹停、行人折返、旁车强行加塞等突发情况,车辆的反应像老司机一样娴熟、敏捷。
在真实道路上,前车可能突然刹停,行人可能临时改变方向,旁车可能突然加塞,预测更多种未来,才能选择更好的行动。模型不仅要识别过往画面和当前画面信息,还需要根据目标物的位置、速度、运动趋势、道路环境等进行预测。小鹏X-Foresight预测世界模型首次上车,可预判6秒内发生什么,推演周边交通参与者未来的可能行为。
记住过去、理解现在、预判未来,小鹏物理世界基座模型从此理解了时间。与此同时,新版模型引入MoT混合架构,通过针对不同任务动态分配模型能力,减少城区道路、园区、泊车等不同场景之间的任务干扰,把不同的问题交给不同的“专家”,让模型在不同驾驶任务之间更加稳定地切换。
Why do we need a larger model? A larger number of parameters is necessary to achieve strong enough generalization ability, making it possible to quickly enter the global market. The new version of Xiaopeng's second-generation VLA has increased the parameter count of the end side model by 3.5 times, which is more than 15 times that of mainstream VLA. Combined with advanced trajectory prediction, ultra long effective time series, and faster decision response, it has increased the multi-dimensional comprehensive security capability by 20 times.
The all-new version of the second-generation VLA not only upgrades intelligent assisted driving, but Xiaopeng chooses to use it as a robot to reconstruct the car's machine brain, launching the Master Agent for the first time, achieving the integration of VLA and VLM in the cockpit, allowing vehicles to further move from "understanding a sentence" to "understanding what users really want to accomplish". The Master Agent is built on top of Xiaopeng's self-developed Omni multimodal model, which not only understands natural language and fuzzy semantics, but also automatically breaks down user intentions into tasks and schedules vertical agents such as intelligent driving, chassis, cockpit, and body control to collaborate and execute, achieving a closed-loop from "understanding" to "action". This means that the Master Agent provides the entire vehicle with a unified brain for the first time, and also enables the vehicle's intelligence to move from responding to a single instruction to further autonomous understanding of intent, task planning, and collaborative execution.
It is worth mentioning that this new version upgrade also brings the functions of "voice side parking" and "voice controlled nearby parking". Users only need to give parking instructions through voice, and the vehicle can autonomously find a suitable position and complete side parking in intelligent driving mode, achieving a complete closed-loop from "voice instructions" to "autonomous execution". The L4 level capability of Xiaopeng Robotaxi is being further transferred to mass-produced cars through homologous technology, bringing car owners a higher-level and more natural driving experience.
“为什么我们需要一个更大的模型?”更大的参数量才能实现足够强的泛化能力,让快速进入全球市场成为可能。小鹏第二代VLA全新版本端侧模型参数量提升3.5倍,是主流VLA的15倍以上,结合超前轨迹预判、超长有效时序、更快决策响应,使得多维综合安全能力提升20倍。
第二代VLA全新版本升级的不只是智能辅助驾驶,小鹏选择用做机器人的方式重构车机大脑,首次推出Master Agent,实现VLA与VLM的驾舱融合,让车辆从“听懂一句话”进一步走向“理解用户真正想完成什么”。Master Agent建立在小鹏自研Omni全模态模型之上,不仅能够理解自然语言和模糊语义,还能将用户意图自动拆解为任务,并调度智驾、底盘、座舱、车身控制等垂直Agent协同执行,实现从“理解”到“行动”的闭环。这意味着,Master Agent让整车第一次拥有了一个统一大脑,也让车辆智能从响应单一指令,进一步迈向自主理解意图、规划任务和协同执行。
值得一提的是,本次全新版本升级还带来了“语音靠边停车”、“语音控制就近泊入”功能,用户只需通过语音下达停车指令,车辆即可在智驾状态下自主寻找合适位置并完成靠边停车,实现从“语音指令”到“自主执行”的完整闭环。小鹏Robotaxi的L4级能力,正通过同源技术进一步下放至量产车,为车主带来更高阶、更自然的用车体验。
A set of AI base connects cars and robots, physical AI enters the stage of large-scale reuse
Xiaopeng Group is positioned as a "travel explorer in the world of physical AI, a embodied intelligence company facing the world". It is not only a car company that can do autonomous driving and build cars, but also a physical AI company. The true value of the physical world AI foundation is not just to do a single task better, but to transfer and reuse the abilities learned by models in the real world to more tasks and ontologies, promoting physical AI from a single product capability to a universal capability.
The second generation of Xiaopeng VLA is based on a unified technology base and has achieved L2 to L4 capability connectivity, further expanding towards higher-level autonomous driving. Robotaxi, equipped with second-generation VLA, has recently obtained the qualification for remote testing of intelligent connected vehicles in Guangzhou. It can conduct road tests without a safety officer on relevant first, second, and third level test roads in Guangzhou, which means that Xiaopeng has entered the key road verification stage of autonomous driving.
一套AI底座打通汽车和机器人,物理AI进入规模化复用阶段
小鹏集团定位为“物理AI世界的出行探索者,面向全球的具身智能公司”,不仅是能做自动驾驶的车企、能造车的自动驾驶公司,更是一家物理AI公司。物理世界AI基座真正的价值,不只是把单一任务做得更好,而是将模型在真实世界中学习到的能力迁移、复用到更多任务和更多本体,推动物理AI从单一产品能力走向通用能力。
小鹏第二代VLA基于统一技术底座,已实现L2至L4能力贯通,进一步向更高阶自动驾驶拓展。搭载第二代VLA的Robotaxi近日获得广州市智能网联汽车远程测试资质,可在广州相关一、二、三级测试道路开展主驾无安全员道路测试,这意味着小鹏开始进入主驾无人的关键道路验证阶段。
In addition to automobiles, this capability is also being applied to more entities such as robots, which is the core significance of Xiaopeng's continuous promotion of a unified physical AI base model. The Xiaopeng Universal Humanoid Robot IRON is equipped with three Turing AI chips, with an effective computing power of up to 2250 TOPS. Relying on the industry-leading computing power foundation, the Xiaopeng Physical World Base Model has achieved end-to-end deployment, and can independently complete complex work tasks without remote operation, while ensuring low latency and data security for inference.
On August 24th, Xiaopeng Robotics completed its first round of equity financing, raising over $900 million and achieving a post investment valuation of over $6.3 billion, setting a new record for single round private equity financing in China's embodied intelligence industry. This round of financing is led by IDG Capital and participated by Gaorong Venture Capital, with support from two strategic investors, Tencent and Alibaba. It fully reflects the high recognition of the capital market for Xiaopeng Group's leading position, technological roadmap, mass production capability, and long-term commercial value in the field of physical AI. It also further verifies Xiaopeng's technology roadmap and industrialization ability extending from automobiles to robots.
The significance of the continuous evolution of physical AI base models is not only to make high-order models stronger, but also to enable advanced intelligent driving capabilities to cover more car models, benefit more old car owners, and achieve technological inclusiveness. Through learning token compression and distillation training, Xiaopeng achieves high-quality visual understanding with fewer effective tokens without sacrificing input information and model capacity. The second-generation VLA base model capability is deployed to platforms with lower computing power. With the enhancement of base model capability, the distilled Turing VLA 2.0 Lite also achieves significant capability improvement, allowing more users to enjoy a safer and more reliable urban assisted driving experience. The first batch of Turing VLA 2.0 Lite is expected to be launched in September, and the Xiaopeng G9L Max version will be equipped with this distilled version for the first time.
除了汽车,这一能力也正在向机器人等更多本体落地,这也是小鹏持续推进统一物理AI基座模型的核心意义。小鹏通用人形机器人IRON搭载3颗图灵AI芯片,有效算力高达2250 TOPS,依托行业领先的算力基础,小鹏物理世界基座模型已实现端侧部署,无需远程遥操作即可自主完成复杂工作任务,同时保障推理低时延与数据安全。
8月24日,小鹏机器人业务完成首轮股权融资,融资超9亿美元、投后估值超63亿美元,刷新中国具身智能行业单轮私募股权融资纪录。本轮融资由IDG资本领投、高榕创投参投,并获得腾讯、阿里巴巴两家战略投资者支持,充分体现了资本市场对小鹏集团在物理AI领域的领先地位、技术路线、规模量产能力和长期商业价值的高度认可,也进一步验证了小鹏从汽车向机器人延伸的技术路线和产业化能力。
物理AI基座模型持续进化的意义,不只是让高阶模型变得更强,也在于让先进智驾能力覆盖更多车型、惠及更多老车主,实现科技普惠。小鹏通过学习式Token压缩与蒸馏训练,在不牺牲输入信息和模型容量的基础上,以更少的有效Token实现高质量视觉理解,将第二代VLA基座模型能力部署到算力更低的平台,伴随基座模型能力的增强,蒸馏后的图灵VLA 2.0 Lite同样获得显著能力提升,让更多用户享受到更安全、更可靠的城区辅助驾驶体验。图灵VLA 2.0 Lite预计9月开启首批推送,小鹏G9L Max版本将首发搭载这一蒸馏版本。
While continuously improving model performance, Xiaopeng is making every effort to promote the global deployment of the second generation VLA. Recently, Xiaopeng completed the localization acceptance test of the second generation VLA in Germany. This model, trained on Chinese data, has a highly similar actual experience on urban roads in Germany to that in China, with almost no increase in localized training data. Xiaopeng aims to obtain regulatory approval in Europe in the first half of next year and gradually deliver the second-generation VLA to overseas users, promoting the global accessibility of advanced intelligent assisted driving.
Physical AI is a long-term project, and the moat lies in the long-term sedimentation and compound interest of AI Infra
Physical AI is a long-term project, and the continuous evolution of model capabilities relies on a complete set of AI Infra that has been built for a long time. The real competition is not just about how big the model parameters are or how strong the single capability is, but about who can continuously build the infrastructure that supports the evolution of the model and transform every real-world feedback into the next round of capability improvement.
在持续提升模型性能的同时,小鹏正全力推进第二代VLA的全球化部署。近期,小鹏在德国完成了第二代VLA的本地化验收测试。这套基于中国数据训练的模型,在几乎不增加本地化训练数据的情况下,在德国城市道路的实际体验与国内高度接近。小鹏目标于明年上半年率先在欧洲获得监管许可,并陆续向海外用户交付第二代VLA,推动高阶智能辅助驾驶的全球普惠。
物理AI是长期工程,护城河在于AI Infra的长期沉淀与复利
物理AI是一场长期工程,模型能力的持续进化,背后依赖的是一整套长期建设的AI Infra。真正的竞争并不只是模型参数有多大、单次能力有多强,而是谁能够持续建设支撑模型进化的基础设施,并把每一次真实世界的反馈转化为下一轮能力提升。
In the past few years, Xiaopeng has continuously invested in building a complete AI Infra system that covers data, training, simulation, evaluation, deployment, and computing power. With autonomous driving as the first large-scale application scenario, it has gradually established a complete closed loop from real-world data collection, model training, testing, deployment, to user feedback and continuous optimization. Unlike digital AI, which mainly relies on Internet data, physical AI data comes from the real world, and models must continue to interact, learn, and verify with the world in the real environment. Therefore, to do a good job in physics AI, in addition to the model itself, two core issues must be addressed: how to efficiently use real-world data, and how to efficiently validate the behavior of the model in the real world.
Xiaopeng relies on a fleet of millions of vehicles and billion level data assets to continuously build a data flywheel. Through unified storage, multimodal retrieval, and large-scale mining, it transforms massive real-world data into data assets for sustainable training and optimization models. Currently, the model's single training data throughput has reached 100 million clips, a tenfold increase compared to six months ago. By structuring, labeling, and semantically processing data, Xiaopeng can efficiently discover unknown problems from massive real-world data and quickly feed back model training and iteration, forming a positive cycle of more data, more thorough problem discovery, faster model evolution, and more data generation.
The data scale is growing, the efficiency of model training is improving, the computing power and toolchain are continuously improving, and more and more real-world problems can be included in model iteration. The moat of data scale is gradually forming. Xiaopeng Group's long-term investment in the AI Infra field has gradually accumulated into a basic ability for sustainable reuse, showing obvious economies of scale and beginning to generate compound interest.
过去几年,小鹏持续投入建设覆盖数据、训练、仿真、评估、部署和算力的完整AI Infra体系,以自动驾驶作为第一个规模化应用场景,逐步建立起从真实世界数据采集,到模型训练、测试、部署,再到用户反馈和持续优化的完整闭环。与数字AI主要依赖互联网数据不同,物理AI的数据来自真实世界,模型必须在真实环境中持续与世界交互、学习和验证。因此,做好物理AI,除了模型本身,还必须解决两个核心问题:如何高效使用真实世界的数据,以及如何高效验证模型在真实世界中的行为。
小鹏依托百万车队和十亿级数据资产,持续构建数据飞轮,通过统一化存储、多模态检索和规模化挖掘,将海量真实世界数据转化为可持续训练和优化模型的数据资产。当前,模型单次训练数据吞吐量已达到1亿clips,数据吞吐量比半年前增长10倍。通过对数据进行结构化、标签化和语义化处理,小鹏能够从海量真实世界数据中高效发现未知问题,并快速反哺模型训练和迭代,形成数据越多、问题发现越充分、模型进化越快、又产生更多数据的正向循环。
数据规模在增长,模型训练效率在提升,算力和工具链持续完善,越来越多真实世界的问题能够被纳入模型迭代,数据规模的护城河正逐步形成,小鹏集团在AI Infra领域的长期投入,已逐步沉淀为可持续复用的基础能力,呈现出明显的规模效应,开始产生复利。
The model will constantly iterate, but what can truly form long-term barriers is the ability to continuously iterate the model; Products will constantly change, but what truly generates long-term value is the infrastructure that supports their continuous evolution. The continuous upgrade of the second generation VLA is the result of the continuous operation of this system, which promotes the collaborative growth of models, data, computing power, and ontology, continuously transforming model capabilities into security, comfort, and efficiency improvements that can be perceived by real users.
As AI Infra continues to accumulate, the evolution of model capabilities will no longer be limited to a single product, but will cross different entities such as cars, robots, and flying cars, continuously expanding the boundaries of interaction between physical AI and the real world. From AI cars to general humanoid robots, and then to more intelligent entities, ultimately moving towards a broader physical world.
模型会不断迭代,但真正能够形成长期壁垒的,是持续迭代模型的能力;产品会不断变化,但真正能够产生长期价值的,是支撑产品持续进化的基础设施。第二代VLA的持续升级,正是这一体系不断运转的结果,通过推动模型、数据、算力和本体的协同增长,将模型能力不断转化为真实用户可感知的安全性、舒适性和效率提升。
随着AI Infra持续沉淀,模型能力的进化将不再局限于单一产品,而是跨越汽车、机器人、飞行汽车等不同本体,持续拓展物理AI与真实世界交互的边界。从AI汽车到通用人形机器人,再到更多的智能本体,最终走向更广阔的物理世界。
(Using AI translation)
(使用AI翻译)