![]() |
市場調查報告書
商品編碼
2129080
2026年汽車與機器人領域大型VLA模型應用調查報告Research Report on Application of VLA Large Model in Automobiles and Robots, 2026 |
||||||
汽車和機器人領域中 VLA 的研究:混合架構正在成為主流,VLA 正在與通用世界模型整合,強化學習正在作為核心引擎發揮作用。
VLA模型是一種整合了視覺、語言和動作三種模態的模型。它採用統一的多模態學習框架,整合感知、推理和控制,能夠直接從視覺輸入(影像和影片)和口頭指令產生可執行的物理世界動作(例如,機器人關節運動、車輛轉向、加速和煞車控制)。
VLA對應於VLM(理解與說明)+ E2E(端對端決策與控制)+ CoT(類人推理鏈)。從功能比較的角度來看,VLA同時實現了精確的3D感知、常識理解、邏輯思考和可解釋性。
自動駕駛中 VLA 的發展可以分為四個階段。
語言用作解釋器的階段(VLA 前):語言模型不參與控制,只產生場景說明。
模組化 VLA:語言作為決策的規劃組成部分,但多階段工作流程會引入延遲。
整合端對端 VLA:感測器輸入透過單次前向傳播直接對應到動作。
增強型推理型虛擬邏輯陣列:LLM 進入封閉回路型控制迴路,進而實現長期推理、記憶與互動功能。力汽車的 MindVLA 就是一個典型的例子。
實施時間表:
分段式端到端架構在 2024 年至 2025 年間實現量產。
2025 年至 2026 年間,採用單一模型的端對端和 VLA 解決方案得到了廣泛應用。
2026 年將是大規模自動駕駛模型的關鍵轉折點,其特點是「多方面競爭加劇和範式融合加速」。
技術狀況和關鍵指標。
模型參數範圍廣泛:NVIDIA Alpamayo 1.5 包含 0.5B/10B 參數;DeepRoute.ai 使用包含 40B 參數的基礎模型;Afari Technology 的 StepVL 基礎模型具有 32B 參數(精簡為 7B 和 3.6B);理髮變體,其中 4B 分配給車輛模型。
汽車即時性能:理想汽車利用稀疏注意力機制和MoE技術,在Orin X平台上實現了10Hz的響應頻率和100ms的延遲。小鵬汽車的第二代VLA實現了低於80ms的延遲。 DeepRoute.ai的40B參數模型利用KV快取、多令牌預測(MTP)、量化以及客製化引擎,將單步延遲控制在60-85ms,封閉回路型在10-15Hz。
運算能力適應性:DeepRoute.ai 可在 100 TOPS 平台上部署純駕駛 VA 模型,並在 500 TOPS 平台上部署支援推理的 VLA 模型。 Leapmotor D19(1280 TOPS)搭載兩顆高通驍龍 8797 晶片,可實現端對端 VLA 輔助駕駛。吉利 H9 則採用兩顆 NVIDIA Thor 晶片(2000 TOPS)。
開放回路性能:基於nuScenes資料集的評估表明,與世界模型相比,虛擬雷射輔助系統(VLA)的軌跡誤差顯著更小。例如,擁有70億個參數的AutoDrive-R2實現了0.19米的L2距離,SENNA實現了0.22米,而即使是性能最佳的世界模型Drive-OccWorld也僅達到了0.32米。這些結果表明,VLA能夠重現真實的人類駕駛軌跡。
主要挑戰
即時效能與運算能力瓶頸:傳統的自回歸 VLA 產生頻率僅達到 3-6 Hz,單次推理延遲通常超過 200 毫秒,消耗大量的車輛運算能力和儲存頻寬。
資料監管不足:雖然甚大陣列(VLA)接收高維度視覺輸入,但其監管卻依賴低維度、稀疏的動作數據,這限制了模型的潛力。為了實現高密度監管,需要一個能夠預測未來影像的世界模型,或需要進行影片預測的預訓練。
缺乏安全冗餘:採用單一端對端模型的垂直陣列(VLA)容錯能力低,需要傳統演算法作為備用方案。例如,廣泛應用的快速投擲雙系統、Horizon Robotics 的 Lite Safety Checker 以及 Bosch 的安全門控獎勵機制。
幻覺和長尾場景:在複雜或未知的場景中,成功率會降低。改進需要利用世界模型和強化學習進行探索,例如博世的 ExploreVLA、華為的 WEWA 和 Momenta 的 R7。
在眾多汽車製造商中,小鵬汽車的第二代 VLA 和理想汽車的 MindVLA 是代表性的例子,它們分別對應於兩種技術方法:「原生多模態現實世界平台模型」和「整合空間、語言和行為 + 隱性世界模型」。
小鵬汽車第二代超音速火箭
小鵬汽車認為自動駕駛本質上是一個實體人工智慧問題。其第二代視覺學習架構(VLA)建構為一個原生多模態物理世界基礎模型。原生多模態分詞器實現了高效的初始融合,避免了單模態偏差。視覺推理的「思考鏈(CoT)」將推理效率提升了32倍。在自適應巡航控制場景中,該模型能夠自動產生諸如變換車道和跟隨前車等建議動作,從而建立一個用於評估的抽象鳥瞰圖。原生駕駛座和駕駛整合使該模型不僅能夠產生動作,還能產生影片和音頻,這既是VLA的基礎,也是世界模型、模擬和強化學習的基礎框架。透過參考兩階段模型、重新訓練單階段基礎模型以及消除語言翻譯環節,延遲控制在80毫秒以下。
李汽車 MindVLA
該架構由三個組件構成。
V(空間智慧):基於BEV和OCC,它採用3D高斯分佈作為中間表示,利用LiDAR點雲作為3D幾何提示,並透過3D ViT編碼器和前饋3DGS進行3D場景重建。靜態環境和動態物體分別建模。
L(語言智慧):重新訓練基於 LLM 的模型(利用 MoE 和稀疏注意力機制)。對於快速思考,並行解碼直接輸出動作標記;而對於慢速思考,則同時輸出 CoT 和動作標記。
A(行動策略):我們採用融合了行動專家的VLA-MoE架構。它利用離散擴展和平行解碼迭代最佳化,輸出高精度的軌跡資料。
工程的四個階段:
基於 VL 的模型的預訓練(蒸餾後到 MoE 3.6B,以適應 32B、Orin-X/Thor-U 的雙重環境)
訓練後透過模仿學習(4B)
強化學習(RLHF + 純強化學習,建構封閉回路型世界模擬器)
驅動代理人機介面
隨後,MindVLA-U1 實作了語言和順序動作的整合式串流和協作建模,並引入了意圖控制語法 (Intent-CFG)。
其他原始設備製造商:
小米 XLA 認知大型模式(VLM + 邊緣雲整合,規劃部署 VLA);
Leapmotor LEAP 4.0 VLA(邊緣的全模態,「理解、計劃、預覽、判斷和糾正」的封閉回路型)
長城CP大師雙VLA(左腦:駕駛代理,右腦:駕駛座代理);
Cherry Falcon 900(VLA + 世界型號,L3 相容)。
供應商中最具代表性的解決方案是 NVIDIA 的 Alpamayo 和 DeepRoute.ai 的 40B VLA,它們分別體現了「基於推理的雙 LLM VLA + 世界模型」和「基於 40B 整合的三階段 VLA」解決方案。
NVIDIA Alpamayo
NVIDIA Alpamayo 是一款系統級 VLA 解決方案,它包含實體 AI 資料集、大規模 VLA 模型和 AlpaSim 模擬框架。它為 OEM 提供三代車輛部署軌跡產生選項:回歸預測(TensorRT,例如 SparseDrive)、擴散生成(TensorRT,例如 DiffusionDrive)和流匹配(TensorRT-Edge-LLM,例如基於 Qwen3 VL 的 Alpamayo)。
以 Alpamayo-R1-10B 為例。其輸入包括歷史影像、使用者指令、歷史軌跡以及帶有雜訊的動作。輸入資料透過 VLM 流匹配分詞器轉換為文字、圖像和軌跡標記。一個 80 億參數的 Qwen3 VL-LLM 執行場景理解和隱式 CoT 推理,輸出推理文本並產生 KV 快取。另一個 20 億參數的 Qwen3 VL-LLM 接收帶雜訊的軌跡和 KV 快取,利用串流匹配進行逐步去噪和校正,最終輸出未來軌跡。 Cloud Cosmos 世界模型產生長尾場景的訓練資料。對於車輛部署,傳統演算法作為安全替代方案,端到端的 VLA 是主要系統,支援 L2-L4 層。未來,潛在空間推理的處理速度可望提升 2-4 倍。
DeepRoute.ai 40B VLA
DeepRoute.ai 將自動駕駛決策過程分解為三個階段。
觀察階段 - 多攝影機影片被編碼成大約 1,000 個視覺標記。
推理階段-模型對場景進行詳細的語意分析,並產生關鍵事件和決策邏輯的說明。推理詞元的數量嚴格控制在10到50之間。
執行階段 - 輸出操作控制指令只需要大約 10 個令牌。
該系統透過統一的40B參數基礎模型,整合了「駕駛員」(根據感測器輸入做出反應)、「分析員」(分析因果關係)和「評論員」(判斷並做出決策)三個功能。為了實現“先思考後駕駛”,系統在V+A、V+A→L和V→L+A三個任務類別中進行協同學習。預訓練從基於軌蹟的監督學習切換到影片預測。透過利用海量影片在像素層級學習物理定律,資料利用率從0.001%提升至100%。部署過程中,鍵值快取、多目標處理(MTP)、量化以及客製化的推理引擎將每步延遲降低至60-85毫秒,實現了10-15赫茲的即時封閉回路型車輛運行。模型蒸餾依運算能力進行,純駕駛VA模型運行在100TOPS平台上,完整的VLA模型運行在500TOPS平台上。
其他供應商:
Afari Technology採用「VLA+E2E」協同封閉回路型。 VLA的低速系統輸出CoT文本,而E2E的高速系統輸出目標偵測/車道偵測結果,兩者融合後用於規劃和控制。 StepVL 4.0已從32B預訓練模型精簡至7B。
QCraft 將於 2026 年升級為「VLA + 世界模型 + 強化學習」的整合架構,從而在單一 Journey 6M 晶片上實現都市區的 NOA。
卓宇發布了 VLA 世界模型(原生多模態基礎模型、世界鏈、具有分離結構和運動的潛在運動表示)。
趨勢一:混合架構正在成為主流。
擴散、 變壓器和 LLM/VLM 的深度融合將成為主流。變壓器擅長生成高品質、連續的運動和軌跡。 Transformer 擅長建模長序列。 LLM/VLM 擅長語意和多模態理解。代表性例子包括力汽車的 MindVLA(3D 高斯分佈 + MoE LLM + 擴散動作專家)、NVIDIA 的 GR00T-N1(快慢雙系統:快速 200Hz 擴散動作,慢速 10Hz VLM)以及 HybridVLA(自回歸 + 協同擴散)。業界已湧現以下三種整合模型:1. 單模型端對端 (E2E) + 世界模型 + 強化學習 (RL)(Momenta、Horizon Robotics);2. VLA + 世界模型(小鵬等); 3. E2E + VLM/VLA 模型(Afari Technology 的 VLA 慢速系統)。
趨勢 2:透過 VLA 和通用世界模型的融合,「世界 VLA/VLA 世界模型」的誕生。
VLA負責認知和行為,而世界模型負責未來預測。透過整合二者,汽車從單純的交通工具轉變為移動機器人,從規則驅動系統轉向認知驅動系統,並獲得自主感知、推理和決策能力以及精準執行能力。卓宇的VLA世界模型已發展成為第三代原生多模態基礎模型。透過“世界鏈”,它能夠在潛在空間中進行多階段世界狀態預測,實現“先思考後行動”。結構與運動的分離以及潛在運動表示降低了重建成本。吉利G-ASD整合了VLA和世界模型,使車輛能夠自動執行任務。 WorldVLA能夠同時理解運動和圖像並進行生成,世界模型和運動模型相互補充。
趨勢 3:將 VLA + 世界模型 + 強化學習 (RL) 整合為三位一體,其中 RL 作為核心引擎。
業界已建立起「預訓練→模擬→強化學習」的三層架構。此架構利用世界模型產生長尾場景,利用虛擬邏輯陣列(VLA)進行循環推理,並透過強化學習在推理空間中迭代推導出最佳策略。代表性案例包括華為的WEWA 2.0(多智慧體博弈+雲端在線強化學習,學習能力提升十倍)、Momenta的R7(三階段流程:預訓練→仿真→強化學習,將AI從“模仿者”轉變為“決策者”)以及Pony.ai的PonyWorld 2.0(自診斷+定向進化+精準飛輪)。
同時,世界模型正從像素級預測演進到潛在空間和因果推理。 NVIDIA 的 Alpamayo 透過潛在空間中的隱式推理實現了 2-4 倍的加速,並透過因果鏈 (CoC) 產生完整的推理鏈。理想汽車將預測性隱式世界模型整合到其 VLA 中。小鵬汽車已將其架構從 VLA 改進為 V/LA,以消除語言轉換環節並減少資訊損失。華為的 DriveVLA-W0 表明,世界模型的整合放大了這些優勢,並強化了數據擴展規律,因為隨著數據量從 70 萬幀擴展到 7000 萬幀,衝突率持續下降。
趨勢四:加速技術實施與安全保障
2025年至2026年間,單一模型端對端和超大規模自動化(VLA)解決方案將廣泛部署。隨著L3/L4級別的演進,安全冗餘至關重要(採用傳統演算法作為備用方案,同時啟用端對端主系統;例如NVIDIA的快慢雙系統、Horizon Robotics的Lite Safety Checker、Bosch的Safety Gating PDMS Reward)。運算能力的分層分配是實現大規模生產的關鍵:靈活部署100至500 TOPS的超大規模自動化(VA)/超大規模自動化(VLA)(例如DeepRoute.ai),以及支援L3等級的高效能運算平台,例如雙Thor/雙8797處理器。
車輛學習架構(VLA)將成為自動駕駛從「端到端感知與控制」向「理解、推理與控制」演進的核心路徑。 2026年,在整車廠商(如小鵬汽車、理想汽車等)和供應商(如英偉達、DeepRoute.ai、Afari Technology、QCraft等)的共同推動下,VLA將與世界模型和強化學習深度融合,形成混合架構,並最終實現「世界級VLA」原型。隨著延遲、運算能力、資料監管和安全冗餘等挑戰的逐步解決,VLA將支援L3級及以上自動駕駛的大規模部署,使車輛能夠演化為現實世界中的通用智慧體。
定義
Research on Automotive and Robot VLA: Hybrid Architectures Become Mainstream, VLA Integrates with General World Models, and Reinforcement Learning Serves as Core Engine
Vision-Language-Action (VLA) model is a model integrating vision, language and action modalities. Adopting a unified multimodal learning framework, it integrates perception, reasoning and control, and generates executable physical world actions (e.g., robot joint motion, and vehicle steering/acceleration/braking control) directly from visual inputs (images/videos) and language instructions.
VLA equals VLM (for understanding and description) plus E2E (end-to-end decision and control) plus CoT (chain-of-thought human-like reasoning). In terms of capability comparison, VLA delivers precise 3D perception, commonsense understanding, logical thinking and interpretability simultaneously.
The evolution of VLA in autonomous driving falls into four stages.
Language as interpreter (Pre-VLA): Language models only generate scene descriptions without participating in control.
Modular VLA: Language acts as a planning component for decision, yet multi-stage workflows incur latency.
Unified end-to-end VLA: Sensor inputs are directly mapped to actions via a single forward propagation.
Reasoning-enhanced VLA: LLMs enter the control closed loop to enable long-term reasoning, memory and interaction capabilities, exemplified by Li Auto MindVLA.
Implementation timeline:
Segmented end-to-end came into mass production from 2024 to 2025.
One-model end-to-end and VLA were largely rolled out between 2025 and 2026.
The year 2026 marks a critical window period of "intensified multi-route competition and accelerated paradigm integration" for intelligent driving large models.
Technical status and core indicators.
Wide-ranging model parameters: NVIDIA Alpamayo 1.5 includes 0.5B/10B parameters; DeepRoute.ai uses a 40B-parameter foundation model; StepVL foundation model from Afari Technology features 32B parameters (distilled to 7B and 3.6B); Li Auto MindVLA 32B-parameter foundation model is distilled into a 3.6B-parameter MoE variant, 4B for vehicle model.
In-vehicle real-time performance: Li Auto leverages sparse attention + MoE to realize 10Hz and 100ms latency on Orin X; Xpeng's second-generation VLA enables <80ms latency; DeepRoute.ai's 40B-parameter model uses KV Cache, Multi-Token Prediction (MTP), quantization and customized engine to achieve single-step latency of 60-85ms and 10-15Hz closed loop.
Computing power adaptation: DeepRoute.ai can deploy pure driving VA models on 100TOPS platforms and reasoning-capable VLA models on 500TOPS platforms; Leapmotor D19 equipped with dual Qualcomm 8797 chips (1280TOPS) realizes end-side VLA-assisted driving; Geely H9 adopts dual NVIDIA Thor chips (2000TOPS).
Open-loop performance: Based on the nuScenes dataset, VLA exhibits notably lower trajectory errors than world models. For instance, AutoDrive-R2 with 7B parameters delivers an L2 distance of 0.19m and SENNA achieves 0.22m; the best-performing world model Drive-OccWorld reaches 0.32m. Such results demonstrate VLA's ability to reproduce human real-world driving trajectories.
Major challenges
Real-time performance and computing power bottlenecks: Traditional auto-regressive VLA generation only reaches 3-6Hz, and single-reasoning latency commonly exceeds 200ms, consuming a lot of vehicle computing power and storage bandwidth.
Data supervision deficit: VLA receives high-dimensional visual inputs yet is supervised by low-dimensional sparse actions, limiting model potential. World models are required to predict future images for dense supervision, or video prediction pre-training shall be adopted.
Lack of safety redundancy: End-to-end single-model VLA has low fault tolerance and requires fallback from traditional algorithms, e.g., the widely adopted fast-slow dual-system, Horizon Robotics Lite Safety Checker and Bosch safety gating reward mechanism.
Hallucinations and long-tail scenarios: Success rates drop in complex or unseen scenarios. World model + reinforcement learning exploration is required to improve, e.g., Bosch ExploreVLA, Huawei WEWA and Momenta R7.
Among OEMs, Xpeng's second-generation VLA and Li Auto MindVLA stand as representative cases, corresponding respectively to two technical routes, namely "native multimodal physical world foundation model" and "space-language-action unification + implicit world model".
Xpeng's Second-Generation VLA
Xpeng holds to the view that intelligent driving is essentially a physical AI problem. Its second-generation VLA is built as a native multimodal physical world foundation model. A native multimodal tokenizer enables highly efficient early-stage fusion to avoid single-modality bias. Visual reasoning chain-of-thought (CoT) boosts reasoning efficiency by 32 times. In car following scenarios, the model automatically generates maneuver proposals such as lane change or car following, and produces abstract bird-eye-view diagrams for scoring. Native cockpit-driving linkage allows the model to generate not only actions but also videos and sounds, serving as the foundation for VLA and the foundation framework for world models, simulation, and reinforcement learning. Referring to two-stage models, it retrains one-stage foundation models and eliminates language translation links to bring latency below 80ms.
Li Auto MindVLA
Its architecture consists of three components.
V (spatial intelligence): Based on BEV and OCC, it adopts 3D Gaussian as intermediate representation, leverages LiDAR point clouds as 3D geometric prompts, and performs 3D scene reconstruction via 3D ViT encoder and feed forward 3DGS. Static environments and dynamic objects are modeled separately.
L (language intelligence): Retrains LLM foundation model (leveraging MoE and sparse attention mechanism). Fast-thinking parallel decoding directly outputs Action Tokens, while slow thinking outputs CoT and Action Tokens simultaneously.
A (action strategy): Adopts VLA-MoE architecture embedded with Action Expert. Discrete diffusion and parallel decoding iterative optimization are used to output high-precision driving trajectories.
Four-phase engineering:
VL foundation model pre-training (32B, post-distilled to 3.6B MoE to adapt to dual Orin-X/Thor-U)
Imitation learning post-training (4B)
Reinforcement training (RLHF + pure RL, building a closed-loop world simulator)
Driver agent HMI
Subsequent MindVLA-U1 enables unified streaming and joint modeling of language and continuous actions, and introduces Intent-CFG.
Other OEMs:
Xiaomi XLA Cognitive Large Model (VLM + edge-cloud integration, VLA planned);
Leapmotor LEAP 4.0 VLA (edge-side full-modality, "understanding-planning-preview-judgment-correction" closed loop);
Great Wall CP Master Dual VLA (left brain driving agent, right brain cockpit agent);
Chery Falcon 900 (VLA + world model, supporting L3).
The most typical solutions of suppliers are NVIDIA Alpamayo and DeepRoute.ai 40B VLA, representing "reasoning-based dual-LLM VLA + world model" and "40B unified base three-stage VLA" solution respectively.
NVIDIA Alpamayo
NVIDIA Alpamayo is a system-level VLA solution encompassing physical AI dataset, VLA large model and AlpaSim simulation framework. It provides OEMs with three generations of trajectory generation options for vehicle deployment: regression prediction (TensorRT, e.g., SparseDrive), diffusion generation (TensorRT, e.g., DiffusionDrive), and flow matching (TensorRT-Edge-LLM, e.g., Alpamayo based on Qwen3 VL).
Take Alpamayo-R1-10B as an example. Inputs include historical images, user instructions, historical trajectories and noisy actions. Inputs are converted into Text, Image and Trajectory Tokens via VLM Flow Matching Tokenizer. The 8B-parameter Qwen3 VL-LLM handles scene understanding and implicit CoT reasoning, outputs reasoning texts and generates KV Cache. The 2B-parameter Qwen3 VL-LLM receives noisy trajectories and KV Cache, conducts progressive denoising and correction via Flow Matching, and finally outputs future trajectories. Cloud Cosmos world model generates training data for long-tail scenarios. Vehicle deployment adopts traditional algorithm as safety fallback + end-to-end VLA as primary system, supporting L2-L4. Future latent space reasoning is expected to accelerate by 2-4 times.
DeepRoute.ai 40B VLA
DeepRoute.ai breaks down autonomous driving decision into three phases.
Observation phase - Multi-camera videos are encoded into approximately 1,000 visual tokens.
Reasoning phase - The model conducts in-depth semantic analysis of scenarios and generates descriptions of key events and decision logics, with the number of reasoning tokens strictly controlled within 10-50.
Execution phase - Outputting driving control commands requires only about 10 tokens.
It integrates three capabilities of "driver (acting based on sensor inputs), analyst (analyzing causality), and commentator (judging and making decisions)" through a unified 40B-parameter foundation model. Joint training is implemented across three task categories, namely, V+A, V+A->L and V->L+A, to realize "thinking before driving". Pre-training switches from trajectory supervision to video prediction. Massive videos are leveraged to learn physical laws at per-pixel level, lifting data utilization rate from 0.001% to 100%. During deployment, KV Cache, MTP, quantization and customized reasoning engines reduce single-step latency to 60-85ms to achieve a 10-15Hz vehicle real-time closed loop. Model distillation is performed according to computing power: pure driving VA models run on the 100TOPS platform, and complete VLA models operate on the 500TOPS platform.
Other suppliers:
Afari Technology adopts the ""VLA+E2E" collaborative closed loop. The VLA slow system outputs CoT texts while the E2E fast system outputs target detection/lane detection results for fusion into planning and control. StepVL 4.0 is distilled from 32B pre-trained to 7B.
QCraft upgrades to the "VLA + world model + reinforcement learning" unified architecture in 2026, and realizes urban NOA on single Journey 6M chip.
Zhuoyu launches VLA World Model (native multimodal foundation model, Chain of World, and structure-motion decoupled latent motion representation).
Trend 1: Hybrid Architectures Become Mainstream
Deep integration of Diffusion, Transformer and LLM/VLM becomes mainstream. Diffusion excels at generating high-quality continuous actions and trajectories. Transformer is good at long sequence modeling. LLM/VLM is skilled in semantic and multimodal understanding. Representative examples include Li Auto MindVLA (3D Gaussian + MoE LLM + Diffusion Action Expert), NVIDIA GR00T-N1 (Fast-Slow Dual System: Fast 200Hz Diffusion Action, Slow 10Hz VLM), HybridVLA (Autoregression + Collaborative Diffusion). The industry has formed three integration models: 1. One-model end-to-end + world model + RL (Momenta, Horizon Robotics); 2. VLA + world model (XPeng, etc.); 3. E2E + VLM/VLA foundation model (Afari Technology VLA slow system + E2E fast system).
Trend 2: VLA and General World Model Integrate into World VLA / VLA World Model
VLA undertakes cognition and action while the world model takes on future prediction. Their unification transforms automobiles from transportation means into mobile robots, shifting from rule-driven to cognition-driven, with capabilities of autonomous perception, reasoning & decision and precise execution. Zhuoyu's VLA World Model has evolved into its third-generation native multimodal foundation model. With Chain of World, it performs multistep world state prediction in latent space, achieving "thinking before acting." Structure-motion decoupling and latent motion representation reduce reconstruction costs. Geely G-ASD integrates VLA and world model, enabling vehicles to automatically perform tasks. WorldVLA jointly understands actions and images for generation, with the world model and action model mutually reinforcing each other.
Trend 3: VLA + World Model + Reinforcement Learning (RL) Trinity Integration, with RL as the Core Engine
The industry forms a "pre-training -> simulation -> reinforcement learning" three-layer architecture. The world model generates long-tail scenarios, VLA conducts in-loop reasoning, and reinforcement learning iterates optimal strategies in the inference space. Representative examples include: Huawei WEWA 2.0 (multi-agent gaming + cloud online RL, training intensity increased by 10 times); Momenta R7 (three-stage process: pre-training -> simulation -> RL, turning AI from "imitator" to "decision-maker"); Pony.ai's PonyWorld 2.0 (self-diagnosis + targeted evolution + precision flywheel).
Meanwhile, world models evolve from pixel-level prediction toward latent space and causal reasoning. NVIDIA Alpamayo achieves 2-4-fold acceleration via implicit reasoning in the latent space, and generates a complete reasoning chain through Chain of Causality (CoC). Li Auto embeds predictive implicit world models into VLA. Xpeng eliminates language translation links and revises architecture from V-L-A to V/L-A to mitigate information loss. Huawei DriveVLA-W0 verifies that with world model integration, collision rates keep decreasing as data volume expands from 0.7 million to 70 million frames and such advantages are amplified, strengthening the data scaling law.
Trend 4: Engineering Implementation and Safety Assurance Accelerate
One-model end-to-end and VLA solutions are largely implemented from 2025 to 2026. The evolution of L3/L4 has driven safety redundancy to become a necessity (traditional algorithm fallback + end-to-end main system, e.g., NVIDIA's fast-slow dual systems, Horizon Robotics' Lite Safety Checker, and Bosch's safety gating PDMS reward). Hierarchical distillation of computing power has become key to mass production: flexible deployment of VA/VLA (DeepRoute.ai) at 100-500 TOPS, and high-performance computing platforms such as dual Thor/dual 8797 supporting L3.
VLA serves as the core route for intelligent driving to evolve from "end-to-end perception-control" toward "understanding-reasoning-control". In 2026, driven by both OEMs (XPeng, Li Auto, etc.) and suppliers (NVIDIA, DeepRoute.ai, Afari Technology, QCraft, etc.), VLA is deeply integrated with world model and reinforcement learning, forming a hybrid architecture, and the prototype of World VLA takes shape. As latency, computing power, data supervision, and security redundancy issues are gradually resolved, VLA will support the large-scale deployment of L3 and above autonomous driving and enable vehicles to evolve into general agents in the physical world.
Definitions