← Back to Index
Daily Autonomous Driving Research Digest

AutoDrive Papers

2026-07-23
9
Papers
4
Topics
9
Translated

感知

4
感知 / 1 / 2607.19528

D3VL: Understanding Driving Scenes from 3D Time Series Data and Video with Language Models

D3VL:通过语言模型理解来自3D时间序列数据和视频的驾驶场景
Heesang Han, A. Lynn Abbott, Abhijit Sarkar
cs.CV · cs.AI
Abstract
Recent advances in Multimodal Large Language Models (MLLMs) have triggered the development of end-to-end MLLMs for autonomous driving. However, the main emphasis to date has been for MLLMs using 2D images and videos. In contrast, this paper considers MLLM effectiveness using 3D sensors, particularly LiDAR and stereo cameras. LiDAR presents unique challenges to integration within an MLLM, largely because of data sparsity and lack of a grid structure for the data. For similar reasons, fusion of camera and LiDAR data within an MLLM pipeline is also uncommon. However, most autonomous systems rely on LiDAR-based sensing, and incorporating 3D data has been proven to improve performance in traditional 3D scene perception tasks. This paper presents D3VL, a novel MLLM framework that integrates 2D and 3D time-series data in a single but simple architecture. The model aims to answer questions involving traffic scene understanding and safety. D3VL shows an 11% improvement in the KITTI Question-Answering (QA) dataset compared to baseline methods in processing 2D and 3D time-series data. This paper further introduces the Waymo QA dataset extension, which assesses models' capabilities in processing 3D and time-series data under diverse driving conditions. D3VL implementation code and WaymoQA extension can be found on our supplemental website: https://automotivesafety-lvlm.github.io
Chinese Translation
近年来,多模态大型语言模型(MLLMs)的进展推动了端到端MLLMs在自动驾驶领域的发展。然而,到目前为止,主要的关注点仍然是使用2D图像和视频的MLLMs。相较之下,本文考虑了使用3D传感器(特别是激光雷达(LiDAR)和立体摄像头)时MLLM的有效性。激光雷达由于数据稀疏性和缺乏网格结构而在MLLM中的集成面临独特挑战。出于类似原因,在MLLM管道中融合摄像头和激光雷达数据也不常见。然而,大多数自动驾驶系统依赖于基于激光雷达的感知,且已证明融合3D数据能够提高传统3D场景感知任务的性能。本文提出了D3VL,一个新颖的MLLM框架,能够在单一且简单的架构中集成2D和3D时间序列数据。该模型旨在回答与交通场景理解和安全相关的问题。与基线方法相比,D3VL在处理2D和3D时间序列数据时,在KITTI问答(QA)数据集上显示出11%的性能提升。本文进一步介绍了Waymo QA数据集扩展,评估模型在多样化驾驶条件下处理3D和时间序列数据的能力。D3VL的实现代码和WaymoQA扩展可以在我们的补充网站上找到:https://automotivesafety-lvlm.github.io
感知 / 2 / 2607.19617

EGRNet: A Lightweight Semantic Segmentation Network with Edge-Gated Refinement and Adversarial Sensing

EGRNet:一种具有边缘门控细化和对抗感知的轻量级语义分割网络
Bareera Qaseem, Mohsin Kamal, Muhammad Naveed Aman
cs.CV
Abstract
As autonomous systems and smart cities continue to evolve, the demand for efficient and robust scene understanding becomes increasingly critical. Semantic segmentation plays a key role in enabling autonomous vehicles to comprehend complex urban environments. However, achieving high accuracy with minimal computational cost remains a significant challenge. In this paper, we present Edge-Gated Refinement Network (EGRNet), a lightweight and efficient deep learning model designed for real-time semantic segmentation in urban scenarios. The model incorporates depthwise separable convolutions to reduce computational complexity and dilated residual blocks for capturing rich multi-scale contextual information. Additionally, we introduce a novel Edge-Gated Refinement (EGR) module, which adaptively fuses original and refined features through a learnable gating mechanism, enhancing boundary preservation and edge-sensitive regions. To further improve feature representation, Squeeze-and-Excitation (SE) attention is applied across the network. With only 0.46M parameters, EGRNet achieves state-of-the-art performance while maintaining low computational overhead. When evaluated on the Cityscapes dataset, the model attains a mean Intersection over Union (mIoU) of 65.28%, demonstrating strong accuracy with minimal resource consumption. Moreover, we introduce a lightweight adversarial attack detection strategy, ensuring robustness against adversarial inputs without compromising real-time performance. By combining efficiency, accuracy, and resilience, EGRNet is well-suited for deployment on edge devices in safety-critical real-time applications.
Chinese Translation
随着自主系统和智慧城市的不断发展,对高效且稳健的场景理解的需求变得愈加重要。语义分割在使自主车辆理解复杂城市环境中发挥着关键作用。然而,以最小的计算成本实现高准确率仍然是一个重大挑战。本文提出了边缘门控细化网络(Edge-Gated Refinement Network,EGRNet),这是一种轻量级且高效的深度学习模型,旨在实现城市场景中的实时语义分割。该模型采用深度可分离卷积以降低计算复杂度,并使用膨胀残差块来捕捉丰富的多尺度上下文信息。此外,我们引入了一种新颖的边缘门控细化(Edge-Gated Refinement,EGR)模块,通过可学习的门控机制自适应融合原始特征和细化特征,增强边界保持和边缘敏感区域。为了进一步改善特征表示,网络中应用了挤压与激励(Squeeze-and-Excitation,SE)注意力机制。EGRNet仅需0.46M参数,即可在保持低计算开销的同时实现最先进的性能。在Cityscapes数据集上的评估中,该模型达到了65.28%的平均交并比(mean Intersection over Union,mIoU),展现出强大的准确性和最低的资源消耗。此外,我们还引入了一种轻量级的对抗攻击检测策略,确保在不影响实时性能的情况下对抗输入的鲁棒性。通过结合效率、准确性和韧性,EGRNet非常适合在安全关键的实时应用中部署于边缘设备。
感知 / 3 / 2607.19781

WASABI: Whole-graph Assignment-based Stabilizer for lAne topology By Inter-frame tracking

WASABI:基于全图分配的稳定器,通过帧间跟踪实现车道拓扑稳定
Tetsuhiro Uchida, Myu Sasaki, Kensho Nakajima, Yasuhiro Shimada, Toru Saito
cs.CV
Abstract
Autonomous driving requires understanding the road as a graph of drivable lanes and their connectivity, beyond the ego lane alone, to follow routes through intersections and reason about cross- and merging-traffic. Recent perception models infer such lane topology, i.e., lane segments together with their inter-lane connectivity (LCLC), from onboard sensors over a 360-degree BEV view. Due to neural perception's imperfections, their outputs retain structural instabilities such as missed detections, lost or incorrect LCLC, over-detection, and label flicker. This paper presents WASABI, a real-time post-processing pipeline that stabilizes lane topology outputs both within and across frames by treating lane segments and their LCLC connectivity as joint tracking targets, under onboard real-time constraints (10 Hz / 20 ms / up to 200 input lanes). The pipeline integrates segment tracking with connectivity, noise-robust topology-aware refinement, and a resource-constrained real-time design. On internal validation data (16 sequences), WASABI improves LCLC detection F1 from 0.834 to 0.948 (+0.114, +13.6%) and reduces centerline lateral error from 2.50 m to 0.95 m, while reducing detection false-positives by 24.6%. Temporal-stability metrics on the same data show LCLC toggle rate reduced by 63.3% and boundary-label flicker rate by 30.2%, confirming across-frame stabilization beyond per-frame accuracy.
Chinese Translation
自主驾驶需要将道路理解为可行驶车道及其连通性的图,不仅仅是自我车道,以便在交叉路口跟随路线并推理交叉和合并交通。最近的感知模型通过车载传感器在360度鸟瞰视图下推断出这种车道拓扑,即车道段及其车道间连通性(LCLC)。由于神经感知的不完美,其输出保留了结构不稳定性,例如漏检、丢失或错误的LCLC、过度检测和标签闪烁。本文提出了WASABI,一个实时后处理管道,通过将车道段及其LCLC连通性视为联合跟踪目标,在车载实时约束下(10 Hz / 20 ms / 最多200个输入车道)稳定车道拓扑输出。该管道将段跟踪与连通性、抗噪声的拓扑感知细化以及资源受限的实时设计相结合。在内部验证数据(16个序列)上,WASABI将LCLC检测F1从0.834提高到0.948(+0.114,+13.6%),并将中心线横向误差从2.50米减少到0.95米,同时将检测假阳性减少了24.6%。在相同数据上的时间稳定性指标显示,LCLC切换率减少了63.3%,边界标签闪烁率减少了30.2%,确认了跨帧稳定性超越逐帧准确性。
感知 / 4 / 2607.20071

GaussianSeed: Hierarchical Gaussian Seeding for High-Resolution 3D Occupancy Prediction

GaussianSeed:用于高分辨率3D占用预测的层次高斯种子
Xinzhuo Li, Xianghui Pan, Jiayuan Du, Wei Wei, Liuyi Wang, Chengju Liu, Qijun Chen
cs.CV
Abstract
Vision-centric 3D occupancy prediction provides dense scene representations essential for autonomous driving and robotic navigation, yet existing methods struggle to scale to high voxel resolutions due to prohibitive computational costs. To address this, we introduce GaussianSeed, a progressive multi-scale Gaussian occupancy prediction framework that organizes primitives into a coarse-to-fine hierarchy. Benefiting from this hierarchical design, GaussianSeed effectively circumvents the memory bottlenecks inherent in dense representations, successfully scaling to a $0.1\text{m}$ spatial resolution while maintaining real-time inference capabilities. To comprehensively evaluate high-resolution geometric perception, we further construct TJScenes, a panoramic six-camera occupancy dataset with highly detailed $0.1\text{m}$ annotations. Extensive experiments on Occ3D-nuScenes and TJScenes demonstrate that GaussianSeed delivers the lowest latency among all evaluated methods while maintaining highly competitive accuracy, advancing the efficiency-quality frontier of high-resolution 3D occupancy prediction.
Chinese Translation
以视觉为中心的3D占用预测提供了对自主驾驶和机器人导航至关重要的密集场景表示,但现有方法由于计算成本过高,难以扩展到高体素分辨率。为了解决这个问题,我们提出了GaussianSeed,一种渐进式多尺度高斯占用预测框架,将原始数据组织成粗到细的层次结构。得益于这种层次设计,GaussianSeed有效地规避了密集表示中固有的内存瓶颈,成功扩展到$0.1 ext{m}$的空间分辨率,同时保持实时推理能力。为了全面评估高分辨率几何感知,我们进一步构建了TJScenes,一个具有高度详细$0.1 ext{m}$标注的全景六摄像头占用数据集。在Occ3D-nuScenes和TJScenes上的大量实验表明,GaussianSeed在所有评估方法中提供了最低的延迟,同时保持了高度竞争的准确性,推动了高分辨率3D占用预测的效率与质量的前沿。

深度/几何

3
深度/几何 / 1 / 2607.19701

SafeGen: Goal-Conditioned Video Diffusion of Safety-Critical Scenarios for VLM-Based Autonomous Driving

SafeGen:基于目标条件的视频扩散框架用于安全关键场景的VLM自主驾驶
Jiangfan Liu, Zexuan Cui, Tianyuan Zhang, Zonglei Jing, Zonghao Ying, Yaoyuan Zhang, Jiakai Wang, Xiaoqi Jiang, Aishan Liu, Xianglong Liu
cs.CV
Abstract
VLMs are increasingly deployed in AD systems, creating an urgent need for rigorous safety evaluation under rare yet safety-critical scenarios. Among these, interactions with vulnerable road users represent a major source of real-world failures. However, existing safety-critical scenario generation methods predominantly rely on simulator-based pipelines, which suffer from a substantial sim-to-real gap and often fail to capture realistic, diverse, and unforeseen human-vehicle interaction dynamics. We present SafeGen, a goal-conditioned diffusion framework for safety-critical scenario generation in VLMADs. Our key insight is to formulate scenario generation as a goal-conditioned diffusion process, where a predefined catastrophic end-state serves as a strong supervisory signal, guiding the generation of temporally coherent video trajectories that naturally evolve toward safety-critical outcomes. Building on this formulation, we introduce Context Grounded End State Reasoning, which leverages VLMs to analyze benign driving contexts and infer latent vulnerabilities in human-vehicle interactions, producing structured end-state specifications that induce high-risk scenarios. Conditioned on these targets, we further propose End State Conditioned Video Evolution, which grounds semantic threats into physically plausible visual dynamics. Specifically, we instantiate high-risk agents within the scene via depth-aware geometric projection, followed by boundary-conditioned diffusion to generate intermediate frames with consistent motion patterns and temporal coherence. Extensive experiments across 3 VLMADs demonstrate that SafeGen increases the Judge Overall Score, a metric using a VLM judge to evaluate VLMADs' understanding and decision-making, by 24.25% on average compared to SoTA baselines. Furthermore, fine-tuning a VLMAD improves performance in real-world driving scenes by an average of 15.9%.
Chinese Translation
随着视觉语言模型(VLM)在自动驾驶(AD)系统中的广泛应用,迫切需要在稀有但安全关键的场景下进行严格的安全评估。在这些场景中,与易受伤害的道路使用者的互动是现实世界失败的主要来源。然而,现有的安全关键场景生成方法主要依赖于基于模拟器的流程,这些流程存在显著的模拟与现实之间的差距,往往无法捕捉到现实、丰富且不可预见的人车互动动态。我们提出了SafeGen,一个用于VLMAD安全关键场景生成的目标条件扩散框架。我们的关键见解是将场景生成形式化为一个目标条件扩散过程,其中预定义的灾难性终态作为强有力的监督信号,指导生成时间上连贯的视频轨迹,自然演变为安全关键的结果。在此基础上,我们引入了上下文基础的终态推理,利用VLM分析良性驾驶上下文并推断人车互动中的潜在脆弱性,生成结构化的终态规范,从而诱导高风险场景。在这些目标的条件下,我们进一步提出了终态条件视频演变,将语义威胁与物理上合理的视觉动态相结合。具体而言,我们通过深度感知几何投影在场景中实例化高风险代理,随后进行边界条件扩散,以生成具有一致运动模式和时间连贯性的中间帧。在3个VLMAD上的广泛实验表明,与最先进的基线相比,SafeGen平均提高了Judge Overall Score(一个使用VLM评估VLMAD理解和决策能力的指标)24.25%。此外,对VLMAD进行微调使其在真实驾驶场景中的表现平均提高了15.9%。
深度/几何 / 2 / 2607.19911

LoRFT: Benchmarking Long-Range Vehicle Trajectory Reconstruction from Fixed Highway Cameras

LoRFT:基于固定高速公路摄像头的长距离车辆轨迹重建基准测试
Yufan Zhu, Kefu Yi, Xueju Zhang, Yunyang Tian, Long Chen, Zixuan Xiao
cs.CV
Abstract
Long-range vehicle trajectories provide important spatio-temporal evidence for traffic safety analysis, autonomous driving evaluation, and data-driven traffic management, yet continuously recovering them from fixed highway cameras remains difficult. As vehicles recede into distant road regions, perspective compression and scale decay often fragment or prematurely terminate automatic tracklets, even when their continuation remains identifiable from motion consistency across neighboring frames. We formulate this problem as recovering the far-range continuation of a vehicle trajectory from a reliable near-field tracklet. We introduce LoRFT, to our knowledge the first open benchmark dedicated to long-range vehicle trajectory reconstruction from fixed highway cameras. LoRFT comprises 22 expressway surveillance scenes, 366,109 video frames, 6,601 manually verified trajectories, 2,694,889 bounding boxes, road-geometry annotations, scene-level splits, and evaluation scripts. We further propose Map-RSTNet, a map-aware residual sequence-to-sequence model that reconstructs distant trajectories in a road-geometry-aligned state space and dynamically refreshes local road geometry during decoding. On LoRFT, Map-RSTNet reduces ADE, FDE, and 5-second RMSE by 11.0%, 15.4%, and 10.5%, respectively, relative to the strongest baseline. These results demonstrate that road-geometry-aware reconstruction can extend usable trajectory records from existing fixed-camera infrastructure. LoRFT provides a reproducible testbed for long-range vehicle trajectory reconstruction.
Chinese Translation
长距离车辆轨迹为交通安全分析、自动驾驶评估和数据驱动的交通管理提供了重要的时空证据,但从固定高速公路摄像头持续恢复这些轨迹仍然困难。当车辆远离时,透视压缩和尺度衰减常常导致自动轨迹片段的碎片化或过早终止,即使它们的继续在相邻帧的运动一致性中仍然可识别。我们将此问题表述为从可靠的近场轨迹片段恢复车辆轨迹的远程延续。我们介绍了LoRFT,这是我们所知的第一个专门用于从固定高速公路摄像头进行长距离车辆轨迹重建的开放基准。LoRFT包含22个高速公路监控场景、366,109个视频帧、6,601个手动验证的轨迹、2,694,889个边界框、道路几何注释、场景级划分和评估脚本。我们进一步提出了Map-RSTNet,这是一种地图感知的残差序列到序列模型,它在与道路几何对齐的状态空间中重建远程轨迹,并在解码过程中动态刷新局部道路几何。在LoRFT上,Map-RSTNet相对于最强基线分别减少了ADE、FDE和5秒RMSE的11.0%、15.4%和10.5%。这些结果表明,基于道路几何的重建可以扩展现有固定摄像头基础设施的可用轨迹记录。LoRFT为长距离车辆轨迹重建提供了一个可重复的测试平台。
深度/几何 / 3 / 2607.19986

STEREOFLOW: Progressive Stereo Matching with StereoDiT and Transition Flow Matching

立体流:基于StereoDiT和过渡流匹配的渐进立体匹配
Hao Wang, Haoran Geng, Xiaotong Yang, Jing Tang, Songlin Wei, Linlong Lang, Yeying Jin, Zheng Zhu, Zhaoxin Fan, Biao Leng
cs.CV
Abstract
Stereo matching is a fundamental task in 3D reconstruction. Despite remarkable advances, the prevailing paradigms formulate stereo matching as a deterministic regression problem, collapsing the multimodal distribution modeling into a single-point estimation. This formulation suffers from a regression-to-mean bias, frequently struggling with ambiguous regions. In contrast, we introduce a prior-guided generative framework that integrates deterministic matching regression and generative distribution modeling within a complementary formulation. Built upon this formulation, we introduce StereoFlow through three key components: (i) a two-stage progressive cascade matching network that progressively produces multi-resolution stereo conditions with complementary matching cues; (ii) a pixel diffusion transformer (termed StereoDiT) with a frequency-decoupled architecture for modeling correspondence ambiguity; (iii) a few-step flow matching objective (termed Transition Flow Matching) for efficient optimization. In summary, \textsc{\textbf{StereoFlow}} achieves strong geometric consistency and rich fine-grained details in ill-posed, discontinuous regions and under zero-shot generalization. Extensive experiments demonstrate that the proposed StereoFlow establishes multiple state-of-the-art results across benchmarks, including Scene Flow, KITTI, ETH3D, and Middlebury.
Chinese Translation
立体匹配是三维重建中的一项基础任务。尽管取得了显著进展,现有的主流范式将立体匹配表述为一个确定性回归问题,将多模态分布建模简化为单点估计。这种表述存在回归到均值的偏差,常常在模糊区域面临挑战。相比之下,我们提出了一种基于先验引导的生成框架,将确定性匹配回归与生成分布建模整合在一个互补的框架中。在这一框架的基础上,我们通过三个关键组件引入了立体流(StereoFlow):(i) 一个两阶段渐进级联匹配网络,逐步生成具有互补匹配线索的多分辨率立体条件;(ii) 一个像素扩散变换器(称为StereoDiT),其具有频率解耦架构,用于建模对应关系的模糊性;(iii) 一个少步流匹配目标(称为过渡流匹配),用于高效优化。总之, extsc{ extbf{StereoFlow}} 在不适定、间断区域以及零样本泛化下实现了强大的几何一致性和丰富的细粒度细节。大量实验表明,所提出的StereoFlow在多个基准测试中建立了多项最先进的结果,包括Scene Flow、KITTI、ETH3D和Middlebury。

预测/规划

1
预测/规划 / 1 / 2607.20175

PerceptDrive: Perception Prior World-Action Modeling with Adaptive Expert Routing for End-to-End Autonomous Driving

PerceptDrive:基于感知先验的世界-动作建模与自适应专家路由的端到端自动驾驶
Yushan Liu, Tianxiong Lv, Bohua Wang, Hangqi Fan, Chenxu Zhao, He Zheng, Xuchang Zhong, Yifan Xie, Congyang Zhao, Zhihao Liao, Leigang Luo, Yang Cai, Xiao-Ping Zhang, Wenbo Ding
cs.CV
Abstract
Frozen perception foundation models encode rich geometric, semantic, and dynamic knowledge. Yet narrow conditioning interfaces may attenuate task-relevant cues, while static fusion cannot adjust expert contributions to each scene. We cast this challenge as the prior-to-plan transfer problem and introduce PerceptDrive, a perception prior world-action modeling framework with adaptive expert routing. PerceptDrive feeds teacher-distilled priors from a frozen, driving-adapted provider and dense observation latents from a frozen self-supervised video encoder into a trainable expert-routed world-action model. Expert-specific query branches process these signals, while a prior-retention objective anchors each branch to its prior. A router predicts soft gates from a shared scene representation and combines the expert conditions before trajectory generation. During training, privileged rule-based sub-metric estimates for branch-specific trajectory drafts provide soft-gate distillation targets. The predicted action-free future latent conditions a flow-matching actor. At inference, privileged components are absent; with one front-facing camera, PerceptDrive generates one trajectory per planning step without test-time scoring, reranking, or search. Experiments show that PerceptDrive achieves state-of-the-art performance with 90.4 PDMS on NAVSIM v1 and 90.2 EPDMS on NAVSIM v2, outperforming existing methods. Ablations confirm complementary gains from prior retention and scene-conditioned routing, alongside differential reliance on the three priors. These results demonstrate that preserving and adaptively routing perception priors improves direct planning without test-time candidate selection.
Chinese Translation
冻结的感知基础模型编码了丰富的几何、语义和动态知识。然而,狭窄的条件接口可能削弱任务相关的线索,而静态融合无法根据每个场景调整专家的贡献。我们将此挑战定义为先验到规划的迁移问题,提出了PerceptDrive,一种具有自适应专家路由的感知先验世界-动作建模框架。PerceptDrive将来自冻结且经过驾驶适配的教师蒸馏先验和来自冻结的自监督视频编码器的密集观测潜变量输入到可训练的专家路由世界-动作模型中。专家特定的查询分支处理这些信号,同时先验保持目标将每个分支锚定于其对应的先验。路由器从共享的场景表示预测软门控,并在轨迹生成前组合专家条件。训练过程中,基于规则的特权子指标对分支特定轨迹草案提供软门控蒸馏目标。预测的无动作未来潜变量条件用于流匹配执行器。推理时,无特权组件;仅使用一台前置摄像头,PerceptDrive在每个规划步骤生成一条轨迹,无需测试时评分、重排序或搜索。实验表明,PerceptDrive在NAVSIM v1上实现90.4 PDMS,在NAVSIM v2上实现90.2 EPDMS,性能优于现有方法。消融实验确认了先验保持与场景条件路由的互补增益,以及对三种先验的差异依赖。这些结果表明,保持并自适应路由感知先验能够提升直接规划性能,无需测试时候选方案选择。

协同/V2X

1
协同/V2X / 1 / 2607.19774

Defer to Plan: Adaptive Multi-Agent Fusion for End-to-End V2X Driving

延迟规划:端到端V2X驾驶的自适应多智能体融合
Nuoran Li, Zhang Zhang, Yueran Zhao, Tianze Wang, Chao Sun
cs.RO
Abstract
Vehicle-to-everything-aided autonomous driving (V2X-AD) significantly enhances driving performance through information sharing. However, existing collaborative perception methods only optimize module-level perception capabilities and fail to effectively serve the ultimate planning and control tasks. We propose an end-to-end collaborative driving system that directly optimizes planning task performance. The system employs MotionNetwork to fuse historical temporal information, utilizes attention mechanisms to efficiently compress spatial features into compact tokens, and adaptively fuses multi-agent features through an autoregressive decoder. Additionally, we introduce Mixture-of-Experts (MoE) architecture to enhance the model's representation capacity for heterogeneous features. Experiments demonstrate that our method achieves a driving score of 79.72, surpassing the state-of-the-art CoDriving baseline (77.15) by 3.33% in closed-loop evaluation while maintaining communication efficiency.
Chinese Translation
基于车与万物(V2X)辅助的自动驾驶(V2X-AD)通过信息共享显著提升了驾驶性能。然而,现有的协同感知方法仅优化模块级感知能力,未能有效服务于最终的规划和控制任务。我们提出了一种端到端的协同驾驶系统,直接优化规划任务的性能。该系统采用MotionNetwork融合历史时间信息,利用注意力机制高效地将空间特征压缩为紧凑的标记,并通过自回归解码器自适应地融合多智能体特征。此外,我们引入了专家混合(Mixture-of-Experts, MoE)架构,以增强模型对异构特征的表征能力。实验表明,我们的方法在闭环评估中实现了79.72的驾驶得分,超过了最先进的CoDriving基线(77.15)3.33%,同时保持了通信效率。