← Back to Index
Daily Autonomous Driving Research Digest

AutoDrive Papers

2026-08-04
22
Papers
6
Topics
22
Translated

感知

5
感知 / 1 / 2608.00588

InstancePin: Instance-Addressable Layout-to-Image Diffusion via Coordinate Pinning

InstancePin:通过坐标固定实现实例可寻址的布局到图像扩散
Chaoyue Wu, Yunfei Zhang, Si Wu
cs.CV
Abstract
Layout-to-image diffusion models have achieved impressive semantic controllability by conditioning generation on category-level segmentation maps. However, such category-aligned control is not necessarily instance-addressable: multiple nearby objects from the same category are often treated as a shared semantic region, leading to ambiguous boundaries, averaged appearances, and feature confusion among instances. This limitation is particularly evident in urban scene synthesis, where small and crowded pedestrians or vehicles require fine-grained instance separation while preserving global scene consistency. In this paper, we propose InstancePin, an instance-addressable layout-to-image diffusion framework that pins each object instance with an explicit coordinate anchor. Instead of directly injecting instance masks into the pretrained backbone, InstancePin introduces an independent instance-aware adapter to preserve the category-level generation prior while learning instance-specific spatial control. For each instance, its center coordinate is encoded with Fourier features and projected into a coordinate token, which serves as a spatial anchor queried by latent image features through coordinate pinning attention. To make these anchors spatially meaningful, we further supervise the coordinate attention maps with instance regions, encouraging each coordinate token to activate its corresponding object area. Finally, an instance-mask guided fusion module routes pretrained backbone features to non-instance regions and adapter features to instance regions, enabling local instance refinement without sacrificing global semantic fidelity. Extensive experiments on Cityscapes demonstrate that InstancePin mitigates instance entanglement in dense layouts and improves both image fidelity and semantic consistency.
Chinese Translation
布局到图像的扩散模型通过以类别级别的分割图为条件,实现了令人印象深刻的语义可控性。然而,这种类别对齐的控制并不一定是实例可寻址的:来自同一类别的多个相邻物体常常被视为共享的语义区域,导致模糊的边界、平均的外观以及实例之间的特征混淆。这一限制在城市场景合成中尤为明显,因为小型且拥挤的行人或车辆需要细粒度的实例分离,同时保持全局场景的一致性。本文提出了InstancePin,一种实例可寻址的布局到图像扩散框架,通过显式的坐标锚点固定每个物体实例。InstancePin并不是直接将实例掩码注入预训练的主干网络,而是引入了一个独立的实例感知适配器,以在学习实例特定的空间控制的同时保留类别级别的生成先验。对于每个实例,其中心坐标通过傅里叶特征编码并投影到一个坐标标记中,该标记作为通过坐标固定注意力查询的空间锚点。为了使这些锚点在空间上具有意义,我们进一步用实例区域监督坐标注意力图,鼓励每个坐标标记激活其对应的物体区域。最后,一个实例掩码引导的融合模块将预训练主干特征路由到非实例区域,并将适配器特征路由到实例区域,从而实现局部实例的细化,而不牺牲全局语义的保真度。在Cityscapes上的大量实验表明,InstancePin减轻了密集布局中的实例纠缠,并提高了图像的保真度和语义一致性。
感知 / 2 / 2608.01535

STAR-VLM: Spatiotemporal Grounding Vision-Language Models for Motion and Velocity Estimation via Automotive Radar Supervision

STAR-VLM:通过汽车雷达监督进行运动和速度估计的时空基础视觉-语言模型
Pou-Chun Kung, Aryaman Rao, Utkrisht Sahai, Hemanth Murali, Yi Liu, Rui-Yu Lin, Katherine A. Skinner
cs.CV · cs.RO
Abstract
Vision-language models (VLMs) are emerging as a key component of embodied intelligence, with growing applications in auto-labeling and end-to-end autonomous driving. However, existing approaches for improving spatiotemporal reasoning in VLMs often rely on complex preprocessing pipelines, expensive human annotations, or synthetic data, which limit scalability and introduce potential sim-to-real gaps. Moreover, although these methods have improved spatiotemporal understanding, they still lack strong metric reasoning capabilities for dynamic scenes, such as estimating object motion in real-world units. Prior work has explored LiDAR-based metric depth supervision to enhance spatial perception, but it does not directly address temporal reasoning. We introduce STAR-VLM, an automotive radar-supervised framework that enhances spatiotemporal VLMs with motion reasoning and metric velocity estimation for autonomous driving. Automotive radar is a low-cost and widely deployed sensor that provides complementary spatiotemporal supervision through range and Doppler measurements. By leveraging these measurements as label-free ground truth during training, STAR-VLM improves the metric spatiotemporal reasoning ability of VLMs. Through experiments on driving scenarios, we show that STAR-VLM achieves state-of-the-art performance on both motion classification and metric velocity estimation, outperforming even task-specific methods designed for each task. These results highlight automotive radar as a scalable and cost-effective source of supervision for building metric-aware spatiotemporal VLMs for real-world autonomous driving.
Chinese Translation
视觉-语言模型(VLMs)正成为具身智能的关键组成部分,广泛应用于自动标注和端到端的自主驾驶。然而,现有的提升VLMs时空推理的方法往往依赖于复杂的预处理流程、昂贵的人类标注或合成数据,这限制了其可扩展性并引入了潜在的模拟与现实之间的差距。此外,尽管这些方法改善了时空理解,但在动态场景中仍缺乏强大的度量推理能力,例如在现实单位中估计物体运动。之前的研究探索了基于激光雷达的度量深度监督以增强空间感知,但未能直接解决时间推理问题。我们提出了STAR-VLM,一个汽车雷达监督框架,增强了VLMs的时空推理能力,并实现了自主驾驶的运动推理和度量速度估计。汽车雷达是一种低成本且广泛部署的传感器,通过距离和多普勒测量提供互补的时空监督。通过在训练过程中利用这些测量作为无标签的真实值,STAR-VLM提高了VLMs的度量时空推理能力。通过在驾驶场景中的实验,我们展示了STAR-VLM在运动分类和度量速度估计方面实现了最先进的性能,甚至超越了为每个任务设计的特定任务方法。这些结果突显了汽车雷达作为构建面向度量的时空VLMs在现实世界自主驾驶中的可扩展和经济有效的监督来源。
感知 / 3 / 2608.02177

GSRAIN: Physically Calibrated High-/Low-Frequency Rainfall Synthesis for 3D Gaussian Driving Scenes

GSRAIN:物理校准的高/低频降雨合成用于3D高斯驱动场景
Fanyu Wang, Longgao Zhang, Junyi Chen
cs.CV
Abstract
Existing rainfall simulation methods for autonomous driving remain limited in physical controllability and multi-view consistency. This paper presents GSRAIN, a high-/low-frequency rainfall synthesis method for 3D Gaussian Splatting (3DGS) driving scenes. GSRAIN constructs a high-frequency raindrop model from measured rainfall data and generates low-frequency rainy appearance using a geometry-aware single-step diffusion model. The two effects are then fused in a unified 3DGS scene, enabling rainfall-intensity control over the range of 0--13~mm/h. The proposed method achieves a Fr\'{e}chet Inception Distance (FID) of 149.09, outperforming CycleGAN-Turbo (155.71) and WeatherEdit (157.94). Object-detection and closed-loop driving experiments further show that the generated scenes expose scene-dependent performance changes of the evaluated algorithms under controllable rainfall. These results indicate that GSRAIN provides an effective approach for constructing physically controllable, repeatable, and closed-loop-compatible rainy-weather test scenes for autonomous driving.
Chinese Translation
现有的自主驾驶降雨模拟方法在物理可控性和多视角一致性方面仍然有限。本文提出了GSRAIN,一种用于3D高斯点云(3DGS)驾驶场景的高/低频降雨合成方法。GSRAIN从测量的降雨数据中构建高频雨滴模型,并使用几何感知的单步扩散模型生成低频雨天外观。然后将这两种效果融合在统一的3DGS场景中,实现了对降雨强度在0--13 mm/h范围内的控制。所提出的方法达到了149.09的Fréchet Inception Distance (FID),优于CycleGAN-Turbo(155.71)和WeatherEdit(157.94)。物体检测和闭环驾驶实验进一步表明,生成的场景在可控降雨下暴露了评估算法的场景依赖性性能变化。这些结果表明,GSRAIN为构建物理可控、可重复且兼容闭环的雨天测试场景提供了一种有效的方法,适用于自主驾驶。
感知 / 4 / 2608.02200

RSC-GestureNet: Reliability-Aware Selective Causal Recognition of Chinese Traffic Police Gestures

RSC-GestureNet:可靠性感知的中国交通警察手势选择性因果识别
Cheng Li, Renjun Gao, Boyi Fu
cs.CV
Abstract
Traffic police gestures are safety-critical perception cues for autonomous driving. A deployable recognizer must infer commands causally from continuous full-frame video, remain stable around transitional arm motion, and avoid over-trusting corrupted pose measurements. This study presents RSC-GestureNet, a reliability-aware selective causal recognizer, for Chinese traffic police gestures. The model treats pose confidence as a first-class signal: unreliable joints are down weighted during graph reasoning, temporal evidence is aggregated causally, and calibrated predictions are selectively emitted through a reliability-aware inference rule. We further introduce CTPGesture-C, a reproducible feature-level corruption benchmark with seven pose/RGB degradation families, and an RGB-level diagnostic in which corrupted frames are reprocessed by MediaPipe before recognition. On the complete official CTPGesture v1 split (134,424 labeled frames and 33,451 causal windows), RSC-GestureNet achieves 93.33+-0.24% accuracy, 91.71+-0.27% macro-F1, 91.69+-0.29% online macro-F1, 98.80+-0.07% Early@10, 0.153+-0.013 s TTC, and the best robust macro-F1 among evaluated methods. Under the same split and causal protocol, it exceeds reproduced traffic-specific MD-GCN and HLP-GCN baselines by 3.23-4.11 macro-F1 points and 2.15-3.07 online-F1 points. These results, together with calibration, selective-risk, statistical, adaptive-branching, and image-level re-extraction analyses, indicate that explicit pose-reliability modeling improves early, stable, and robust traffic-command recognition.
Chinese Translation
交通警察手势是自动驾驶中安全关键的感知线索。一个可部署的识别器必须能够从连续的全帧视频中因果推断命令,在过渡的手臂运动中保持稳定,并避免对损坏的姿态测量过于信任。本研究提出了RSC-GestureNet,一种针对中国交通警察手势的可靠性感知选择性因果识别器。该模型将姿态置信度视为一类重要信号:在图推理过程中,对不可靠的关节进行降权,因果地聚合时间证据,并通过可靠性感知推理规则选择性地输出校准预测。我们进一步引入了CTPGesture-C,这是一个可重复的特征级损坏基准,包含七种姿态/RGB降级类别,以及一个RGB级诊断,其中损坏的帧在识别前通过MediaPipe重新处理。在完整的官方CTPGesture v1划分(134,424个标记帧和33,451个因果窗口)上,RSC-GestureNet达到了93.33±0.24%的准确率,91.71±0.27%的宏F1,91.69±0.29%的在线宏F1,98.80±0.07%的Early@10,0.153±0.013秒的TTC,以及在评估方法中最佳的鲁棒宏F1。在相同的划分和因果协议下,它超越了再现的交通特定MD-GCN和HLP-GCN基线,宏F1提高了3.23-4.11点,在线F1提高了2.15-3.07点。这些结果,以及校准、选择性风险、统计、适应性分支和图像级重新提取分析,表明明确的姿态可靠性建模改善了早期、稳定和鲁棒的交通命令识别。
感知 / 5 / 2608.02449

MoRAL: Sensor-Grounded BEV Reasoning for Compact VLMs toward Edge-Oriented Autonomous Driving

MoRAL:面向边缘自主驾驶的传感器基础鸟瞰视图推理紧凑型视觉语言模型
Ambarish Govindarajulu Kaliamurthi, Kaikai Liu
cs.CV · cs.RO
Abstract
Deploying vision-language models (VLMs) for safety-critical spatial reasoning on resource-constrained autonomous driving platforms requires both compact model size and reliable metric grounding. We present MoRAL (Multimodal Reasoning for Autonomous Language Models), a two-stage fine-tuning pipeline that teaches Cosmos-Reason2-2B to first read a physics-encoded Bird's Eye View (BEV) representation and then reason over it for driving decisions. The BEV image encodes LiDAR metric distance as color bands, object class as cluster morphology, and radar Doppler velocity as directional wedge overlays, externalizing spatial perception into the input image so that no learned 3D backbone is required at inference. Stage 1 fine-tunes the vision encoder on 60,000 grounding records; zero-shot baselines produce no parseable BEV outputs, confirming the vocabulary requires explicit training. Stage 2 fine-tunes the full model (52M parameters, 2.4% of total) on 57,696 chain-of-thought records generated by Cosmos-Reason2-8B as teacher, spanning eight driving question types. On 2,304 held-out nuScenes frames evaluated by Gemma 4 (31B) calibrated against human review, MoRAL wins seven of eight question types over a zero-shot 8B baseline despite using four times fewer parameters, with the largest margins on question types requiring structured multi-step physics reasoning. Emergency braking recall improves from 10.8% to 47.8%, output degeneration falls from 94.1% to 20.8%, and the full pipeline fits a consumer 8 GB GPU at 42 tok/s without quantization. These results establish a reproducible foundation for compact, physics-grounded VLM reasoning on mobile edge platforms.
Chinese Translation
在资源受限的自主驾驶平台上部署视觉语言模型(VLM)以进行安全关键的空间推理,需要模型体积紧凑且度量基础可靠。我们提出了MoRAL(多模态推理用于自主语言模型),这是一种两阶段的微调管道,旨在教会Cosmos-Reason2-2B首先读取物理编码的鸟瞰视图(BEV)表示,然后基于此进行驾驶决策推理。BEV图像将激光雷达的度量距离编码为颜色带,将物体类别编码为聚类形态,将雷达多普勒速度编码为方向楔形叠加,从而将空间感知外部化到输入图像中,使得推理时无需学习的3D骨干网络。第一阶段在60,000个基础记录上微调视觉编码器;零样本基线未产生可解析的BEV输出,确认词汇需要明确的训练。第二阶段在57,696个由Cosmos-Reason2-8B生成的思维链记录上微调完整模型(5200万参数,占总数的2.4%),涵盖八种驾驶问题类型。在经过人类审查校准的Gemma 4(310亿)评估的2304个保留nuScenes帧中,尽管使用的参数数量是零样本8B基线的四分之一,MoRAL在八种问题类型中赢得了七种,且在需要结构化多步骤物理推理的问题类型上具有最大的优势。紧急制动的召回率从10.8%提高到47.8%,输出退化率从94.1%降低到20.8%,整个管道在不进行量化的情况下以42 tok/s的速度适配消费者8 GB GPU。这些结果为在移动边缘平台上进行紧凑的、基于物理的VLM推理奠定了可重复的基础。

BEV/Occupancy

2
BEV/Occupancy / 1 / 2608.01338

Driver2Map: Imitating Human Driving for Online High-Definition Map Construction

Driver2Map:模仿人类驾驶进行在线高清地图构建
Pan Yin, Runtian Xia, Weisong Kuang, Kaiyu Li, Cong Zhao, Xiangyong Cao
cs.CV
Abstract
High-definition (HD) maps are essential for autonomous driving systems. In constructing such maps, onboard multi-view camera images, standard-definition maps and satellite images provide crucial information. However, due to the modality and perspective differences among these data sources, existing methods often struggle to effectively align and fuse them, making online HD map construction still challenging. To address these issues, we propose Driver2Map, an online HD map construction model inspired by human drivers. Unlike existing HD map construction models that utilize only two modalities, our Driver2Map can simultaneously exploit three modalities. Specifically, we propose a "two-stage alignment" strategy to reduce spatial misalignment across different modalities. Additionally, we introduce "Pose-Guided BEV Fusion", a BEV (bird's-eye-view) generation module that leverages camera pose information to adaptively weight multi-view features, thereby effectively suppressing cross-view feature overlap during BEV generation. Also, we design a "Pretrained Prior for Map Refinement" module to refine the initial prediction by learning map structure priors, thus improving the HD map prediction under dynamic occlusions. Extensive experiments demonstrate that Driver2Map outperforms existing methods on both IoU and AP metrics.
Chinese Translation
高清(HD)地图对于自动驾驶系统至关重要。在构建此类地图时,车载多视角相机图像、标清地图和卫星图像提供了重要信息。然而,由于这些数据源之间的模态和视角差异,现有方法往往难以有效对齐和融合它们,使得在线高清地图构建仍然面临挑战。为了解决这些问题,我们提出了Driver2Map,这是一种受人类驾驶启发的在线高清地图构建模型。与仅利用两种模态的现有高清地图构建模型不同,我们的Driver2Map能够同时利用三种模态。具体而言,我们提出了一种“二阶段对齐”策略,以减少不同模态之间的空间错位。此外,我们引入了“姿态引导的鸟瞰图融合”(Pose-Guided BEV Fusion),这是一个利用相机姿态信息自适应加权多视角特征的鸟瞰图(BEV)生成模块,从而在BEV生成过程中有效抑制视角间特征重叠。同时,我们设计了一个“用于地图精细化的预训练先验”模块,通过学习地图结构先验来精细化初始预测,从而提高动态遮挡下的高清地图预测。大量实验表明,Driver2Map在IoU和AP指标上均优于现有方法。
BEV/Occupancy / 2 / 2608.02309

CalibBEV: LiDAR-Camera Calibration via BEV Alignment

CalibBEV:基于鸟瞰视角(BEV)对LiDAR与相机的标定
Filippo D'Addeo, Lorenzo Cipelli, Adriano Cardace, Emanuele Ghelfi, Andrea Zinelli, Massimo Bertozzi
cs.CV
Abstract
We present CalibBEV, a novel Bird's Eye View (BEV) alignment approach for LiDAR-camera calibration. Our method unifies LiDAR and camera data into a shared 3D spatial representation, enabling accurate and robust cross-modal calibration. CalibBEV extracts sensor-wise BEV features from each modality using domain-specific architectures and estimates the calibration matrix through a two-step alignment process. First, we perform an implicit alignment by regressing a coarse calibration matrix directly from the BEV features. To ease this alignment, we enforce semantic consistency between BEV representations across modalities using a contrastive loss inspired by CLIP, guiding both networks toward a unified feature space. In the second step, we leverage our BEV formulation to explicitly align the features of one modality with the other, refining the initial coarse estimate into a final, more accurate calibration matrix. CalibBEV significantly outperforms prior point-to-pixel matching methods, achieving state-of-the-art calibration accuracy. On the KITTI and nuScenes benchmarks, our method reduces the Relative Rotation Error (RRE) by 51% and 68%, and the Relative Translation Error (RTE) by 80% and 91%, respectively, compared to previous methods.
Chinese Translation
我们提出了CalibBEV,一种新颖的基于鸟瞰视角(BEV)对LiDAR与相机进行标定的方法。我们的方法将LiDAR和相机数据统一为一个共享的三维空间表示,从而实现准确且稳健的跨模态标定。CalibBEV通过特定领域的架构从每种模态中提取传感器级的BEV特征,并通过两步对齐过程估计标定矩阵。首先,我们通过直接从BEV特征回归一个粗略的标定矩阵来执行隐式对齐。为了简化这一对齐过程,我们使用受CLIP启发的对比损失强制不同模态间的BEV表示保持语义一致性,引导两个网络朝着统一的特征空间发展。在第二步中,我们利用我们的BEV公式显式地将一种模态的特征与另一种模态对齐,将初始的粗略估计精炼为最终的更准确的标定矩阵。CalibBEV在标定精度上显著优于先前的点到像素匹配方法,达到了最先进的标定精度。在KITTI和nuScenes基准测试中,我们的方法分别将相对旋转误差(RRE)降低了51%和68%,将相对平移误差(RTE)降低了80%和91%,相比于之前的方法。

深度/几何

5
深度/几何 / 1 / 2608.02320

TravKAN: Fast and Interpretable Nonlinear Traversability Analysis with Kolmogorov-Arnold Networks

TravKAN:基于Kolmogorov-Arnold网络的快速且可解释的非线性可通行性分析
Daniel Fusaro, Simone Mosco, Wanmeng Li, Alberto Pretto
cs.RO · cs.CV
Abstract
Traversability analysis is a fundamental capability for autonomous mobile robots operating in unstructured environments. While modern machine learning approaches such as deep neural networks and gradient-boosted trees achieve strong predictive performance, they lack interpretability and provide limited insight into the underlying terrain-robot interaction dynamics. In this paper, we propose TravKAN, a Kolmogorov-Arnold Network-based framework for fast, scalable, and interpretable traversability estimation. TravKAN represents multivariate decision functions through compositions of learnable univariate functions, enabling compact architectures and symbolic extraction of analytic expressions after training. In addition, we introduce a novel set of handcrafted features derived from the reflectivity channel of LiDAR sensors. To the best of our knowledge, reflectivity has not been systematically exploited for handcrafted traversability descriptors, despite its potential to capture material and surface properties complementary to geometric cues. We evaluate TravKAN on public, real-world urban and off-road datasets and compare it against strong baselines. TravKAN achieves strong performance across all metrics, outperforming conventional deep models and approaching the performance of XGBoost. TravKAN-Lite, i.e., TravKAN's symbolic representation, reveals meaningful nonlinear feature interactions and provides a compact, deployment-friendly, and fast analytic model. Ablation studies further show the robustness of our method to architectural variations and quantify the contribution of the proposed reflectivity-based features. These properties make TravKAN attractive for robotic systems requiring transparency, real-time computational efficiency, and interpretability in safety-critical decision-making.
Chinese Translation
可通行性分析是自主移动机器人在非结构化环境中操作的基本能力。尽管现代机器学习方法如深度神经网络和梯度提升树在预测性能上表现优异,但它们缺乏可解释性,且对基础的地形-机器人交互动态提供的洞察有限。本文提出了TravKAN,一个基于Kolmogorov-Arnold网络的框架,用于快速、可扩展且可解释的可通行性估计。TravKAN通过可学习的单变量函数的组合表示多变量决策函数,使得在训练后能够实现紧凑的架构和符号提取的解析表达。此外,我们引入了一组新颖的手工特征,这些特征源自LiDAR传感器的反射率通道。尽我们所知,反射率尚未被系统性地用于手工可通行性描述符,尽管它在捕捉材料和表面特性方面具有潜力,这些特性与几何线索互为补充。我们在公共的真实世界城市和越野数据集上评估了TravKAN,并与强基线进行了比较。TravKAN在所有指标上均表现出色,超越了传统的深度模型,并接近XGBoost的性能。TravKAN-Lite,即TravKAN的符号表示,揭示了有意义的非线性特征交互,并提供了一个紧凑、适合部署且快速的解析模型。消融研究进一步表明我们的方法对架构变化的鲁棒性,并量化了所提反射率特征的贡献。这些特性使TravKAN在需要透明性、实时计算效率和安全关键决策可解释性的机器人系统中具有吸引力。
深度/几何 / 2 / 2608.00237

Latent-Centroid Steering: Single-Pass Classifier-Free Guidance for Command-Aligned Autonomous Driving

潜在中心引导:命令对齐自主驾驶的单次通行无分类器引导
Meibo Hu, Jiamian Wang, Pichao Wang, Zhiqiang Tao
cs.CV · cs.RO
Abstract
Vision-language models (VLMs) have recently emerged as a promising paradigm for end-to-end autonomous driving, enabling agents to map multimodal inputs and high-level navigation instructions directly to executable trajectories. However, in practice, these models exhibit a persistent command-following gap: predicted trajectories often show weak sensitivity to navigation commands, resulting in incorrect behavior at critical decision points. We identify this issue as a form of conditional policy collapse, where regression-based training under multimodal trajectory distributions encourages the model to rely on dominant visual priors while marginalizing the language-conditioned signal. To address this issue, we introduce a principled formulation of classifier-free guidance (CFG) for regression-based vision-language driving. We show that CFG can be interpreted as isolating the instruction-induced residual in the action space by contrasting conditional and unconditional predictions, thereby explicitly amplifying the effect of the navigation command at inference time. However, a standard two-pass CFG introduces prohibitive latency for real-time control and produces noisy instance-level guidance directions. Building on a mean-shift interpretation of CFG, we propose Latent-Centroid Steering (LCS), a single-pass guidance mechanism that replaces instance-level residuals with class-level latent shifts. By projecting conditional representations toward precomputed command-specific centroids, LCS performs class-level latent steering based on cluster geometry that is both more stable and computationally efficient. We demonstrate that LCS reduces inference latency by approximately 50% while achieving stronger command adherence and improved driving performance on both closed-loop (Bench2Drive) and open-loop (nuScenes) benchmarks. Code will be released.
Chinese Translation
视觉-语言模型(VLMs)最近作为一种有前景的端到端自主驾驶范式出现,使得智能体能够将多模态输入和高层次导航指令直接映射到可执行轨迹。然而,在实践中,这些模型表现出持续的命令跟随差距:预测的轨迹往往对导航指令的敏感性较弱,导致在关键决策点出现不正确的行为。我们将此问题识别为一种条件策略崩溃的形式,其中基于回归的训练在多模态轨迹分布下促使模型依赖于主导的视觉先验,同时边缘化语言条件信号。为了解决这个问题,我们引入了一种无分类器引导(CFG)的原则性公式,用于基于回归的视觉-语言驾驶。我们展示了CFG可以被解释为通过对比条件和无条件预测来隔离动作空间中的指令诱导残差,从而在推理时明确放大导航指令的效果。然而,标准的双次CFG为实时控制引入了过高的延迟,并产生了嘈杂的实例级引导方向。在对CFG的均值漂移解释的基础上,我们提出了潜在中心引导(LCS),这是一种单次通行的引导机制,它用类级潜在偏移替代实例级残差。通过将条件表示投影到预计算的命令特定中心,LCS基于聚类几何进行类级潜在引导,这种方法更加稳定且计算效率更高。我们证明LCS将推理延迟减少了约50%,同时在闭环(Bench2Drive)和开环(nuScenes)基准测试上实现了更强的命令遵循和改进的驾驶性能。代码将会发布。
深度/几何 / 3 / 2608.00687

Proteus: A Truncation-Robust Entropy Model for Progressive LiDAR Compression

Proteus:一种针对渐进式LiDAR压缩的截断鲁棒熵模型
Yihan Qiu, Xiaodong Lin, Baoquan Zhao, Hailong Jiao, Ge Li
cs.CV
Abstract
LiDAR point clouds provide explicit, deterministic physical boundaries critical for collaborative safety-critical perception. However, wireless channels inherently impair and corrupt transmitted signals. Existing robust frameworks (such as deep JSCC or MDC) attempt to counter these channel impairments through statistical or parametric estimation, turning exact physical measurements into unverified algorithmic estimates. To address this, we propose Proteus, a learned LiDAR codec operating on 2D range images. By decoupling the frame representation into independent coders for the \textbf{sig}nificant range bit-planes (SIG) and the \textbf{ins}ignificant range bit-planes and attributes (INS), Proteus achieves overall stream-level truncation robustness. The non-truncatable SIG block encodes the most significant range bit-planes to establish a necessary, self-contained perceptual lower bound, below which the reconstructed point cloud is severely degraded. Meanwhile, INS employs bit-plane slicing representation and coding, ensuring that range truncation mathematically maps to a deterministic spatial precision degradation. Subordinate attributes are reconstructed via a hybrid lossless-predictive method, leveraging the decoded geometry as a strong structural prior for fine-grained approximation. Furthermore, strategic ordering within INS prioritizes geometry over attributes under bandwidth drops. Experimental results on the Waymo Open Dataset and SemanticKITTI demonstrate that Proteus tolerates up to approximately 70\% bitstream truncation, while outperforming established standards (G-PCC, Draco, and JPEG XL) and the representative learned compressor Unicorn under ideal channel conditions.
Chinese Translation
LiDAR点云提供了明确的、确定性的物理边界,这对于协作安全感知至关重要。然而,无线信道固有地会损害和破坏传输信号。现有的鲁棒框架(如深度JSCC或MDC)试图通过统计或参数估计来抵消这些信道损害,将精确的物理测量转变为未经验证的算法估计。为了解决这个问题,我们提出了Proteus,一种基于2D范围图像的学习型LiDAR编解码器。通过将帧表示解耦为独立的编码器,分别针对显著范围比特平面(SIG)和不显著范围比特平面及属性(INS),Proteus实现了整体流级别的截断鲁棒性。不可截断的SIG块编码了最显著的范围比特平面,以建立一个必要的、自包含的感知下限,低于该下限,重建的点云将严重退化。同时,INS采用比特平面切片表示和编码,确保范围截断在数学上映射为确定性的空间精度退化。附属属性通过混合无损预测方法重建,利用解码后的几何形状作为细粒度近似的强结构先验。此外,INS中的战略排序在带宽下降时优先考虑几何形状而非属性。在Waymo开放数据集和SemanticKITTI上的实验结果表明,Proteus能够容忍约70%的比特流截断,同时在理想信道条件下超越了现有标准(G-PCC、Draco和JPEG XL)以及代表性的学习压缩器Unicorn。
深度/几何 / 4 / 2608.01761

DecoupleGS: Interactive 3D Gaussian Splatting for End-to-End Autonomous Driving Testing

DecoupleGS:用于端到端自主驾驶测试的交互式3D高斯点云渲染
Siying Li, Ying Ni, Jie Sun, Jian Sun, Haotian Shi
cs.CV
Abstract
End-to-end (E2E) autonomous driving algorithms require rigorous closed-loop validation in simulation environments offering high visual fidelity, strong interactivity, and real-time performance. Existing approaches, from game engines to static neural rendering, inherently trade off these requirements and struggle with the dynamic scene composition essential for E2E testing. To bridge this gap, we propose a novel decoupled 3D Gaussian Splatting (3DGS) framework tailored for large-scale E2E evaluation. We fundamentally decompose scenes into a high-fidelity static background and manipulable dynamic agents using an object-centric canonical representation. To resolve resulting representational conflicts, we introduce three targeted modules: (1) asset compression via perceptual pruning and vector quantization for real-time traffic rendering; (2) map-guided geometric registration leveraging semantic topology to strictly align trajectories; and (3) proxy-based relighting transferring ambient illumination for seamless photometric integration. Extensive experiments demonstrate that DecoupleGS achieves a balanced fidelity-efficiency trade-off, improves metric and photometric consistency, and provides a practical closed-loop sensor simulation platform for E2E autonomous driving evaluation.
Chinese Translation
端到端(E2E)自主驾驶算法需要在提供高视觉保真度、强交互性和实时性能的仿真环境中进行严格的闭环验证。现有的方法,从游戏引擎到静态神经渲染,固有地在这些要求之间进行权衡,并在E2E测试所需的动态场景组合方面面临挑战。为了解决这一问题,我们提出了一种新颖的解耦3D高斯点云渲染(3DGS)框架,专为大规模E2E评估而设计。我们从根本上将场景分解为高保真静态背景和可操控的动态代理,采用以对象为中心的标准表示法。为了解决由此产生的表示冲突,我们引入了三个针对性的模块:(1)通过感知剪枝和向量量化实现资产压缩,以实现实时交通渲染;(2)利用语义拓扑的地图引导几何注册,严格对齐轨迹;(3)基于代理的重光照,转移环境光照以实现无缝的光度集成。大量实验表明,DecoupleGS实现了保真度与效率的平衡权衡,提高了度量和光度一致性,并提供了一个实用的闭环传感器仿真平台,用于E2E自主驾驶评估。
深度/几何 / 5 / 2608.02191

DerainSplat: Feed-Forward Clean 3D Gaussian Splatting from Sparse Rainy Views

DerainSplat:从稀疏雨天视图中前馈清晰的3D高斯点云重建
Fuzhen Jiang, Changyue Shi, Chuxiao Yang, Xinyuan Hu, Wenjie Ye, Minghao Chen
cs.CV
Abstract
Although image deraining has advanced substantially, existing methods mainly focus on 2D image restoration. As spatial intelligence applications such as embodied AI and autonomous driving continue to emerge, reconstructing clean 3D scenes from sparse rainy views in a feed-forward manner becomes increasingly important. Existing feed-forward 3D Gaussian Splatting (3DGS) methods often assume clean inputs and collapse under rainy conditions. To this end, we present \textbf{\textit{DerainSplat}}, a feed-forward framework that reconstructs clean 3D scenes from only a few rainy views. To support this task, we build a large-scale multi-view derain dataset through a four-stage synthesis pipeline that sequentially models overcast illumination, depth-dependent haze, rain streaks, and lens raindrops, producing privileged weather factors. We introduce a weather net that predicts the weather factors from rainy context and yields two support maps. Scene support modulates cross-view cost-volume matching, while radiance support drives depth-aligned appearance fusion to fill corrupted pixels. The derived geometry evidence further attenuates Gaussian opacity to reduce spurious structures. A rainy cycle consistency re-renders clean views using the predicted factors and aligns them with rainy inputs. Extensive experiments show that \textbf{\textit{DerainSplat}} outperforms existing methods on various datasets, including RealEstate10K, ACID, Mip-NeRF360, and real-world rainy scenes, with strong cross-dataset generalization.
Chinese Translation
尽管图像去雨技术已经取得了显著进展,但现有方法主要集中在2D图像恢复上。随着具身人工智能和自动驾驶等空间智能应用的不断涌现,从稀疏的雨天视图中以前馈方式重建清晰的3D场景变得愈发重要。现有的前馈3D高斯点云重建(3DGS)方法通常假设输入为清晰图像,在雨天条件下会失效。为此,我们提出了 extbf{ extit{DerainSplat}},这是一个前馈框架,能够仅从少量雨天视图中重建清晰的3D场景。为了支持这一任务,我们通过四阶段合成管道构建了一个大规模的多视角去雨数据集,该管道依次模拟阴天照明、深度依赖雾霭、雨滴条纹和镜头雨滴,生成特权天气因子。我们引入了一个天气网络,该网络从雨天上下文中预测天气因子,并生成两个支持图。场景支持调节跨视图成本体积匹配,而辐射支持驱动深度对齐的外观融合,以填补受损像素。推导出的几何证据进一步减弱高斯不透明度,以减少虚假结构。雨天循环一致性使用预测因子重新渲染清晰视图,并将其与雨天输入对齐。大量实验表明, extbf{ extit{DerainSplat}}在多个数据集上超越了现有方法,包括RealEstate10K、ACID、Mip-NeRF360和真实世界的雨天场景,并展现出强大的跨数据集泛化能力。

预测/规划

7
预测/规划 / 1 / 2608.00113

Track-Guided Hierarchical Reinforcement Learning for Autonomous Vehicle Drifting with Minimum-Lap-Time Planning

基于轨迹引导的层次强化学习在自主车辆漂移中的最小圈速规划
Sheng Zhao, Bolin Zhao, Xiaodong Wu, Chen Lv
cs.RO
Abstract
In Formula 1, drivers optimize racing lines within tire grip limits to minimize lap times; however, in rally racing, drivers intentionally break traction to drift on loose surfaces. This maneuver rapidly aligns the vehicle for corner exits, ultimately reducing lap time. Autonomously executing such maneuvers formulates a complex dual-objective control problem: stabilizing highly nonlinear drift dynamics while strictly minimizing lap time. Addressing this challenge motivates the development of advanced Minimum-Lap-Time (MLT) drift control architectures. This paper proposes a planning-control framework specifically designed for MLT drifting scenario. First, we formulate an optimal control problem to generate a MLT drift planning trajectory, which is used as prior data to train a deep reinforcement learning drift controller. Given that drifting involves extremely large sideslip angles and is therefore challenging to learn directly, a Track-guided Reinforcement Learning (TgRL) drift control method is proposed to enable progressive training in a step-by-step manner, from drift control policy, to drift corner policy, and finally to a comprehensive drift race policy. The reward function incorporates both an instant reward term and an end reward term derived from the Minimum-Lap-Time objective. Simulation results demonstrate that the proposed framework enables the agent to learn a drift racing policy that not only ensures vehicle motion control performance but also effectively reduces lap time.
Chinese Translation
在一级方程式赛车中,车手在轮胎抓地力的限制内优化赛车线路,以最小化圈速;然而,在拉力赛中,车手故意打滑以在松散的路面上漂移。这一操作迅速使车辆在转角出口处对齐,从而最终减少圈速。自主执行此类操作形成了一个复杂的双重目标控制问题:在严格最小化圈速的同时稳定高度非线性的漂移动态。应对这一挑战促使了先进的最小圈速(Minimum-Lap-Time, MLT)漂移控制架构的发展。本文提出了一种专门为MLT漂移场景设计的规划控制框架。首先,我们构建了一个最优控制问题,以生成MLT漂移规划轨迹,该轨迹作为先验数据用于训练深度强化学习漂移控制器。考虑到漂移涉及极大的侧滑角,因此直接学习具有挑战性,提出了一种轨迹引导的强化学习(Track-guided Reinforcement Learning, TgRL)漂移控制方法,以逐步训练的方式,从漂移控制策略到漂移转角策略,最终到全面的漂移比赛策略。奖励函数结合了即时奖励项和源自最小圈速目标的最终奖励项。仿真结果表明,所提出的框架使代理能够学习到一种漂移比赛策略,不仅确保了车辆运动控制性能,还有效减少了圈速。
预测/规划 / 2 / 2608.00625

Learning-Based Motion Planning for Dynamic Environments: From Foundational Algorithms to Emerging Paradigms

基于学习的动态环境运动规划:从基础算法到新兴范式
Zongyuan Shen, Shalabh Gupta, Shancheng Zhao, Dehua Zhou, Gao Wang, Rui Cheng, Yaming Ou, Zhongqiang Ren, Yikui Zhai, C. L. Philip Chen
cs.RO · cs.AI
Abstract
Motion planning in dynamic environments is a fundamental problem in robotics, aiming to generate safe and efficient paths, trajectories, or control actions in the presence of moving obstacles, uncertain predictions, and multi-agent interactions. It has broad applications in autonomous driving, service robotics, warehouse logistics, human-robot collaboration, crowd navigation, and multi-robot systems. This survey reviews representative works published primarily between 2015 and 2025, with a particular focus on how recent learning-based advances extend, complement, or interact with classical planning foundations. We first revisit classical planning methods as algorithmic foundations and reference frameworks for learning-based extensions. We then propose a role-of-learning taxonomy that categorizes existing methods according to how learning participates in the planning pipeline, including direct policy learning, learning-augmented classical planning, hybrid planning, and training enhancement methods. For each category, we summarize the main problem settings, representative algorithms, key ideas, integration mechanisms, strengths, and limitations. We further analyze how observation representations, prediction uncertainty, interaction modeling, planner integration, safety constraints, and training strategies shape learning-based motion planning in dynamic environments. Finally, we discuss open challenges and future directions, including sim-to-real gap, safe and certifiable planning, dense crowd navigation, perception-planning coupling, and embodied AI.
Chinese Translation
动态环境中的运动规划是机器人学中的一个基本问题,旨在在移动障碍物、不确定预测和多智能体交互的情况下生成安全和高效的路径、轨迹或控制动作。它在自动驾驶、服务机器人、仓储物流、人机协作、拥挤导航和多机器人系统等领域具有广泛的应用。本文综述了2015年至2025年间发表的代表性研究,特别关注近期基于学习的进展如何扩展、补充或与经典规划基础相互作用。我们首先回顾经典规划方法,作为算法基础和基于学习的扩展的参考框架。然后,我们提出了一种学习角色分类法,根据学习在规划流程中的参与方式对现有方法进行分类,包括直接策略学习、学习增强的经典规划、混合规划和训练增强方法。对于每个类别,我们总结了主要问题设置、代表性算法、关键思想、集成机制、优点和局限性。我们进一步分析了观察表示、预测不确定性、交互建模、规划者集成、安全约束和训练策略如何塑造动态环境中的基于学习的运动规划。最后,我们讨论了开放挑战和未来方向,包括模拟与现实之间的差距、安全和可证明的规划、密集人群导航、感知-规划耦合和具身人工智能。
预测/规划 / 3 / 2608.00779

SIPTraj: Map-Free End-to-End Trajectory Prediction via Physics-Guided Scene Interaction

SIPTraj:通过物理引导的场景交互实现无地图的端到端轨迹预测
Feifei Liu, Zejun Wei, Haozhe Wang, Yazhi Ye, Yuying Zhang, Jintao Cheng, Chi Man Vong, Xieyuanli Chen, Xiaoyu Tang
cs.RO
Abstract
Trajectory prediction of surrounding agents is a prerequisite for safe planning and decision making in autonomous driving. Without high-definition (HD) maps, sensor-derived bird's-eye-view (BEV) features provide no explicit lane topology or drivable-area priors, making it inherently difficult to ground each agent in its surrounding scene context. Moreover, physical feasibility remains difficult to capture through data-driven learning alone, as kinematic constraints on agent motion cannot be explicitly encoded without structured supervision. Existing map-free predictors extract scene context in an agent-agnostic manner through a single fusion step and treat physical constraints only as output-level penalties, leaving both challenges unaddressed. We propose SIPTraj, a map-free trajectory prediction framework that jointly addresses scene grounding and physical feasibility. SIPTraj introduces a Hierarchical Agent-Scene Encoder (HASE) progressively grounding each agent in agent-guided scene evidence and refining inter-agent relations within the scene-grounded space. To tackle physical infeasibility in predicted trajectories, we develop a Physics-Guided Iterative Decoder (PGID). It conditions decoding on instantaneous kinematic states, propagating physical supervision into internal representations rather than output trajectories alone. Extensive experiments on nuScenes and Argoverse 2 Sensor show that SIPTraj surpasses prior map-free predictors and strong map-based baselines without any HD map at inference. Our code will be released as open-source.
Chinese Translation
周围智能体的轨迹预测是自动驾驶中安全规划和决策的前提。没有高清(HD)地图时,传感器获取的鸟瞰图(BEV)特征无法提供明确的车道拓扑或可行驶区域先验,使得将每个智能体定位于其周围场景上下文中变得极为困难。此外,仅通过数据驱动学习捕捉物理可行性也很困难,因为智能体运动的运动学约束无法在没有结构化监督的情况下明确编码。现有的无地图预测器通过单一融合步骤以与智能体无关的方式提取场景上下文,并仅将物理约束视为输出级别的惩罚,从而未能解决这两个挑战。我们提出了SIPTraj,一种无地图的轨迹预测框架,联合解决场景定位和物理可行性问题。SIPTraj引入了分层智能体-场景编码器(HASE),逐步将每个智能体定位于智能体引导的场景证据中,并在场景定位空间内细化智能体之间的关系。为了应对预测轨迹中的物理不可行性,我们开发了物理引导的迭代解码器(PGID)。它将解码条件化于瞬时运动学状态,将物理监督传播到内部表示中,而不仅仅是输出轨迹。对nuScenes和Argoverse 2 Sensor的广泛实验表明,SIPTraj超越了先前的无地图预测器和强大的基于地图的基线,在推理时无需任何HD地图。我们的代码将作为开源发布。
预测/规划 / 4 / 2608.01201

PRISM: Privileged Probabilistic Latent Supervision for End-to-End Autonomous Driving Motion Planning

PRISM:用于端到端自主驾驶运动规划的特权概率潜在监督
Volodymyr Havrylov, Faris Janjoš, Andreas Look, Jürgen Mathes, Andreas Geiger
cs.RO · cs.CV
Abstract
End-to-end autonomous driving (E2E AD) systems integrate perception, prediction, and planning into a single differentiable architecture. While these models show great promise, their standard training often relies on output-only supervision, which can lead to weak gradients for the hidden layers of increasingly complex models. Recent works have integrated vision-language model (VLM) supervision for latent features to address this, yielding substantial empirical gains, yet leaving the underlying theoretical mechanisms poorly understood. Our investigation into this methodology reveals that the resulting performance gains stem not from VLM reasoning capabilities, as previously assumed, but rather from the latent connections forged between the E2E AD model and ground-truth (GT) data during training. Building on this insight, we propose a probabilistic deep supervision framework that regularizes intermediate latent representations directly from GT data. By treating model latents as reparameterizable distributions, we optimize the architecture via the Evidence Lower Bound (ELBO). Our evaluations conducted on the nuScenes dataset demonstrate that supervising trajectory-related latents with future GT paths consistently improves planning performance. Using identical training data and E2E architectures, our method achieves an 8% reduction in planning L2 error and a 3% decrease in collision rates compared to competitive vectorized baselines, all while incurring negligible computational overhead.
Chinese Translation
端到端自主驾驶(E2E AD)系统将感知、预测和规划整合到一个单一的可微分架构中。尽管这些模型展现出巨大的潜力,但其标准训练通常依赖于仅输出监督,这可能导致对于日益复杂模型的隐藏层产生较弱的梯度。近期的研究将视觉-语言模型(VLM)监督整合到潜在特征中,以解决这一问题,取得了显著的实证增益,但其背后的理论机制仍然不够清晰。我们对这一方法的调查表明,性能提升的原因并非源于VLM的推理能力,如之前所假设的,而是源于在训练过程中E2E AD模型与真实数据(GT)之间建立的潜在联系。基于这一见解,我们提出了一种概率深度监督框架,直接从GT数据对中间潜在表示进行正则化。通过将模型潜在视为可重新参数化的分布,我们通过证据下界(ELBO)优化架构。我们在nuScenes数据集上的评估表明,使用未来GT路径对与轨迹相关的潜在进行监督,能够持续改善规划性能。在使用相同训练数据和E2E架构的情况下,我们的方法相比于竞争性的向量化基线实现了8%的规划L2误差降低和3%的碰撞率下降,同时几乎没有增加计算开销。
预测/规划 / 5 / 2608.02289

Extended Field of View Analysis for VideoGAN-based Trajectory Generation

基于VideoGAN的轨迹生成的扩展视野分析
Annajoyce Mariani, Kira Maag, Hanno Gottschalk
cs.CV · cs.LG
Abstract
Realistic and diverse trajectory generation is central to enabling higher levels of vehicle automation. While rule-based and classical learning-based methods may struggle to capture the complexity of traffic behavior, generative models have already demonstrated in other fields that they can handle a comparable level of complexity. In this paper, we build upon previous work on generative adversarial network (GAN)-based semantic bird's-eye-view traffic generation and extend the proposed framework in several key aspects. We improve the semantic representation, replace the trajectory extraction procedure with a graph-based association method, and systematically investigate increasingly larger fields of view. In addition, we introduce a quantitative evaluation framework to assess hallucinations and object permanence in generated videos. Our experiments demonstrate that the framework generalizes to larger and more complex traffic scenes while maintaining statistically realistic trajectories and coherent spatial relationships between traffic participants. Within 150GPU hours of training and with inference times below 20ms for scenes of up to 20s, our results demonstrate that video-based GANs remain an efficient and scalable approach for realistic trajectory generation, even in substantially larger traffic scenes, making them well suited for downstream tasks such as prediction, planning, and simulation in automated driving.
Chinese Translation
现实且多样的轨迹生成是实现更高水平车辆自动化的核心。虽然基于规则和经典学习的方法可能难以捕捉交通行为的复杂性,但生成模型在其他领域已经证明能够处理相当复杂的情况。在本文中,我们在基于生成对抗网络(GAN)的语义鸟瞰视图交通生成的先前工作基础上,扩展了所提出的框架的几个关键方面。我们改进了语义表示,使用基于图的关联方法替换了轨迹提取过程,并系统地研究了越来越大的视野。此外,我们引入了一个定量评估框架,以评估生成视频中的幻觉和物体持久性。我们的实验表明,该框架能够推广到更大和更复杂的交通场景,同时保持统计上现实的轨迹和交通参与者之间一致的空间关系。在150个GPU小时的训练和20秒场景推理时间低于20毫秒的情况下,我们的结果表明,基于视频的GAN仍然是实现现实轨迹生成的高效且可扩展的方法,即使在显著更大的交通场景中,也使其非常适合于自动驾驶中的预测、规划和仿真等下游任务。
预测/规划 / 6 / 2608.00298

WM-Cov: Test Adequacy for Interactive World-Model-Style Autonomous Driving Simulation

WM-Cov:交互式世界模型风格自主驾驶仿真测试充分性
Jianxun Cui, Ping Wu, Stanisa Peric, Marko Milojkovic, Vladan Devedzic
cs.AI
Abstract
World models and generative simulators are emerging as interactive testing infrastructure for autonomous driving because they can react to the ego planner and produce counterfactual, rare, and safety-critical rollouts. This changes a test scenario from a fixed replayed trajectory into an interactive scenario family whose realized evolution depends on the planner under test. The unresolved question is therefore not only whether dangerous rollouts can be generated, but what valid closed-loop evidence is enough to support a specified testing intent and stopping decision. This paper formulates interactive world-model-style testing adequacy and introduces WM-Cov, a provider-agnostic evaluation layer that converts raw provider outputs into requested, realized, and valid evidence. WM-Cov reports adequacy through coverage growth, valid-failure discovery, failure-mode diversity, realism, artifact suppression, duplicate accounting, and valid-evidence precision. Studies on executed TeraSim/SUMO events, WM-like mixed trace pools, and a real DriveArena TrafficManager--WorldDreamer matrix show that dangerous-looking events can include valid ADS failures, duplicates, partial realizations, and artifacts. The DriveArena matrix evaluates two planners, two horizons, six prompt conditions, and 360 ego-route requests; 304 attempts become fully realized evidence and 56 remain partial. A disjoint 80-request route-slice check yields 74 fully realized and 6 partial attempts. The results support evaluating world-model-style testing by convergence of valid interactive evidence under budget, rather than by raw generated failures or prompt coverage alone.
Chinese Translation
世界模型和生成模拟器正逐渐成为自主驾驶的交互测试基础设施,因为它们能够对自我规划者做出反应,并生成反事实、稀有和安全关键的演示。这将测试场景从固定的重放轨迹转变为一个交互场景家族,其实现的演变依赖于被测试的规划者。因此,未解决的问题不仅是是否可以生成危险的演示,而是支持特定测试意图和停止决策所需的有效闭环证据的量。本文提出了交互式世界模型风格测试的充分性,并引入了WM-Cov,一个与提供者无关的评估层,将原始提供者输出转换为请求的、实现的和有效的证据。WM-Cov通过覆盖增长、有效失败发现、失败模式多样性、现实性、伪影抑制、重复计数和有效证据精度来报告充分性。对执行的TeraSim/SUMO事件、WM类混合轨迹池以及真实的DriveArena TrafficManager-WorldDreamer矩阵的研究表明,外观危险的事件可能包括有效的ADS失败、重复、部分实现和伪影。DriveArena矩阵评估了两个规划者、两个时间范围、六个提示条件和360个自我路线请求;304次尝试变成完全实现的证据,56次保持部分实现。一个不相交的80请求路线片段检查产生了74次完全实现和6次部分尝试。结果支持通过在预算下有效交互证据的收敛来评估世界模型风格的测试,而不是仅通过原始生成的失败或提示覆盖。
预测/规划 / 7 / 2608.01755

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs

未来轨迹的延迟暴露用于自主驾驶可验证推理的视觉语言模型
Zixuan Huang, Yang Zhou, Kaixuan Wang, Guli Zhang, Hongyan Xie, Yakun Zhu, Hao Geng, Xiaozhi Chen, Yikun Ban, Deqing Wang
cs.AI
Abstract
Recent Vision-Language-Action (VLA) models for autonomous driving (AD) increasingly utilize chain-of-thought (CoT) supervision to enhance the reasoning capabilities of their Vision-Language Model (VLM) components, yet existing annotation pipelines commonly expose the teacher model to the logged ground-truth (GT) future trajectory. We empirically show that this induces trajectory anchoring bias: teacher models rationalize the revealed outcome rather than infer a decision from scene evidence, producing less causally faithful CoTs and substantially more severe hallucinations, especially in causally challenging scenes. Removing the GT trajectory eliminates this shortcut, but open-ended trajectory generation entangles high-level decision-making with precise geometric synthesis and low-level dynamics. To make trajectory-level driving decisions verifiable without requiring open-ended trajectory synthesis, we introduce Autonomous-Driving Multiple-Choice Question (AD-MCQ), which casts planning as selection among explicit trajectory candidates. Taking this a step further, we propose Deferred Exposure of Future Trajectories for RLVR (DEFT-RLVR) to transform future trajectories from pre-decision anchors into post-decision verification targets. Experimental results show that DEFT-RLVR improves AD reasoning while preserving or even enhancing general visual capabilities. With VLM-only inference and controllable difficulty through candidate construction, AD-MCQ provides a flexible, scalable, and extensible foundation for future research on verifiable AD reasoning.
Chinese Translation
近期的自主驾驶(AD)视觉-语言-动作(VLA)模型越来越多地利用思维链(CoT)监督来增强其视觉-语言模型(VLM)组件的推理能力,但现有的注释流程通常将教师模型暴露于记录的真实(GT)未来轨迹中。我们实证表明,这会导致轨迹锚定偏差:教师模型会合理化所揭示的结果,而不是从场景证据中推断决策,从而产生因果关系不够忠实的CoT,并在因果关系具有挑战性的场景中产生更严重的幻觉。去除GT轨迹消除了这一捷径,但开放式轨迹生成将高层决策与精确几何合成和低层动态纠缠在一起。为了使轨迹级驾驶决策可验证而不需要开放式轨迹合成,我们引入了自主驾驶多项选择题(AD-MCQ),将规划视为在明确轨迹候选中进行选择。更进一步,我们提出了未来轨迹的延迟暴露用于RLVR(DEFT-RLVR),将未来轨迹从决策前的锚点转变为决策后的验证目标。实验结果表明,DEFT-RLVR在提高AD推理的同时保持或甚至增强了一般视觉能力。通过仅使用VLM推理和通过候选构建可控的难度,AD-MCQ为未来可验证AD推理的研究提供了灵活、可扩展和可扩展的基础。

VLM/VLA

1
VLM/VLA / 1 / 2608.01035

WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA

WAM-Diff2:用于高效自主驾驶的层次化自回归到扩散蒸馏
Zhihao Zhu, Hanlin Shang, Mingwang Xu, Feipeng Cai, Zhuolin He, Yaoyi Li, Jianhua Han, Hang Xu, Siyu Zhu
cs.RO · cs.AI · cs.CV
Abstract
Vision-Language-Action (VLA) models have emerged as a prominent paradigm for end-to-end autonomous driving; however, their efficient deployment is severely constrained by high computational latency and exposure bias arising from sequential autoregressive decoding. Conversely, while specialized diffusion policies enable low-latency, parallel execution, training them from scratch typically yields narrow, single-task architectures that lack holistic visual-linguistic reasoning. Successfully transforming pre-trained autoregressive generalists into parallel diffusion models could combine multi-task cognitive intelligence with execution efficiency, yet this transition presents a formidable architectural challenge due to mismatched attention patterns (causal versus bidirectional) and divergent optimization objectives. To bridge this divide, we introduce WAM-Diff2, a multi-task discrete diffusion VLA framework powered by a three-stage hierarchical distillation strategy. By structuring the architectural shift through progressive block-wise adaptation, block-wise distillation, and model-wise cross-scale distillation, WAM-Diff2 preserves the underlying semantic foundations of the base model while accelerating inference. Extensive evaluations across driving understanding, perception, and planning benchmarks demonstrate that WAM-Diff2 effectively mitigates exposure bias and achieves performance parity with autoregressive baselines. Crucially, the autoregressive-to-diffusion transition yields a 2.8x decoding speedup, which scales to an ultimate 15.1x acceleration when combined with system-level optimizations including FlashInfer and CUDA Graphs.
Chinese Translation
视觉-语言-行动(VLA)模型已成为端到端自主驾驶的一个重要范式;然而,由于顺序自回归解码所导致的高计算延迟和暴露偏差,其高效部署受到严重限制。相反,尽管专门的扩散策略能够实现低延迟的并行执行,但从零开始训练它们通常会产生狭窄的单任务架构,缺乏整体的视觉-语言推理。成功地将预训练的自回归通用模型转变为并行扩散模型,能够将多任务认知智能与执行效率结合起来,但这一转变由于注意力模式(因果与双向)不匹配和优化目标的差异,带来了巨大的架构挑战。为了解决这一问题,我们提出了WAM-Diff2,一个由三阶段层次化蒸馏策略驱动的多任务离散扩散VLA框架。通过逐步的块级适应、块级蒸馏和模型级跨尺度蒸馏来构建架构转变,WAM-Diff2在加速推理的同时保留了基础模型的语义基础。在驾驶理解、感知和规划基准上的广泛评估表明,WAM-Diff2有效减轻了暴露偏差,并与自回归基线达成了性能平衡。关键的是,自回归到扩散的转变实现了2.8倍的解码加速,当与包括FlashInfer和CUDA Graphs在内的系统级优化结合时,最终可实现15.1倍的加速。

仿真/数据

2
仿真/数据 / 1 / 2608.01336

Asleep at the Wheel: JEPA's Limitations in Evaluating Novel Driving Data

驾驶中的失误:JEPA在评估新颖驾驶数据中的局限性
Advait Pavuluri, Shamik Karkhanis, Uzma Mushtaque
cs.CV · cs.LG
Abstract
Modern autonomous-driving fleets record far more video than human reviewers can inspect. This motivates the need for an automatic clip triage mechanism, to surface rare and review-worthy clips, so that driving models can be fine-tuned to better handle unideal circumstances. We test a label-free approach that scores clips by the prediction-error "novelty" of a self-supervised joint-embedding predictive architecture (JEPA); a frozen V-JEPA video encoder is paired with a lightweight predictor head to reconstruct masked clip embeddings, and clips whose embeddings are hard to predict are flagged as interesting. Evaluated under a realistic protocol that trains on one dataset and tests against footage from others, this approach appears highly effective. We show that this apparent success is actually a domain-shift consequence: on a fair benchmark drawn from a single dataset, this mechanism collapses to chance and is on par with simple no-training baselines. A lightly supervised probe on the same frozen embeddings results in almost double the average precision, indicating that the bottleneck is indeed the self-supervised objective, rather than the representation. We present this as a study for evaluating the effectiveness of self-supervised learning, where cross-dataset protocols can silently reward domain separation over novelty.
Chinese Translation
现代自动驾驶车队记录的视频量远超人类评审员能够检查的数量。这促使我们需要一种自动剪辑筛选机制,以便提取稀有且值得审查的剪辑,从而使驾驶模型能够更好地处理不理想的情况。我们测试了一种无标签的方法,通过自监督的联合嵌入预测架构(JEPA)对剪辑进行预测误差“新颖性”评分;一个冻结的 V-JEPA 视频编码器与一个轻量级预测头配对,以重建被遮蔽的剪辑嵌入,而那些嵌入难以预测的剪辑则被标记为有趣。在一个现实的协议下进行评估,该协议在一个数据集上训练并在其他数据集的录像上测试,这种方法似乎非常有效。我们展示了这种表面上的成功实际上是领域转移的结果:在一个来自单一数据集的公平基准上,这一机制崩溃至随机水平,并与简单的无训练基线相当。对相同冻结嵌入进行轻度监督的探测结果显示平均精度几乎翻倍,表明瓶颈确实在于自监督目标,而非表示本身。我们将此作为评估自监督学习有效性的研究,其中跨数据集协议可能在不知不觉中奖励领域分离而非新颖性。
仿真/数据 / 2 / 2608.01572

Enhancing Visual Perception in Foggy Conditions via Multiclass Fog Density Modeling

通过多类雾密度建模增强雾天条件下的视觉感知
Mohamad Mofeed Chaar, Galia Weidl
cs.CV · cs.AI
Abstract
Autonomous driving (AD) systems have advanced rapidly over the past decade; however, robust perception under adverse weather conditions remains a major challenge, particularly in dense fog. In this work, we investigate fog-aware perception using synthetically generated fog data derived from the Waymo dataset. To support fog simulation, depth images are generated using an iterative learning approach. We consider five fog-density levels: clear, light fog, moderate fog, heavy fog, and very heavy fog. Instead of training a single unified model across all conditions, we train separate perception models for each fog-density level. Experimental results show that density-specific training improves performance in severe fog conditions. In particular, for the very heavy fog class, recall improves from 0.076 to 0.232, corresponding to an absolute gain of 15.6 percentage points. These findings suggest that deploying multiple specialized models, rather than a single general-purpose model, can improve perception robustness for autonomous vehicles under challenging visibility conditions. Future work will extend this strategy to additional sensing modalities, including LiDAR and radar, and evaluate generalization across diverse weather scenarios.
Chinese Translation
自主驾驶(AD)系统在过去十年中迅速发展;然而,在恶劣天气条件下保持稳健的感知仍然是一个主要挑战,尤其是在浓雾环境中。在本研究中,我们利用从Waymo数据集中生成的合成雾数据来研究雾感知。为了支持雾模拟,采用迭代学习方法生成深度图像。我们考虑了五种雾密度水平:清晰、轻雾、中等雾、重雾和极重雾。我们不是在所有条件下训练一个统一的模型,而是为每种雾密度水平训练单独的感知模型。实验结果表明,针对特定密度的训练在严重雾天条件下提高了性能。特别是在极重雾类别中,召回率从0.076提高到0.232,绝对增益为15.6个百分点。这些发现表明,部署多个专用模型而不是单一通用模型,可以提高自主车辆在挑战性能见度条件下的感知鲁棒性。未来的工作将把这一策略扩展到包括LiDAR和雷达在内的其他传感模式,并评估在不同天气场景下的泛化能力。