← Back to Index
Daily Autonomous Driving Research Digest

AutoDrive Papers

2026-07-16
10
Papers
7
Topics
10
Translated

感知

1
感知 / 1 / 2607.13927

Cyclone: Diffusion Model for Cycle-Consistent Weather Editing from Unpaired Driving Data

Cyclone:基于无配对驾驶数据的循环一致天气编辑扩散模型
Thang-Anh-Quan Nguyen, Moussab Bennehar, Luis Guillermo Roldao Jimenez, Nathan Piasco, Dzmitry Tsishkou, Laurent Caraffa, Jean-Philippe Tarel, Roland Brémond
cs.CV
Abstract
Reliable perception under diverse weather conditions remains a major challenge for autonomous driving systems. A common strategy to improve robustness is either to synthesize adverse weather conditions for training perception models or to apply weather-removal techniques to recover clean inputs. However, existing approaches typically rely on synthetic data augmentation or physics-based, task-specific models that require paired training data and often struggle to generate realistic weather effects or generalize robustly to out-of-domain scenarios. Toward this problem, we present Cyclone, a unified framework for weather editing based on latent diffusion, equipped with cycle-consistent constraints and knowledge from image-text models. Cyclone enables the generation of multiple weather conditions across diverse scenes while eliminating the need for paired data. Experimental results show that our approach produces more realistic, structure-preserving outputs than existing baselines and leads to consistent improvements across several downstream driving perception tasks. Furthermore, we demonstrate that Cyclone can be distilled to a video diffusion model for temporally consistent weather editing.
Chinese Translation
在多样化天气条件下可靠感知仍然是自动驾驶系统面临的主要挑战。提高鲁棒性的常见策略是合成恶劣天气条件以训练感知模型,或应用天气去除技术以恢复干净输入。然而,现有方法通常依赖于合成数据增强或基于物理的、任务特定的模型,这些模型需要配对的训练数据,并且往往难以生成真实的天气效果或在域外场景中稳健地泛化。针对这一问题,我们提出了Cyclone,一个基于潜在扩散的统一天气编辑框架,配备循环一致性约束和来自图像-文本模型的知识。Cyclone能够在多样场景中生成多种天气条件,同时消除了对配对数据的需求。实验结果表明,我们的方法生成的输出比现有基线更具真实感且保留结构,并在多个下游驾驶感知任务中实现了一致的改进。此外,我们展示了Cyclone可以被提炼为一个视频扩散模型,以实现时间一致的天气编辑。

BEV/Occupancy

1
BEV/Occupancy / 1 / 2607.13410

Ego-Dynamics-Augmented World Model for Autonomous Driving with Zero-Shot Cross-Chassis Adaptation

增强自我动态的世界模型用于零样本跨底盘适应的自主驾驶
Zhidong Wang, Jingsong Liang, Zirui Li, Zhan Chen, Han Yu, Chen Lv
cs.RO
Abstract
World model (WM)-based reinforcement learning enables sample-efficient end-to-end autonomous driving learning by imagining long-horizon trajectories in latent space. However, most driving WMs operate on bird's-eye-view (BEV) representations that are inherently egocentric: the transition between consecutive frames entangles the ego vehicle's own motion with scene dynamics. As a result, the WM devotes significant capacity to recovering ego-motion from warped observations, at the cost of scene modeling fidelity and imagination accuracy. This work proposes DynaDreamer, a dynamics-augmented Dreamer-style reinforcement learning method to address this problem by augmenting the WM with an explicit ego-dynamics prior. A physics-informed ego-dynamics encoder-decoder extracts the ego-state history into a compact and identifiable context, which modulates a causal Transformer WM to condition both its prior and posterior latents. During imagination, the ego-dynamics predictor propagates this context forward to keep the ego-dynamics prior synchronized with the rollout. An information-theoretic analysis shows that conditioning on this context reduces both the predictive entropy of the observation transition and the prior--posterior Kullback--Leibler divergence, confining the WM's modeling burden to the scene dynamics beyond ego-motion. An additional benefit is zero-shot cross-chassis adaptation: the ego-dynamics context depends on identifiable chassis parameters, so that a vehicle with previously unseen dynamic characteristics can adapt the WM to the new chassis without retraining. Experiments demonstrate that DynaDreamer improves task success rates over the strongest baseline by 28% and 61% in urban and highway driving scenarios, respectively, with the advantage rising to 73% when extrapolating to unseen chassis.
Chinese Translation
基于世界模型(WM)的强化学习通过在潜在空间中想象长时间轨迹,实现了样本高效的端到端自主驾驶学习。然而,大多数驾驶WM操作于鸟瞰图(BEV)表示,这种表示本质上是自我中心的:连续帧之间的过渡将自我车辆的运动与场景动态纠缠在一起。因此,WM在从扭曲的观测中恢复自我运动方面投入了大量能力,牺牲了场景建模的保真度和想象的准确性。本研究提出了DynaDreamer,一种增强动态的Dreamer风格强化学习方法,通过用显式的自我动态先验增强WM来解决这一问题。一个物理信息驱动的自我动态编码器-解码器将自我状态历史提取为紧凑且可识别的上下文,调节因果Transformer WM以条件其先验和后验潜变量。在想象过程中,自我动态预测器向前传播该上下文,以保持自我动态先验与展开同步。信息论分析表明,基于该上下文的条件减少了观测过渡的预测熵和先验-后验Kullback-Leibler散度,将WM的建模负担限制在超出自我运动的场景动态上。另一个好处是零样本跨底盘适应:自我动态上下文依赖于可识别的底盘参数,因此具有先前未见动态特征的车辆可以在不重新训练的情况下将WM适应于新底盘。实验表明,DynaDreamer在城市和高速公路驾驶场景中,任务成功率分别比最强基线提高了28%和61%,当推广到未见底盘时,优势上升至73%。

LiDAR/点云

1
LiDAR/点云 / 1 / 2607.13405

WNOJ-LIO: A White-Noise-on-Jerk Motion-Prior EKF for High-Dynamic LiDAR-IMU Fusion

WNOJ-LIO:一种基于白噪声-加速度变化运动先验的高动态激光雷达-惯性测量单元融合扩展卡尔曼滤波器
Junning Lyu, Qizhi Guo, Xia Ning, Tao Song, Shaoming He
cs.RO
Abstract
LiDAR-inertial odometry (LIO) is a key component of autonomous navigation, but high-dynamic driving exposes two coupled challenges: intra-scan motion distortion and vibration-contaminated inertial measurements. Most real-time LiDAR-inertial pipelines propagate the system state by integrating raw IMU measurements and then use the propagated trajectory for point cloud de-distortion, thereby propagating inertial noise into both the corrected scan and the subsequent scan-to-map registration. This paper presents WNOJ-LIO, a LiDAR-IMU fusion framework based on a White-Noise-on-Jerk (WNOJ) Extended Kalman Filter (EKF). WNOJ-LIO employs a decoupled WNOJ prior on $\R^3 \times \SO(3)$ for state prediction and treats the IMU as a high-frequency measurement source rather than the driver of state propagation. The resulting posterior state history is then used for LiDAR scan de-distortion and subsequent point-to-plane LiDAR updates. The decoupled process model enables closed-form covariance propagation, thereby bridging the gap between batch WNOJ Gaussian process (GP) trajectory priors and recursive filtering. Simulation results demonstrate improvements in acceleration and angular-velocity denoising, scan de-distortion, and localization accuracy over a FAST-LIO-style baseline. Real-world experiments were conducted using an autonomous racing car on four driving segments with maximum speeds ranging from 53 to 208~km/h, covering a wide range of vehicle vibration levels. The experiments further validate the proposed method and provide a comprehensive evaluation of its performance in estimating acceleration, angular velocity, body-frame linear velocity, attitude, and position under highly dynamic driving. The source code of WNOJ-LIO is publicly available at https://github.com/LvJohny/wnoj-ekf-lio.git.
Chinese Translation
激光雷达惯性里程计(LIO)是自主导航的关键组成部分,但高动态驾驶带来了两个相互关联的挑战:扫描内运动失真和振动污染的惯性测量。大多数实时激光雷达-惯性管道通过整合原始惯性测量单元(IMU)测量值来传播系统状态,然后使用传播的轨迹进行点云去失真,从而将惯性噪声传播到校正后的扫描和后续的扫描与地图配准中。本文提出了WNOJ-LIO,一种基于白噪声-加速度变化(WNOJ)扩展卡尔曼滤波器(EKF)的激光雷达-IMU融合框架。WNOJ-LIO在$ ext{R}^3 imes ext{SO}(3)$上采用解耦的WNOJ先验进行状态预测,并将IMU视为高频测量源,而非状态传播的驱动因素。由此得到的后验状态历史用于激光雷达扫描去失真和后续的点对平面激光雷达更新。解耦的过程模型实现了闭式形式的协方差传播,从而弥合了批处理WNOJ高斯过程(GP)轨迹先验与递归滤波之间的差距。仿真结果表明,与FAST-LIO风格基线相比,在加速度和角速度去噪、扫描去失真和定位精度方面都有所改善。通过使用一辆自主赛车在四个驾驶段进行的实地实验,最大速度范围从53到208 km/h,覆盖了广泛的车辆振动水平。这些实验进一步验证了所提出的方法,并全面评估了其在高动态驾驶下估计加速度、角速度、机体坐标系线速度、姿态和位置的性能。WNOJ-LIO的源代码已公开,地址为https://github.com/LvJohny/wnoj-ekf-lio.git。

深度/几何

1
深度/几何 / 1 / 2607.13481

GPOcc++: Unified Sparse Gaussian Occupancy Prediction with Visual Geometry Priors

GPOcc++:结合视觉几何先验的统一稀疏高斯占用预测
Changqing Zhou, Yueru Luo, Yulan Guo, Bing Wang, Jie Qin, Changhao Chen
cs.CV
Abstract
Accurate 3D scene understanding is fundamental to embodied intelligence and autonomous driving, where 3D occupancy provides a unified representation of objects, structures, and free space. However, recovering such a complete volumetric representation from visual observations remains challenging, particularly in occluded and unobserved regions. Visual geometry priors offer strong and generalizable geometric cues for addressing this challenge, but their outputs are inherently surface-centric, whereas occupancy prediction requires reasoning about volumetric interiors and free space. To bridge this gap, we introduce GPOcc, which transforms visual geometry priors into occupancy-aware sparse Gaussian representations for efficient and expressive volumetric scene modeling. Building on GPOcc, GPOcc++ models multi-view observations and temporal sequences within a unified framework, allowing spatial and temporal evidence to be handled through the same representation. We further extend GPOcc++ from indoor scenes to outdoor occupancy prediction. Extensive experiments on both indoor and outdoor benchmarks demonstrate consistently strong performance across both multi-view and temporal settings, together with favorable efficiency and generalization. Code will be released at https://github.com/JuIvyy/GPOcc.
Chinese Translation
准确的三维场景理解是具身智能和自动驾驶的基础,其中三维占用提供了对象、结构和自由空间的统一表示。然而,从视觉观测中恢复这样一个完整的体积表示仍然具有挑战性,特别是在被遮挡和未观察到的区域。视觉几何先验为解决这一挑战提供了强大且可推广的几何线索,但它们的输出本质上是以表面为中心的,而占用预测需要对体积内部和自由空间进行推理。为了解决这一问题,我们引入了GPOcc,它将视觉几何先验转化为占用感知的稀疏高斯表示,以实现高效且富有表现力的体积场景建模。在GPOcc的基础上,GPOcc++在统一框架内建模多视角观测和时间序列,允许通过相同的表示处理空间和时间证据。我们进一步将GPOcc++从室内场景扩展到室外占用预测。在室内和室外基准上的大量实验表明,在多视角和时间序列设置中均表现出一致的强大性能,同时具有良好的效率和泛化能力。代码将发布在 https://github.com/JuIvyy/GPOcc。

预测/规划

4
预测/规划 / 1 / 2607.13348

Safe Overtaking for Autonomous Racing Using Hierarchical Optimization and Learning-Based Control

基于层次优化和学习控制的自主赛车安全超车
Hassan Jardali, Kai Yin, Lantao Liu
cs.RO
Abstract
Autonomous racing overtaking requires balancing competitive performance with safety under nonlinear vehicle dynamics and real-time constraints. Model Predictive Control (MPC) combined with Control Barrier Functions (CBFs) provides a principled mechanism for certifying forward invariance of a safe set. However, commonly used fixed-decay discrete-time CBF formulations can become overly conservative in interactive racing scenarios, limiting overtaking performance and requiring manual tuning across track conditions. This paper proposes a hierarchical overtaking framework that explicitly separates maneuver-level decision making from safety-certified trajectory control, reducing conservatism while preserving safety. A high-level Mixed-Integer Quadratic Program (MIQP) resolves the combinatorial passing-side selection problem by selecting a feasible overtaking topology, while a nonlinear Frenet-frame MPC enforces vehicle dynamics and safety through embedded discrete-time CBF constraints. This decomposition isolates the combinatorial complexity of maneuver selection from the continuous trajectory optimization. To further mitigate the sensitivity of fixed-decay barrier constraints, a reinforcement learning policy adapts the discrete-time CBF decay parameter online, enabling context-dependent modulation of safety margins without directly controlling vehicle inputs. Simulation and scaled-hardware experiments show that no single fixed decay parameter achieves uniformly strong performance across tracks, whereas the adaptive strategy attains the highest aggregate success rate and consistently strong safety--performance trade-offs without per-track tuning, improving robustness to environment variation while maintaining safety constraint satisfaction in nominal operation.
Chinese Translation
自主赛车超车需要在非线性车辆动力学和实时约束下平衡竞争性能与安全性。模型预测控制(Model Predictive Control, MPC)结合控制屏障函数(Control Barrier Functions, CBFs)为验证安全集的前向不变性提供了一种原则性机制。然而,常用的固定衰减离散时间CBF公式在交互式赛车场景中可能变得过于保守,从而限制超车性能,并需要在不同赛道条件下进行手动调节。本文提出了一种层次化超车框架,明确将机动级决策与安全认证轨迹控制分开,减少保守性同时保持安全性。高层次的混合整数二次规划(Mixed-Integer Quadratic Program, MIQP)通过选择可行的超车拓扑来解决组合性超车侧选择问题,而非线性Frenet框架MPC则通过嵌入的离散时间CBF约束来强制执行车辆动力学和安全性。这种分解将机动选择的组合复杂性与连续轨迹优化隔离开来。为了进一步减轻固定衰减屏障约束的敏感性,强化学习策略在线调整离散时间CBF衰减参数,实现安全边际的上下文依赖调节,而无需直接控制车辆输入。仿真和缩放硬件实验表明,没有单一的固定衰减参数能够在各赛道上实现均匀强劲的性能,而自适应策略则在不需要逐赛道调节的情况下,获得了最高的整体成功率,并持续实现强大的安全性与性能权衡,提高了对环境变化的鲁棒性,同时在名义操作中保持了安全约束的满足。
预测/规划 / 2 / 2607.13354

A Hybrid Sampling-Based Trajectory Planner with Game-Theoretic Guidance for Autonomous Racing

一种基于混合采样的轨迹规划器,结合博弈论指导用于自主赛车
Alexander Langmann, Frederico Pita de Araujo, Mattia Piccinini, Johannes Betz
cs.RO
Abstract
Autonomous racing demands planning algorithms that balance vehicle dynamics at the limits of handling with strategic decision-making in competitive multi-agent scenarios. Game theory provides a mathematical framework for modeling these interactions, enabling interactive trajectory planning and strategic behaviors, such as blocking. However, directly solving full dynamic games online is computationally prohibitive and challenging to integrate into robust, high-frequency autonomous software stacks. This paper proposes a hybrid architecture that integrates game-theoretic reasoning into a sampling-based motion planner, combining strategic interactions with robust trajectory generation. Building upon an $\alpha$-potential game formulation, we utilize an offline-learned potential function to capture multi-agent interactions. During online operation, a gradient-based optimization dynamically refines interaction parameters to generate an \textit{Interaction Reference Path}. This path serves as a dynamic cost bias within a high-frequency sampling planner. We evaluate our approach in a high-fidelity simulation environment on the Yas Marina Circuit. Qualitative and quantitative results demonstrate that our approach successfully induces defensive behaviors like blocking without carrying the computational burden of full dynamic game solvers.
Chinese Translation
自主赛车要求规划算法在车辆动态与竞争多智能体场景中的战略决策之间取得平衡。博弈论提供了一种数学框架,用于建模这些交互,支持交互式轨迹规划和战略行为,如阻挡。然而,在线直接求解完整动态博弈在计算上是不可行的,并且难以集成到稳健的高频自主软件堆栈中。本文提出了一种混合架构,将博弈论推理集成到基于采样的运动规划器中,结合战略交互与稳健的轨迹生成。在$eta$-潜力博弈的基础上,我们利用离线学习的潜力函数来捕捉多智能体交互。在在线操作期间,基于梯度的优化动态地细化交互参数,以生成 extit{交互参考路径}。该路径作为高频采样规划器中的动态成本偏置。我们在雅斯玛丽娜赛道的高保真仿真环境中评估了我们的方法。定性和定量结果表明,我们的方法成功地诱导了如阻挡等防御行为,而无需承担完整动态博弈求解器的计算负担。
预测/规划 / 3 / 2607.13704

nuTruck: Benchmarking Autonomous Driving Planning for Distributed Electric-drive Trucks

nuTruck:分布式电驱动卡车自主驾驶规划的基准测试
Jinyu Miao, Pu Zhang, Yifei He, Chengyao Zhang, Kun Jiang, Ke Wang, Mengmeng Yang, Diange Yang
cs.RO
Abstract
The dominance of traditional rule-based methods in autonomous driving has gradually been replaced by learning-based approaches. While learning-based planners have achieved considerable success in passenger vehicles, their performance on heavy-duty trucks, particularly modern distributed electric-drive trucks (DETs), remains largely unexplored. To facilitate research and application of learning-based planners in DETs, this letter presents the first high-fidelity benchmark, called nuTruck, designed to support large-scale neural network training and closed-loop evaluation. Given the complex dynamics and high rollover susceptibility of DETs, we first incorporate a highly accurate nonlinear truck dynamical model into the simulation, which enables independent driving and steering of all wheels and captures dynamic load transfer caused by acceleration, deceleration, and cornering, thereby allowing quantitative assessment of rollover risk in closed-loop simulation. Second, we adapt several rule-based and learning-based planners as baselines for DETs and evaluate their performance in closed-loop simulation. Finally, using real-world driving scenarios from the nuPlan dataset, we conduct extensive closed-loop evaluations, analyzing not only conventional collision-free planning performance, but also the dynamical safety of the planned trajectories. The proposed nuTruck benchmark is expected to serve as a new standard for fair and realistic evaluation of autonomous driving planners on DETs.
Chinese Translation
传统基于规则的方法在自主驾驶领域的主导地位逐渐被基于学习的方法所取代。尽管基于学习的规划器在乘用车上取得了显著成功,但它们在重型卡车,特别是现代分布式电驱动卡车(DETs)上的表现仍然很大程度上未被探索。为了促进基于学习的规划器在DETs中的研究和应用,本文提出了第一个高保真基准,称为nuTruck,旨在支持大规模神经网络训练和闭环评估。考虑到DETs的复杂动态特性和高翻车敏感性,我们首先将一个高度准确的非线性卡车动力学模型纳入模拟中,该模型能够实现所有车轮的独立驱动和转向,并捕捉因加速、减速和转弯引起的动态载荷转移,从而允许在闭环模拟中对翻车风险进行定量评估。其次,我们将几种基于规则和基于学习的规划器调整为DETs的基准,并评估它们在闭环模拟中的表现。最后,利用来自nuPlan数据集的真实驾驶场景,我们进行了广泛的闭环评估,不仅分析了传统的无碰撞规划性能,还分析了规划轨迹的动态安全性。所提出的nuTruck基准预计将作为对DETs上的自主驾驶规划器进行公平和现实评估的新标准。
预测/规划 / 4 / 2607.13926

S-squared-VLA: Decoupling Semantic and Spatial Streams in Vision-Language-Action Models for Autonomous Driving

S平方-VLA:在自动驾驶的视觉-语言-动作模型中解耦语义和空间流
Jianguo Yu, Rukang Wang, Duanfeng Chu, Chen Wang, Renju Feng, Liping Lu
cs.RO
Abstract
Vision-Language Models (VLMs) have demonstrated remarkable potential for high-level reasoning in autonomous driving, yet they fundamentally struggle to generate precise, low-level control actions. This limitation is rooted in a semantic-physical gap caused by the inherent mismatch between discrete language tokens and continuous trajectory planning. While Vision-Language-Action (VLA) architectures attempt to bridge this gap by unifying perception and control into a single policy, this entanglement creates a new bottleneck. Standard VLAs experience a severe spatial representation collapse, which irreversibly degrades the fine-grained spatial and geometric priors essential for safe, boundary-aware navigation. To address this limitation, we propose the S-squared-VLA, which explicitly decouples the semantic and spatial streams in Vision-Language-Action models. The semantic stream leverages hierarchical bridging to extract multi-scale VLM features for robust intent reasoning. In parallel, an independent spatial stream bypasses the autoregressive language bottleneck, directly preserving uncompressed spatial features from the visual encoder. By integrating auxiliary perception supervision, this stream explicitly equips the model with rich spatial and geometric priors. Finally, a dual-stream planning adapter fuses high-level semantic intent with precise spatial constraints via cascaded attention mechanisms. Evaluations on the NAVSIM closed-loop benchmark show that S-squared-VLA achieves a Predictive Driver Model Score (PDMS) of 87.1, establishing a new state-of-the-art for VLA models under a purely supervised fine-tuning (SFT) setting. By mitigating the spatial representation collapse of traditional VLMs, our framework significantly outperforms baselines, achieving the highest No Collision (NC) rate of 98.4 among all evaluated methods.
Chinese Translation
视觉-语言模型(VLMs)在自动驾驶中的高层次推理方面展现了显著的潜力,但它们在生成精确的低层次控制动作方面存在根本性的困难。这一限制源于语义与物理之间的差距,这种差距是由于离散语言符号与连续轨迹规划之间的固有不匹配造成的。尽管视觉-语言-动作(VLA)架构试图通过将感知与控制统一为单一策略来弥合这一差距,但这种纠缠又创造了新的瓶颈。标准的VLA经历了严重的空间表示崩溃,这不可逆转地降低了安全且边界感知导航所必需的细粒度空间和几何先验。为了解决这一限制,我们提出了S平方-VLA,它明确地在视觉-语言-动作模型中解耦语义和空间流。语义流利用层次桥接提取多尺度VLM特征,以实现稳健的意图推理。与此同时,一个独立的空间流绕过自回归语言瓶颈,直接保留来自视觉编码器的未压缩空间特征。通过整合辅助感知监督,该流明确地为模型提供丰富的空间和几何先验。最后,一个双流规划适配器通过级联注意机制将高层次语义意图与精确的空间约束融合在一起。在NAVSIM闭环基准测试中的评估显示,S平方-VLA在纯监督微调(SFT)设置下达到了87.1的预测驾驶员模型分数(PDMS),为VLA模型建立了新的最先进水平。通过缓解传统VLM的空间表示崩溃,我们的框架显著优于基线,在所有评估方法中实现了98.4的最高无碰撞(NC)率。

仿真/数据

1
仿真/数据 / 1 / 2607.13674

WAVE-Stereo: Warp-Aligned Volume Encoding for Stereo Matching

WAVE-Stereo:用于立体匹配的扭曲对齐体积编码
Zehan Liu, Yage He, Xianwu Gong
cs.CV · cs.RO
Abstract
Existing iterative stereo matching methods primarily adopt two types of correspondence representation: explicit matching search via correlation volumes and local residual refinement via warped features, yet the two remain separately modeled. We propose WAVE-Stereo, built on a core insight: correlation volumes and feature warping provide complementary matching cues. \textbf{GeoWarp Correspondence Encoder (GWCE)} encodes matching search, residual alignment, and disparity prior in parallel at the ConvGRU input. To mitigate matching degradation in textureless regions, we propose \textbf{Periodic Global Context Propagation (PGCP)}, which propagates global spatial information in a periodic manner. On five real-world benchmarks -- Middlebury, ETH3D, KITTI 2012, KITTI 2015, and Booster -- WAVE-Stereo achieves competitive zero-shot generalization accuracy without any external foundation model prior, achieving 3.18\% D1-all on KITTI 2015, 4.42\% Bad-2.0 on Booster, and 66ms real-time inference, striking a favorable balance between accuracy and efficiency. Our code is available at https://github.com/yamanoko-do/WAVE-Stereo.
Chinese Translation
现有的迭代立体匹配方法主要采用两种对应关系表示:通过相关体积进行显式匹配搜索,以及通过扭曲特征进行局部残差精细化,但这两者仍然是分开建模的。我们提出了WAVE-Stereo,基于一个核心见解:相关体积和特征扭曲提供了互补的匹配线索。\textbf{GeoWarp Correspondence Encoder (GWCE)}在ConvGRU输入中并行编码匹配搜索、残差对齐和视差先验。为了减轻在无纹理区域的匹配退化,我们提出了\textbf{Periodic Global Context Propagation (PGCP)},它以周期性方式传播全局空间信息。在五个真实世界基准测试上——Middlebury、ETH3D、KITTI 2012、KITTI 2015和Booster——WAVE-Stereo在没有任何外部基础模型先验的情况下,实现了具有竞争力的零-shot泛化精度,在KITTI 2015上达到了3.18\% D1-all,在Booster上达到了4.42\% Bad-2.0,并实现了66毫秒的实时推理,在准确性和效率之间取得了良好的平衡。我们的代码可在https://github.com/yamanoko-do/WAVE-Stereo获取。

协同/V2X

1
协同/V2X / 1 / 2607.13806

A Deployed Hybrid Vehicle-in-the-Loop Platform for Validating Cooperative Perception

一种部署的混合车辆环路平台用于验证协同感知
Anastasia Bolovinou, Giorgos Hadjipavlis, Markos Antonopoulos, Panagiotis Tachtalis, Konstantinos Petousakis, Konstantinos Lazaridis, Alexandros Siskos, Bill Roungas, Angelos Amditis
cs.RO · cs.MA
Abstract
European safety regulation now permits a large share of automated-driving homologation evidence to be produced virtually, provided a validated physical-virtual facility generates it. We present a deployed hybrid Vehicle-in-the-Loop (ViL) platform that couples a real instrumented vehicle with a CARLA-based digital twin (DT) through a V2X message pipeline, and we report its first integrated operation on a public-road-representative test track. A real vehicle streams ETSI-compliant CAM/CPM messages into the DT, where a GPU-accelerated Cooperative Perception (CP) module fuses them into a probabilistic occupancy grid during scenario runtime. We demonstrate the platform on a multi-vehicle double T-intersection scenario, characterise the CP workload across nominal, rain and night conditions and five localization-noise levels, and discuss the platform's current architectural limits and the engineering targets they define. The results show that CP substantially widens field-of-view (FoV) coverage and improves occupied-cell recall, and that beyond a moderate localization-noise threshold, positioning uncertainty, and not weather, becomes the dominant error source. We outline the platform's trajectory toward a Mediterranean operational design domain (ODD) testing service.
Chinese Translation
欧洲安全法规现在允许通过经过验证的物理-虚拟设施在虚拟环境中生成大量自动驾驶认证证据。我们提出了一种部署的混合车辆环路(Vehicle-in-the-Loop, ViL)平台,该平台通过V2X消息管道将一辆真实的仪器化车辆与基于CARLA的数字双胞胎(Digital Twin, DT)相结合,并报告其在公共道路代表性测试轨道上的首次集成操作。一辆真实车辆将符合ETSI标准的CAM/CPM消息流入DT,在场景运行期间,GPU加速的协同感知(Cooperative Perception, CP)模块将这些消息融合为概率占用网格。我们在一个多车辆双T型交叉口场景中演示了该平台,表征了在正常、雨天和夜间条件下以及五种定位噪声水平下的CP工作负载,并讨论了平台当前的架构限制及其定义的工程目标。结果表明,CP显著扩大了视场(Field-of-View, FoV)覆盖范围,并提高了占用单元的召回率,并且在超过适度的定位噪声阈值后,定位不确定性而非天气成为主要误差来源。我们概述了该平台朝向地中海操作设计域(Operational Design Domain, ODD)测试服务的轨迹。