← Back to Index
Daily Autonomous Driving Research Digest

AutoDrive Papers

2026-07-17
6
Papers
5
Topics
6
Translated

感知

2
感知 / 1 / 2607.14248

3D Lane Detection with Odometry for High-Speed Vehicle Racing

基于里程计的高速赛车3D车道检测
Omoruyi Atekha, John Subosits, Marcus Greiff
cs.CV · eess.IV
Abstract
Lane boundary detection is a critical component in autonomous driving systems and has been rigorously studied in regular driving scenarios. However, it is less explored in vehicle racing, where the car moves at higher speeds across more extreme road geometries. To study this problem, we introduce a new dataset for 3D lane detection in racing, featuring >$250$k images from multiple camera feeds and inertial measurements taken with a Lexus LC 500 driving on a closed circuit. With this dataset, we compare various approaches to 3D lane detection and propose modifications that permit frames to be processed at rates of almost 300Hz while retaining high predictive performance in the racing application. This facilitates a multi-camera ensemble approach that is validated on hardware. We show that sensing modalities such as inertial measurements can be leveraged for pre-integration to regress road geometries over both cameras and time, yielding improvements in key metrics. Compared to methods such as BevLaneDet, adding odometry and ensemble predictions improves the F1 score by 3 points and reduces near-vehicle mean absolute errors (MAEs) by $>30 \%$. We show F1 scores $>$0.9 and lateral MAEs of $<$0.18m in vehicle deployments.
Chinese Translation
车道边界检测是自动驾驶系统中的关键组成部分,并且在常规驾驶场景中得到了严格的研究。然而,在赛车中,由于车辆以更高的速度行驶并穿越更极端的道路几何形状,这一领域的研究相对较少。为了解决这个问题,我们引入了一个新的数据集,用于赛车中的3D车道检测,该数据集包含来自多个摄像头和使用雷克萨斯 LC 500 在封闭赛道上行驶时获取的惯性测量的超过250,000张图像。利用该数据集,我们比较了多种3D车道检测方法,并提出了修改方案,使得帧处理速率接近300Hz,同时在赛车应用中保持高预测性能。这促进了一种多摄像头集成方法,并在硬件上进行了验证。我们展示了惯性测量等传感方式可以用于预集成,以回归摄像头和时间上的道路几何形状,从而在关键指标上取得了改进。与 BevLaneDet 等方法相比,添加里程计和集成预测将 F1 分数提高了3分,并将近车道的平均绝对误差(MAE)减少了超过30%。我们在车辆部署中展示了 F1 分数超过0.9,横向 MAE 小于0.18米。
感知 / 2 / 2607.14710

Variational Inference for Bird's Eye View Segmentation in Autonomous Driving

用于自动驾驶的鸟瞰图分割的变分推断
Jingyue Shi, Huaicheng Li, Junhui Zhao, Yanxiang Jiang
cs.CV · eess.SY
Abstract
The bird's eye view (BEV) has emerged as a pivotal approach for environmental perception in autonomous driving, providing a unified spatial representation for vehicles. Nevertheless, despite BEV's significance in addressing the challenges inherent to autonomous driving, effectively fusing data from multiple camera sensors and operating in complex external driving environments remains a considerable challenge. To mitigate this issue, we recast the BEV segmentation problem within a variational inference framework. In this paper, we propose a novel transformer-based variational flow transformation network for BEV segmentation, denoted as TVB. Our architecture implicitly learns the mapping from multiple camera views to a unified canonical BEV map during training by exploiting posterior BEV supervision. TVB employs a conditional variational auto encoder (CVAE) as its backbone and produces multiple BEV map candidates. To augment the realism of the generated BEV maps, we integrate normalizing flows into the map generation process, enabling the construction of more complex and expressive probability distributions. Furthermore, we design a BEV-attention fusion (BAF) module that harnesses attention mechanisms to adaptively integrate the multiple candidate BEV maps. Experimental results, evaluated on both the nuScenes and OPV2Vdatasets, demonstrate that our proposed method achieves superior performance in multi-camera view BEV segmentation and lane environment perception.
Chinese Translation
鸟瞰图(BEV)已成为自动驾驶环境感知的关键方法,为车辆提供了统一的空间表示。然而,尽管BEV在解决自动驾驶固有挑战方面具有重要意义,但有效融合来自多个摄像头传感器的数据并在复杂的外部驾驶环境中操作仍然是一个相当大的挑战。为了解决这个问题,我们将BEV分割问题重新表述为变分推断框架。在本文中,我们提出了一种新颖的基于变压器的变分流变换网络用于BEV分割,称为TVB。我们的架构在训练过程中通过利用后验BEV监督隐式学习从多个摄像头视图到统一标准BEV图的映射。TVB采用条件变分自编码器(CVAE)作为其主干,并生成多个BEV图候选。为了增强生成BEV图的真实感,我们将归一化流整合到图生成过程中,使得构建更复杂和更具表现力的概率分布成为可能。此外,我们设计了一个BEV注意力融合(BAF)模块,利用注意力机制自适应地整合多个候选BEV图。实验结果在nuScenes和OPV2V数据集上评估,表明我们提出的方法在多摄像头视图BEV分割和车道环境感知方面实现了优越的性能。

深度/几何

1
深度/几何 / 1 / 2607.15048

RoGS: Adaptive Meshgrid Gaussian for Large-Scale Road Surface Mapping

RoGS:用于大规模道路表面映射的自适应网格高斯模型
Tianchen Deng, Zhiheng Feng, Wenhua Wu, Ziming Li, Siting Zhu, Hesheng Wang
cs.CV
Abstract
Road surface mapping plays a crucial role in autonomous driving, supporting high-definition map generation, lane-level perception, and automatic road annotation. Recent mesh-based road surface reconstruction methods have shown promising results, but they still suffer from limited reconstruction quality and high optimization cost, especially in large-scale driving scenarios. To address these limitations, we propose ROADGS-T, a robust and efficient large-scale road surface mapping framework based on adaptive meshgrid Gaussian representation. Specifically, we model the road surface by placing 2D Gaussian surfels on a meshgrid, where each surfel explicitly stores color, semantic, and geometric information. Compared with conventional mesh-based representations and 3D Gaussian primitives, the proposed meshgrid Gaussian representation better matches the thin-surface property of roads while significantly reducing redundant primitives and overlap during optimization. To further improve representation efficiency and structural fidelity, we introduce a road-structure-aware adaptive meshgrid strategy, which allocates denser Gaussian surfels to geometrically or semantically complex regions, such as lane markings, road boundaries, and height discontinuities, while maintaining a compact representation in flat road areas. Moreover, instead of relying on a single nearest vehicle pose, we design a trajectory-consistency-guided pose-robust refinement strategy, which estimates local surface priors from multiple neighboring poses and adaptively weights pose-guided height regularization according to their geometric consistency.
Chinese Translation
道路表面映射在自动驾驶中起着至关重要的作用,支持高清地图生成、车道级感知和自动道路标注。最近的基于网格的道路表面重建方法显示出良好的结果,但在大规模驾驶场景中,它们仍然面临重建质量有限和优化成本高的问题。为了解决这些局限性,我们提出了ROADGS-T,一个基于自适应网格高斯表示的稳健高效的大规模道路表面映射框架。具体而言,我们通过在网格上放置二维高斯表面元素(surfel)来建模道路表面,其中每个表面元素明确存储颜色、语义和几何信息。与传统的基于网格的表示和三维高斯原语相比,所提出的网格高斯表示更好地匹配了道路的薄表面特性,同时在优化过程中显著减少了冗余原语和重叠。为了进一步提高表示效率和结构保真度,我们引入了一种道路结构感知的自适应网格策略,该策略将更密集的高斯表面元素分配给几何或语义复杂的区域,如车道标记、道路边界和高度不连续性,同时在平坦道路区域保持紧凑的表示。此外,我们设计了一种轨迹一致性引导的姿态鲁棒性优化策略,而不是依赖单一的最近车辆姿态,该策略从多个邻近姿态中估计局部表面先验,并根据几何一致性自适应地加权姿态引导的高度正则化。

预测/规划

1
预测/规划 / 1 / 2607.14727

WorkDrive: Roadwork Chain of Causation for Autonomous Driving

WorkDrive:自主驾驶的道路施工因果链
Tianyi Jiang, Wen Zhang, Sihan Yang, Ming Lu, Wentao Zhang
cs.CV
Abstract
Autonomous driving vision-language models (VLMs) struggle in roadwork zones, where familiar visual cues such as lane markings and permanent signs are altered or absent, and temporary devices such as cones and barriers redefine the drivable corridor. VLMs can detect these objects, but without explicit guidance they anchor their reasoning on familiar elements from pre-training and fail to connect work-zone observations to correct planning decisions. We propose WorkDrive, a framework that constructs perception-grounded causal reasoning for work zones and aligns it with trajectory prediction. An automated multitask perception pipeline extracts structured scene facts and injects them into a Chain-of-Causation (CoC) annotation pipeline, redirecting the annotator's attention to domain-specific elements. The resulting reasoning labels are used for supervised fine-tuning, followed by reinforcement learning with a single reward: consistency between lateral meta-actions and the predicted trajectory. On ROADWork, the largest public work-zone dataset, the proposed roadwork CoC reduces trajectory average displacement error (ADE) by 9.0\%, and consistency-based GRPO yields a further 3.0\%, achieving progressive improvement over the trajectory-only baseline. Code and data will be publicly released.
Chinese Translation
自主驾驶的视觉-语言模型(VLMs)在道路施工区域面临挑战,因为熟悉的视觉线索如车道标记和永久性标志被改变或缺失,而临时设备如锥形标和障碍物重新定义了可驾驶走廊。VLMs可以检测这些物体,但在没有明确指导的情况下,它们将推理锚定在预训练时熟悉的元素上,无法将施工区的观察与正确的规划决策联系起来。我们提出了WorkDrive,一个为施工区构建感知基础的因果推理框架,并将其与轨迹预测对齐。一个自动化的多任务感知管道提取结构化场景事实,并将其注入因果链(Chain-of-Causation, CoC)注释管道,重新引导注释者的注意力到特定领域的元素。生成的推理标签用于监督微调,随后进行强化学习,采用单一奖励:横向元动作与预测轨迹之间的一致性。在ROADWork这个最大的公共施工区数据集上,所提出的道路施工因果链将轨迹平均位移误差(ADE)降低了9.0%,而基于一致性的GRPO进一步提高了3.0%,实现了相对于仅基于轨迹的基线的逐步改进。代码和数据将公开发布。

仿真/数据

1
仿真/数据 / 1 / 2607.14387

Chat2Scenic: An Iterative RAG-Based Framework for Scenario Generation in Autonomous Driving

Chat2Scenic:一种基于迭代检索增强的自主驾驶场景生成框架
Yuan Gao, Wenting Miao, Mattia Piccinini, Haoyu Wang, Qunying Song, Johannes Betz
cs.AI · cs.RO
Abstract
Validating autonomous driving systems requires diverse, regulation-compliant test scenarios. In simulation-based testing, scenarios are defined as executable scripts. Yet automatically generating such scripts from regulatory descriptions remains an open challenge, and existing approaches face fundamental trade-offs. Retrieval-assemble methods achieve reasonable compilation rates but lack scalability, whereas retrieval-based full-script generation suffers from low compilation success rates. We present Chat2Scenic, the first iterative retrieval-augmented framework to generate scenario scripts in Domain Specific Language (DSL). Specifically, Chat2Scenic provides a chatbot interface that supports interactive scenario refinement and integrates Retrieval-augmented Generation (RAG) to ground scenario generation in regulatory knowledge and DSL syntax. Furthermore, we propose an open benchmark for scenario generation comprising 123 scenarios from various regulations, including NHTSA and United Nations Vehicle Regulations, as well as other sources. Extensive evaluation with State-of-the-Art (SOTA) Large Language Models (LLMs) demonstrates that Chat2Scenic achieves 76.42% Compilation Success Rate (CSR) and 58.17% Framework Accuracy (FA), outperforming existing methods (Retrieval Assemble with 30.08% CSR, 11.03% FA and Retrieval full script generation with 16.26% CSR, 10.86% FA). To facilitate future research, we release our code as open source at https://github.com/TUM-AVS/chat2scenic.
Chinese Translation
验证自主驾驶系统需要多样化且符合规定的测试场景。在基于仿真的测试中,场景被定义为可执行脚本。然而,从监管描述中自动生成此类脚本仍然是一个未解决的挑战,现有方法面临基本的权衡。检索-组装方法实现了合理的编译率,但缺乏可扩展性,而基于检索的完整脚本生成则遭遇低编译成功率。我们提出了Chat2Scenic,这是第一个迭代检索增强框架,用于生成领域特定语言(Domain Specific Language, DSL)中的场景脚本。具体而言,Chat2Scenic提供了一个聊天机器人接口,支持交互式场景细化,并整合了检索增强生成(Retrieval-augmented Generation, RAG),以将场景生成基于监管知识和DSL语法。此外,我们提出了一个开放的场景生成基准,包含来自各种法规的123个场景,包括美国国家公路交通安全管理局(NHTSA)和联合国车辆法规,以及其他来源。与最先进的大型语言模型(State-of-the-Art Large Language Models, SOTA LLMs)的广泛评估表明,Chat2Scenic实现了76.42%的编译成功率(Compilation Success Rate, CSR)和58.17%的框架准确率(Framework Accuracy, FA),优于现有方法(检索组装的CSR为30.08%,FA为11.03%;检索完整脚本生成的CSR为16.26%,FA为10.86%)。为了促进未来的研究,我们将代码作为开源发布,地址为https://github.com/TUM-AVS/chat2scenic。

协同/V2X

1
协同/V2X / 1 / 2607.14688

MIND-CAVs: Multi-Intelligence Negotiation and Decision System for CAVs based on Intent-Driven Autonomy

MIND-CAVs:基于意图驱动自主性的连接自动驾驶车辆的多智能体协商与决策系统
Mainak Mondal, Yihang Feng, Yangchao Luo, Song Han
cs.RO
Abstract
Modern autonomous vehicles largely operate as isolated agents: they rely on on-board perception and decision modules and broadcast Basic Safety Messages (BSMs) that expose only low-level kinematic state. While existing cooperative driving frameworks enable limited sensor sharing, they rarely communicate high-level maneuver intentions, and edge computing is primarily used for content delivery rather than decision arbitration. As a result, current connected autonomy lacks a principled mechanism for making globally consistent, intent-aware coordination decisions across vehicles. To address this gap, we propose MIND-CAVs, a Multi-Intelligence Negotiation and Decision framework for connected autonomous vehicles (CAVs) based on intent-driven autonomy. Each vehicle abstracts raw sensor observations into structured intent representations, exchanges them over V2X links, and receives globally consistent coordination plans from roadside edge servers. Edge agents combine learned and rule-based arbitration mechanisms to negotiate conflicting intents among vehicles, while a cloud platform records decisions for auditing and continual retraining. We implement MIND-CAVs in a CARLA-based AI-in-the-loop platform and evaluate it in multi-lane highway scenarios involving conflicting maneuvers and route-constrained exits. Experimental results show improved maneuver completion time and reduced unsafe proximity and unnecessary braking compared with isolated autonomy, first-come-first-served arbitration, and multi-agent reinforcement learning baselines.
Chinese Translation
现代自动驾驶车辆在很大程度上作为孤立的智能体运作:它们依赖于车载感知和决策模块,并广播基本安全消息(Basic Safety Messages, BSMs),这些消息仅暴露低级运动状态。虽然现有的协作驾驶框架允许有限的传感器共享,但它们很少传达高级机动意图,而边缘计算主要用于内容传递而非决策仲裁。因此,当前的连接自主性缺乏一个原则性的机制来在车辆之间做出全球一致、关注意图的协调决策。为了解决这一问题,我们提出了MIND-CAVs,一个基于意图驱动自主性的连接自动驾驶车辆(Connected Autonomous Vehicles, CAVs)多智能体协商与决策框架。每辆车将原始传感器观察抽象为结构化的意图表示,通过车对一切(Vehicle-to-Everything, V2X)链接进行交换,并从路边边缘服务器接收全球一致的协调计划。边缘代理结合学习和基于规则的仲裁机制,协商车辆之间的冲突意图,而云平台则记录决策以便审计和持续再训练。我们在基于CARLA的AI-in-the-loop平台上实现了MIND-CAVs,并在涉及冲突机动和路线受限出口的多车道高速公路场景中进行了评估。实验结果表明,与孤立自主性、先到先服务仲裁和多智能体强化学习基线相比,机动完成时间有所改善,且不安全接近和不必要制动的情况减少。