← Back to Index
Daily Autonomous Driving Research Digest

AutoDrive Papers

2026-07-30
10
Papers
6
Topics
10
Translated

感知

4
感知 / 1 / 2607.26165

DVPSFormer: Efficient Online Depth-aware Video Panoptic Segmentation for Autonomous Driving

DVPSFormer:用于自主驾驶的高效在线深度感知视频全景分割
Yung-Hsu Yang, Luigi Piccinelli, Siyuan Li, Mattia Segu, Lei Ke, Martin Danelljan, Yuqian Fu, Zuria Bauer, Fisher Yu, Hermann Blum, Marc Pollefeys
cs.CV
Abstract
Safe autonomous navigation requires a holistic understanding of dynamic environments, necessitating the simultaneous estimation of metric depth, semantic segmentation, and instance trajectories. While depth-aware video panoptic segmentation (DVPS) unifies these tasks, existing approaches often rely on computationally expensive, multi-stage pipelines or offline tracking, rendering them unsuitable for real-time decision-making. To address this, we propose DVPSFormer, a unified online architecture designed for efficient 4D scene understanding. Central to our approach is explicit scene discretization (ESD), a novel mechanism that leverages segmentation queries to represent foreground and background regions, enabling a discrete-to-continuous (D2C) depth head to decode metric depth in a single pass. This tightly couples semantic and geometric learning while significantly reducing latency. Furthermore, we propose an online majority voting (OMV) mechanism that exploits temporal consistency to refine classification during instance tracking. DVPSFormer establishes a new state-of-the-art on the Cityscapes-DVPS and SemKITTI-DVPS benchmarks, offering a streamlined solution for online robotic perception. Code and models are available at https://royyang0714.github.io/DVPSFormer.
Chinese Translation
安全的自主导航需要对动态环境的整体理解,这要求同时估计度量深度、语义分割和实例轨迹。虽然深度感知视频全景分割(DVPS)将这些任务统一起来,但现有方法通常依赖于计算开销大的多阶段流程或离线跟踪,使其不适合实时决策。为了解决这个问题,我们提出了DVPSFormer,这是一种旨在高效进行4D场景理解的统一在线架构。我们方法的核心是显式场景离散化(ESD),这是一种新颖的机制,利用分割查询来表示前景和背景区域,使得离散到连续(D2C)深度头能够在一次传递中解码度量深度。这紧密结合了语义和几何学习,同时显著降低了延迟。此外,我们提出了一种在线多数投票(OMV)机制,利用时间一致性在实例跟踪期间细化分类。DVPSFormer在Cityscapes-DVPS和SemKITTI-DVPS基准测试中建立了新的最先进水平,为在线机器人感知提供了一种简化的解决方案。代码和模型可在 https://royyang0714.github.io/DVPSFormer 获取。
感知 / 2 / 2607.26283

HeteroPROPMT: A Real-time and Privacy-Preserving Heterogeneous Collaborative Perception Framework

HeteroPROPMT:一种实时且保护隐私的异构协同感知框架
Armin Maleki, Hayder Radha
cs.CV · cs.RO
Abstract
Collaborative Perception (CP) improves autonomous systems' awareness of their surroundings by sharing sensor data, intermediate features, and detection results. In real-world deployments, however, collaborating vehicles often use heterogeneous sensors, perception models, datasets, and training domains, creating feature-space shifts that degrade downstream fusion and detection. Existing approaches typically retrain fusion and detection components or introduce modality-specific feature interpreters. These methods scale poorly to newly joining agents and often require access to proprietary metadata, raising privacy concerns. We propose HeteroPROMPT, a real-time and privacy-preserving framework for heterogeneous collaborative perception. HeteroPROMPT rapidly aligns each heterogeneous agent's features with an ego-centric unified feature space through modular prompts and lightweight learning-based tuning, while keeping agent encoders and the collaborative fusion and detection stacks frozen. Its visual prompt-based training and inference modulate Bird's Eye View (BEV) features across channels and spatial locations with low computational overhead. For metadata-free deployment, an autoencoder learns a compact unified representation and extracts modality cues from shared features, enabling real-time modality classification and routing to the appropriate HeteroPROMPT modules without exposing proprietary agent information. Experiments on the OPV2V-H and V2XSet datasets show that HeteroPROMPT improves Average Precision over state-of-the-art heterogeneous CP methods while using orders of magnitude fewer trainable parameters. This offers a scalable and practical CP solution. The proposed modality classifier also predicts the joining agent's modality from compact features with greater than 99.99 percent accuracy during deployment. Code will be available at https://github.com/arminmaleki007/HeteroPROMPT.
Chinese Translation
协同感知(Collaborative Perception, CP)通过共享传感器数据、中间特征和检测结果,提高了自主系统对周围环境的感知。然而,在实际部署中,协作车辆通常使用异构传感器、感知模型、数据集和训练领域,导致特征空间的偏移,从而降低了下游融合和检测的效果。现有方法通常需要重新训练融合和检测组件,或引入特定模态的特征解释器。这些方法在新加入的代理上扩展性较差,且通常需要访问专有元数据,从而引发隐私问题。我们提出了HeteroPROMPT,一种实时且保护隐私的异构协同感知框架。HeteroPROMPT通过模块化提示和轻量级学习调优,快速将每个异构代理的特征与以自我为中心的统一特征空间对齐,同时保持代理编码器和协作融合及检测堆栈的冻结。其基于视觉提示的训练和推理在通道和空间位置上调节鸟瞰图(Bird's Eye View, BEV)特征,计算开销低。为了实现无元数据的部署,自动编码器学习紧凑的统一表示,并从共享特征中提取模态线索,使得实时模态分类和路由到适当的HeteroPROMPT模块成为可能,而无需暴露专有代理信息。在OPV2V-H和V2XSet数据集上的实验表明,HeteroPROMPT在使用数量级更少的可训练参数的情况下,提升了与最先进的异构CP方法相比的平均精度。这提供了一种可扩展且实用的CP解决方案。所提出的模态分类器在部署期间也能以超过99.99%的准确率从紧凑特征中预测加入代理的模态。代码将可在 https://github.com/arminmaleki007/HeteroPROMPT 获取。
感知 / 3 / 2607.26651

Physically Real-time Infrared Attack against Optical Flow Estimation Networks

针对光流估计网络的物理实时红外攻击
Shen You, Wei Jiang, Jiarui Liu, Yijian Ye, Qiuzhen Lin, Xiangtao Li, Ka-Chun Wong
cs.CV · cs.AI
Abstract
With the promising performance of deep neural networks on image-based tasks, different real-world applications such as autonomous driving and motion detection have become increasingly mature and relevant to human lives. In particular, Optical Flow Estimation Networks (OFENs), as upstream models, play a critical role in different domains. Its outputs are heavily assumed and adopted for different downstream tasks, and it is essential to test its robustness to prevent safety accidents. We present an approach for real-time attacks on OFENs in the physical world, leveraging infrared lights for their stealthiness. By generating a large number of Adversarial Examples in advance, our approach computes AEs in real time and dynamically displays them, which allows our method to facilitate precise and targeted attacks without modifying the victim system. Unlike previous digital-to-physical attack techniques, our method directly attacks victim models within the physical world, thereby overcoming the limitations associated with the ineffectiveness of AEs. Experimental results demonstrate the efficacy of our approach in compromising OFENs across diverse lighting conditions, varying object motion velocities, and different object placements, ultimately impairing the network's ability to accurately estimate optical flow.
Chinese Translation
随着深度神经网络在基于图像任务中的出色表现,自动驾驶和运动检测等不同的现实应用变得愈加成熟,并与人类生活息息相关。特别是光流估计网络(Optical Flow Estimation Networks, OFENs)作为上游模型,在不同领域中发挥着关键作用。其输出被广泛假设和应用于不同的下游任务,因此测试其鲁棒性以防止安全事故至关重要。我们提出了一种针对OFENs的物理世界实时攻击方法,利用红外光的隐匿性。通过提前生成大量对抗样本(Adversarial Examples, AEs),我们的方法能够实时计算对抗样本并动态展示,从而实现精确和有针对性的攻击,而无需修改受害系统。与之前的数字到物理攻击技术不同,我们的方法直接在物理世界中攻击受害模型,从而克服了对抗样本无效性所带来的局限性。实验结果表明,我们的方法在不同光照条件、变化的物体运动速度和不同物体放置情况下有效地破坏了光流估计网络的性能,最终削弱了网络准确估计光流的能力。
感知 / 4 / 2607.27058

Object Detection for Autonomous Driving in Chinese Rural Scenes: An Experimental Study on Real-Synthetic Data Mixing and Model Evaluation

中国农村场景下的自动驾驶物体检测:真实与合成数据混合及模型评估的实验研究
Danning Zhu, Ziyan Lin, Jing Wu
cs.CV
Abstract
Currently, autonomous driving object detection models face significant data scarcity and generalization challenges when navigating complex Chinese rural traffic scenarios. To address these limitations, we propose a novel real-synthetic mixed object detection dataset tailored specifically for Chinese rural roads and systematically evaluate the performance of 13 mainstream detectors under different real-to-synthetic data ratios, thereby providing empirical evidence for model selection and data strategy design in rural autonomous driving scenarios. Our dataset combines real-world images captured in Weishi County, Henan Province, with parameterized virtual scenes generated via Unreal Engine. To accurately reflect the unique realities of rural traffic, we define a comprehensive 14-category object system encompassing region-specific elements such as electric tricycles, low-speed vehicles (LSVs), and roadside stalls. Under a unified training protocol, we systematically evaluate 13 mainstream detectors -- spanning the YOLOv5, YOLOv8, YOLO11, and YOLO26 series, as well as RT-DETR-L -- across three data configurations: an all-real baseline, a 1:0.5 real-to-virtual mix, and a 1:1 mix. Experimental results demonstrate that a moderate injection of synthetic data (1:0.5 ratio) effectively enhances detection performance, with YOLO11m achieving the highest [email protected] of 0.758. However, a higher proportion of synthetic data (1:1) introduces domain shifts that offset the benefits of data scaling. While most models reliably identify distinct local vehicles, significant perceptual bottlenecks remain for long-tail, non-standard objects like stalls and railings. This research provides crucial empirical evidence and novel insights for model selection and synthetic data strategies, facilitating the practical deployment of autonomous driving perception systems in rural areas.
Chinese Translation
目前,自动驾驶物体检测模型在复杂的中国农村交通场景中面临显著的数据稀缺和泛化挑战。为了解决这些限制,我们提出了一种新颖的专门针对中国农村道路的真实-合成混合物体检测数据集,并系统地评估了13种主流检测器在不同真实与合成数据比例下的性能,从而为农村自动驾驶场景中的模型选择和数据策略设计提供实证依据。我们的数据集结合了在河南省卫辉市捕获的真实世界图像与通过虚幻引擎(Unreal Engine)生成的参数化虚拟场景。为了准确反映农村交通的独特现实,我们定义了一个涵盖区域特定元素的综合14类物体系统,包括电动三轮车、低速车辆(Low-Speed Vehicles, LSVs)和路边摊等。在统一的训练协议下,我们系统地评估了13种主流检测器——涵盖YOLOv5、YOLOv8、YOLO11和YOLO26系列,以及RT-DETR-L——在三种数据配置下的表现:全真实基线、1:0.5的真实与虚拟混合,以及1:1的混合。实验结果表明,适度注入合成数据(1:0.5比例)有效提升了检测性能,其中YOLO11m达到了最高的 [email protected] 为0.758。然而,更高比例的合成数据(1:1)引入了领域转移,抵消了数据扩展的好处。尽管大多数模型能够可靠地识别出不同的本地车辆,但对于长尾的非标准物体,如摊位和护栏,仍然存在显著的感知瓶颈。本研究为模型选择和合成数据策略提供了重要的实证依据和新颖的见解,促进了自动驾驶感知系统在农村地区的实际部署。

LiDAR/点云

2
LiDAR/点云 / 1 / 2607.26855

NeoRacer: An Open, Standardized 1:12 Scale Autonomous Race Car for Benchmarking and Education

NeoRacer:一个开放、标准化的1:12比例自主赛车平台,用于基准测试和教育
Koneshka Bandyopadhyay, Ansh Mehta, Bassel El Mabsout, Renato Mancuso
cs.RO · eess.SY
Abstract
Many scientific fields rely on standard benchmarks and shared platforms to improve review and reproducibility, but autonomous systems research still lacks widely accepted open hardware. Where standardization has emerged, progress has accelerated. This is especially evident in autonomous racing, where teams often build custom systems or buy niche, expensive vehicles, making control and robotics research and education hard to compare and reproduce. High costs also limit access outside well-funded labs, while affordable educational robots are often underpowered. To address this gap, we present NeoRacer, an open-source 1:12 scale autonomous racing platform. It is built around an NVIDIA Jetson Orin Nano (67 TOPS), a 270{\deg} LiDAR, a 120 fps global-shutter camera, and a 9-axis IMU. NeoRacer ships pre-assembled for USD 2,699, offering over 3x the compute of comparable platforms at less than half the cost of the nearest pre-assembled alternative. Co-developed by the Neobotics Foundation and Seeed Studio, and manufactured by Seeed Studio, NeoRacer combines open hardware and software design with scalable, repeatable production. The modular, extensible platform provides a standardized benchmarking environment for autonomous racing algorithms across institutions. We describe the hardware/software architecture, design decisions from two pilot deployments (MIT IAP, 15 students; BU CPS Lab, 10 students), and key cost-performance tradeoffs. Hardware is licensed under CERN-OHL-S v2 and software under GPLv3, with all design files, firmware, and ROS2 packages publicly accessible.
Chinese Translation
许多科学领域依赖标准基准和共享平台来提高评审和可重复性,但自主系统研究仍然缺乏广泛接受的开放硬件。在标准化出现的地方,进展得以加速。这在自主赛车中尤为明显,团队通常构建定制系统或购买小众、昂贵的车辆,使得控制和机器人研究及教育难以比较和重复。高昂的成本也限制了资金不足的实验室的访问,而价格适中的教育机器人往往性能不足。为了解决这一差距,我们推出了NeoRacer,一个开放源代码的1:12比例自主赛车平台。它基于NVIDIA Jetson Orin Nano(67 TOPS)、270° LiDAR、120 fps全局快门相机和9轴IMU构建。NeoRacer以2,699美元的价格预装配出售,提供超过3倍于可比平台的计算能力,而成本不到最近预装配替代品的一半。NeoRacer由Neobotics基金会和Seeed Studio共同开发,并由Seeed Studio制造,结合了开放硬件和软件设计与可扩展、可重复的生产。该模块化、可扩展的平台为各机构的自主赛车算法提供了标准化的基准测试环境。我们描述了硬件/软件架构、两个试点部署(MIT IAP,15名学生;BU CPS实验室,10名学生)的设计决策,以及关键的成本-性能权衡。硬件遵循CERN-OHL-S v2许可证,软件遵循GPLv3许可证,所有设计文件、固件和ROS2包均可公开访问。
LiDAR/点云 / 2 / 2607.26645

FPSGen: Flexible Point Cloud Scene Generation with BEV-Supported Transport Flows

FPSGen:基于鸟瞰视图支持的运输流的灵活点云场景生成
Wenzhe He, Meng Wang, JiaWei Qian, Jinfeng Xu, Ying Liu, Ruihui Li
cs.CV · cs.AI
Abstract
Existing point-based generative methods for outdoor scenes primarily focus on LiDAR-conditioned completion. During training, noisy point clouds are constructed by perturbing complete ground-truth scenes, whereas during inference, they are initialized by adding noise to duplicated partial scans. This train-inference mismatch inherits the sparsity and visibility bias of partial scans, leading to sparse distant regions and incomplete geometry in occluded areas. Moreover, the reliance on partial scans restricts generation when LiDAR observations are unavailable or replaced by layout cues. We present FPSGen, a flexible framework that constructs point sources independently of partial scans. FPSGen first predicts a bird's-eye-view (BEV) prior with density, height, and mask channels from the active cues. The density map is then sampled to form a BEV-supported point source, enabling both unconditional and conditioned initialization. A teacher-student approximate optimal transport scheme then uses teacher-predicted endpoints to learn a velocity field that induces straighter transport paths. By integrating BEV point source construction with path-straightening transport, FPSGen provides a unified framework for unconditional and flexible cue-conditioned scene generation. Extensive experiments show that FPSGen achieves state-of-the-art JSD and voxel IoU performance on SemanticKITTI completion while maintaining strong performance with a single point transport step. On KITTI-360 unconditional generation, it also achieves the best Coverage (COV) among the compared methods.
Chinese Translation
现有的基于点的户外场景生成方法主要集中在基于激光雷达(LiDAR)条件的补全。在训练过程中,通过扰动完整的真实场景构建噪声点云,而在推理过程中,则通过向重复的部分扫描添加噪声来初始化。这种训练-推理不匹配继承了部分扫描的稀疏性和可见性偏差,导致远处区域稀疏以及遮挡区域几何形状不完整。此外,依赖部分扫描限制了在没有LiDAR观测或被布局线索替代时的生成能力。我们提出了FPSGen,一个独立于部分扫描构建点源的灵活框架。FPSGen首先从活跃线索中预测一个包含密度、高度和掩膜通道的鸟瞰视图(BEV)先验。然后对密度图进行采样,以形成一个支持BEV的点源,从而实现无条件和有条件的初始化。接着,教师-学生近似最优运输方案利用教师预测的端点学习一个速度场,以诱导更直的运输路径。通过将BEV点源构建与路径直线化运输相结合,FPSGen提供了一个统一的框架,用于无条件和灵活的线索条件场景生成。大量实验表明,FPSGen在SemanticKITTI补全任务上实现了最先进的JSD和体素IoU性能,同时在单点运输步骤下保持强劲表现。在KITTI-360无条件生成中,它在比较方法中也实现了最佳覆盖率(COV)。

深度/几何

1
深度/几何 / 1 / 2607.26600

JEPADepth: Masked Predictive Representation Learning for Self-Supervised Monocular Depth Estimation

JEPADepth:用于自监督单目深度估计的掩码预测表示学习
Ionuţ Grigore, Călin-Adrian Popa
cs.CV
Abstract
Self-supervised monocular depth estimation typically relies on photometric reconstruction losses that couple depth, pose, and appearance assumptions. In this paper, we propose JEPADepth, a self-supervised monocular depth framework that incorporates a complementary training objective inspired by Image Joint-Embedding Predictive Architectures (I-JEPA) for self-supervised depth learning. Our method augments a standard photometric pipeline with a masked prediction loss computed in the representation space of a pretrained DINOv3 Vision Transformer encoder. A predictor infers target-region embeddings from visible context-region embeddings under structured masking, and is discarded along with the target encoder at inference time, adding no deployment cost. On KITTI, adding the JEPA objective consistently improves performance over the same DINOv3-based photometric baseline, without changing the inference-time architecture. Compared to prior monocular self-supervised methods, JEPADepth is competitive with state-of-the-art transformer-based approaches and outperforms strong CNN-based baselines on the standard benchmark. In zero-shot transfer (trained on KITTI and evaluated without fine-tuning), JEPADepth achieves the best or near-best performance among the compared methods on both Make3D and Cityscapes across multiple metrics.
Chinese Translation
自监督单目深度估计通常依赖于光度重建损失,这些损失将深度、姿态和外观假设结合在一起。在本文中,我们提出了JEPADepth,一种自监督单目深度框架,结合了受图像联合嵌入预测架构(Image Joint-Embedding Predictive Architectures, I-JEPA)启发的互补训练目标,用于自监督深度学习。我们的方法在标准光度管道中增加了一个在预训练DINOv3视觉变换器编码器的表示空间中计算的掩码预测损失。预测器在结构化掩码下从可见上下文区域嵌入推断目标区域嵌入,并在推理时与目标编码器一起被丢弃,不增加部署成本。在KITTI数据集上,添加JEPA目标始终改善了相同基于DINOv3的光度基线的性能,而不改变推理时的架构。与之前的单目自监督方法相比,JEPADepth在与最先进的基于变换器的方法竞争时表现出色,并在标准基准上超越了强大的基于卷积神经网络(CNN)的基线。在零样本迁移(在KITTI上训练并在不进行微调的情况下评估)中,JEPADepth在Make3D和Cityscapes的多个指标上,在比较的方法中实现了最佳或接近最佳的性能。

预测/规划

1
预测/规划 / 1 / 2607.26802

Risk-Aware Motion Planning with Learned Trajectory Primitives and Probabilistic Safety Assessment

基于学习的轨迹原语和概率安全评估的风险感知运动规划
Marc Kaufeld, Dian Zhuang, Johannes Betz
cs.RO
Abstract
This paper presents a radial basis function network (RBFN)-informed motion planning framework for safe and efficient urban autonomous driving. The proposed approach combines RBFN-based candidate trajectory generation with an analytic collision probability assessment and optimization-based trajectory refinement. The network learns jerk-minimal trajectories, enabling the MPC to operate within a reduced and dynamically consistent search space. Candidate motion primitives are selected based on an accurate probabilistic risk measure. This design decreases solver complexity while preserving safety and constraint satisfaction. The framework is evaluated in numerous urban driving scenarios. Results demonstrate improved risk awareness and fewer vehicle-limit violations compared to benchmark methods. The proposed approach integrates learning-based trajectories into optimization-based motion planning, thereby ensuring safety and interpretability.
Chinese Translation
本文提出了一种基于径向基函数网络(RBFN)的运动规划框架,以实现安全高效的城市自动驾驶。所提出的方法结合了基于RBFN的候选轨迹生成、解析碰撞概率评估和基于优化的轨迹优化。该网络学习最小颠簸的轨迹,使得模型预测控制(MPC)能够在一个减少且动态一致的搜索空间内运行。候选运动原语的选择基于准确的概率风险度量。该设计在保持安全性和约束满足的同时,降低了求解器的复杂性。该框架在多个城市驾驶场景中进行了评估。结果表明,与基准方法相比,风险意识得到了改善,车辆限值违规情况减少。所提出的方法将基于学习的轨迹集成到基于优化的运动规划中,从而确保安全性和可解释性。

世界模型/生成

1
世界模型/生成 / 1 / 2607.27036

Mitigating Compounding Error via Video Representation Regularization

通过视频表示正则化减轻复合误差
Taiye Chen, Qi Zhang, Yisen Wang
cs.CV · cs.LG
Abstract
Video diffusion-based world models enable long autoregressive video generation for robotics, autonomous driving and simulation tasks, yet sliding-window autoregressive inference suffers from severe error accumulation that degrades frame quality over time. Although this phenomenon has been widely observed, the underlying mechanism of compounding error and how to achieve stable long-horizon generation remain largely unresolved. In this paper, we investigate the internal representation dynamics of video world models and discover that compounding error is tightly coupled with dimensional collapse of hidden representations. Specifically, the effective rank of model representations sharply decreases at the onset of generation drift, revealing a strong connection between representational degradation and long-term rollout instability. Furthermore, we find that pure training data scaling fails to boost model resistance to error drift, contradicting mainstream scaling paradigms. To address this problem, we propose video representation regularization, a lightweight training constraint that stabilizes latent representations and suppresses iterative error accumulation. Compared with Diffusion Forcing, our method achieves improvements from 38.65 to 55.56 and from 44.37 to 72.08 on the Aesthetic Quality and Imaging Quality metrics of VBench. Our work establishes the first connection between autoregressive video drifting and model internal representations, adopts erank as a quantitative metric for error accumulation, reveals counterintuitive scaling limitations for video world models, and presents a simple yet effective regularization strategy to improve long video generation robustness.
Chinese Translation
基于视频扩散的世界模型使得机器人、自动驾驶和仿真任务能够进行长时间的自回归视频生成,但滑动窗口自回归推理在时间上严重的误差累积导致帧质量下降。尽管这一现象已被广泛观察,但复合误差的潜在机制以及如何实现稳定的长时间生成仍然在很大程度上未得到解决。本文研究了视频世界模型的内部表示动态,发现复合误差与隐藏表示的维度崩溃紧密相关。具体而言,模型表示的有效秩在生成漂移开始时急剧下降,揭示了表示退化与长期展开不稳定性之间的强关联。此外,我们发现单纯扩大训练数据并未能增强模型对误差漂移的抵抗力,这与主流的扩展范式相悖。为了解决这一问题,我们提出了视频表示正则化,这是一种轻量级的训练约束,能够稳定潜在表示并抑制迭代误差累积。与Diffusion Forcing相比,我们的方法在VBench的美学质量和成像质量指标上分别实现了从38.65到55.56和从44.37到72.08的提升。我们的工作首次建立了自回归视频漂移与模型内部表示之间的联系,采用有效秩(erank)作为误差累积的定量指标,揭示了视频世界模型的反直觉扩展限制,并提出了一种简单而有效的正则化策略,以提高长视频生成的鲁棒性。

仿真/数据

1
仿真/数据 / 1 / 2607.27085

Controlled Experiments on Lane Changing by Transitional Autonomous Vehicle: Dataset and Behavioral Insights

过渡性自动驾驶车辆的变道控制实验:数据集与行为洞察
Abhinav Sharma, Md Abdullah Al Hasan, Danjue Chen, George F. List
cs.RO · eess.SY
Abstract
This paper presents the North Carolina Transitional Autonomous Vehicle Lane-Changing (NC-tALC) dataset and uses it to characterize mandatory lane-changing behavior of transitional automated vehicles (tAVs). It quantifies the evolution of lead--lag gaps throughout the lane-change process and examines how potential collision risk develops during the maneuver. A controlled field experiment comprising 78 mandatory lane-change trials was conducted on a public roadway in Apex, North Carolina. Four instrumented vehicles created repeatable traffic conditions while varying the lane changer's initial position within the candidate target gap. High-resolution RTK-GNSS/INS trajectories were processed to identify key timestamps, calculate lead, lag, and lane-change gaps, and estimate interactions using time-gap- and speed-based surrogate safety measures. Despite substantial differences in initial conditions, lead and lag gaps consistently converged toward a relatively narrow range near lane crossing. Potential collision risk increased as the maneuver progressed, peaked near physical lane entry, and was dominated by interactions with the target-lane leader. Lane-change completion did not necessarily coincide with the disappearance of collision risk. This study provides one of the first controlled empirical characterizations of the complete mandatory lane-change process of tAVs using repeatable public-road experiments. The NC-tALC dataset supports analysis of behavioral and safety evolution throughout the maneuver rather than only at the gap-acceptance instant. The dataset and findings provide empirical benchmarks for evaluating automated lane-changing behavior, calibrating behavioral models, and validating simulation and safety assessment methods for mandatory lane-change scenarios.
Chinese Translation
本文介绍了北卡罗来纳州过渡性自动驾驶车辆变道(NC-tALC)数据集,并利用该数据集对过渡性自动化车辆(tAVs)的强制变道行为进行了特征描述。研究量化了变道过程中前后间隙的演变,并考察了在该操作过程中潜在碰撞风险的发展。我们在北卡罗来纳州阿佩克斯的一条公共道路上进行了包含78次强制变道试验的受控实地实验。四辆配备仪器的车辆创造了可重复的交通条件,同时改变了变道者在候选目标间隙中的初始位置。高分辨率RTK-GNSS/INS轨迹被处理以识别关键时间戳,计算前后间隙和变道间隙,并使用基于时间间隙和速度的替代安全措施来估计交互作用。尽管初始条件存在显著差异,前后间隙在接近变道时始终趋向于相对狭窄的范围。随着操作的进行,潜在碰撞风险增加,在物理车道入口附近达到峰值,并主要受目标车道领头车辆的交互影响。变道的完成并不一定与碰撞风险的消失同时发生。本研究提供了对tAVs完整强制变道过程的首次受控实证特征描述,基于可重复的公共道路实验。NC-tALC数据集支持在整个操作过程中分析行为和安全的演变,而不仅仅是在间隙接受的瞬间。该数据集及其研究结果为评估自动变道行为、校准行为模型以及验证强制变道场景的仿真和安全评估方法提供了实证基准。