← Back to Index
Daily Autonomous Driving Research Digest

AutoDrive Papers

2026-08-06
5
Papers
3
Topics
5
Translated

LiDAR/点云

2
LiDAR/点云 / 1 / 2608.04130

Radar4D-VLM: Proposal-Grounded Temporal 4D Radar Reasoning Across Frozen Language Models

Radar4D-VLM:基于提议的时间性4D雷达推理在冻结语言模型中的应用
Jiaju Han, Xuemeng Sun, Qike Zhang, Xiang Chen, Luwei Yang, Jiahuan Long, Yiwei Wei, Jiujiang Guo, Chengyin Hu
cs.CV
Abstract
Vision-language models for autonomous driving primarily rely on cameras and LiDAR, leaving 4D radar largely unexplored as a standalone perceptual modality despite its robustness to adverse visibility and direct measurement of radial velocity. We introduce Radar4D-VLM, a radar-only temporal vision-language model that reasons from ten consecutive 4D-radar point-cloud sweeps without camera or LiDAR input. Radar4D-VLM extracts geometrically grounded object proposals and organizes radar evidence into a compact hierarchy of object, scene, and kinematic tokens. A parameter-efficient projector maps these tokens into frozen language backbones, while auditable prediction heads jointly model object count, spatial distribution, motion state, collision risk, semantic category, and radial velocity. Radar4D-VLM combines proposal-grounded temporal object tokenization, global scene context, and explicit kinematic tokens within a unified frozen-backbone interface. On sequence-isolated K-Radar development validation, its Top-64 proposal recall reaches 98.13% at 4 m, exceeding fixed-lattice and uniform-random controls by 6.40 and 22.83 percentage points, respectively. We further evaluate 24 matched runs spanning eight frozen Qwen, Phi, Mistral, Llama, and Gemma backbones under an identical adaptation budget. The radar-token interface remains compatible across all five language-model families, while matched aligned, permuted, and no-language controls show sensor dependence but no stable direct-head gain from aligned language supervision. These results establish a reproducible foundation for radar-only multimodal scene and motion reasoning while separating interface compatibility from the benefit of language supervision.
Chinese Translation
自主驾驶的视觉-语言模型主要依赖于摄像头和激光雷达(LiDAR),而4D雷达作为一种独立的感知模态,尽管其在恶劣能见度下的鲁棒性和径向速度的直接测量能力,仍然未得到充分探索。我们提出了Radar4D-VLM,这是一种仅基于雷达的时间性视觉-语言模型,它从连续十个4D雷达点云扫描中进行推理,而不依赖于摄像头或激光雷达输入。Radar4D-VLM提取几何基础的物体提议,并将雷达证据组织成物体、场景和运动学标记的紧凑层次结构。一个参数高效的投影器将这些标记映射到冻结的语言骨干网络中,而可审计的预测头共同建模物体数量、空间分布、运动状态、碰撞风险、语义类别和径向速度。在统一的冻结骨干接口中,Radar4D-VLM结合了基于提议的时间性物体标记化、全局场景上下文和显式运动学标记。在序列隔离的K-Radar开发验证中,其Top-64提议召回率在4米处达到了98.13%,超过了固定格子和均匀随机控制,分别提高了6.40和22.83个百分点。我们进一步评估了在相同适应预算下跨越八个冻结的Qwen、Phi、Mistral、Llama和Gemma骨干的24次匹配运行。雷达标记接口在所有五个语言模型家族中保持兼容,而匹配对齐、置换和无语言控制显示出传感器依赖性,但没有从对齐语言监督中获得稳定的直接头增益。这些结果为基于雷达的多模态场景和运动推理建立了可重复的基础,同时将接口兼容性与语言监督的收益分开。
LiDAR/点云 / 2 / 2608.04568

Talk2Sensors: 3D Visual Grounding in Autonomous Driving via Sensor-Adaptive Physical Cue Matching

Talk2Sensors:通过传感器自适应物理线索匹配实现自主驾驶中的3D视觉定位
Runwei Guan, Di Tian, Ningwei Ouyang, Ruixiao Zhang, Shaofeng Liang, Haocheng Zhao, Lianqing Zheng, Xiaokai Bai, Guotao Wang, Daizong Liu, Henghui Ding, Hui Xiong
cs.CV
Abstract
As a key capability for embodied intelligence, 3D visual grounding (3DVG) has been predominantly studied in indoor scenes with RGB-D or point-cloud inputs, while existing outdoor extensions largely rely on monocular images alone. Both settings fall short of real-world outdoor perception, where heterogeneous sensors capture complementary yet distinct physical properties, such as visual texture, 3D geometry, and object kinematics, that are indispensable for flexible and robust query-adaptive grounding but remain under-exploited. To bridge this gap, we introduce Talk2Sensors, the first multi-sensor 3D visual grounding dataset built upon camera, LiDAR, and 4D radar. It contains 8,682 language instructions and 20,558 referred objects, with diverse prompts explicitly aligned with sensor-specific physical cues. Furthermore, we propose TSFormer, a unified Transformer-based framework for language-guided 3D visual grounding in autonomous driving. TSFormer adopts a coarse-to-fine property-aware fusion strategy: the Language-Routed Property Sampler first performs coarse text-conditioned feature retrieval by modulating sensor sampling weights with query-level linguistic cues, while the subsequent Sparse-Preserving Modality Arbiter module conducts fine-grained modality arbitration and text-guided refinement to determine the precise referred spatial location. This design enables dynamic routing of appearance, geometry, and motion cues according to the semantic requirements of each prompt, preventing dense modalities from overwhelming sparse but critical sensor signals. Extensive experiments demonstrate that TSFormer achieves state-of-the-art performance across multiple benchmarks: it improves over the strongest baseline by 8.05 mAP on Talk2Sensors, and transfers to the monocular Mono3DRefer benchmark with 53.05\% [email protected] .
Chinese Translation
作为具身智能的关键能力,3D视觉定位(3DVG)主要在室内场景中使用RGB-D或点云输入进行研究,而现有的户外扩展大多仅依赖单目图像。这两种设置都未能充分反映现实世界的户外感知,其中异构传感器捕获互补但不同的物理属性,如视觉纹理、3D几何形状和物体运动学,这些属性对于灵活和稳健的查询自适应定位至关重要,但尚未得到充分利用。为了解决这一问题,我们提出了Talk2Sensors,这是第一个基于相机、激光雷达和4D雷达构建的多传感器3D视觉定位数据集。该数据集包含8,682条语言指令和20,558个参考对象,具有与传感器特定物理线索明确对齐的多样化提示。此外,我们提出了TSFormer,一个基于Transformer的统一框架,用于自主驾驶中的语言引导3D视觉定位。TSFormer采用粗到细的属性感知融合策略:语言引导属性采样器首先通过调节传感器采样权重与查询级语言线索进行粗略的文本条件特征检索,而后续的稀疏保留模态仲裁模块则进行细粒度的模态仲裁和文本引导的精细化,以确定精确的参考空间位置。该设计使得外观、几何和运动线索能够根据每个提示的语义需求进行动态路由,防止密集模态淹没稀疏但关键的传感器信号。大量实验表明,TSFormer在多个基准测试中实现了最先进的性能:在Talk2Sensors上比最强基线提高了8.05 mAP,并在单目Mono3DRefer基准上转移到53.05\% [email protected] 。

预测/规划

1
预测/规划 / 1 / 2608.04453

TwinIR: Coordinated Invisible Dual-Point Attacks on Online HD Map Construction

TwinIR:在线高清地图构建中的协调隐形双点攻击
Haibo Hu, Jianghuai Deng, Chen Tang, Yang Lou, Qian Xu, Jianping Wang
cs.CV · cs.AI
Abstract
Online HD map construction is critical to prediction and planning in autonomous driving. We find that existing physical attacks against online map construction are limited by a cross-boundary compensation effect: after the target boundary is perturbed, another visible boundary may retain sufficient geometric cues for the model to recover the original road geometry. Based on this observation, we propose TwinIR, a new mechanism-guided physical attack methodology for online map construction. TwinIR jointly optimizes attack effectiveness and point sparsity, seeking the minimum number of attack points needed to suppress compensating geometric cues from surrounding boundaries. To reduce the perceptibility of multi-point attacks, TwinIR models camera responses to near-infrared illumination and maps optimized attack points to feasible physical placements, producing camera-visible interference with minimal visible-spectrum changes. Experiments on nuScenes across state-of-the-art online map construction models show that TwinIR reduces mAP by 8.18-8.96 percentage points under RSA and 2.84-5.62 points under ETA, while increasing the unreachable-goal rate by 25-28 points and the unsafe-planned-trajectory rate by 19-20 points over clean inputs. These attacks are also validated on a real-world testbed AV, where TwinIR successfully induces both road straightening and early-turn deformations while remaining inconspicuous in full-color views.
Chinese Translation
在线高清地图构建对自动驾驶中的预测和规划至关重要。我们发现,现有针对在线地图构建的物理攻击受到跨边界补偿效应的限制:在目标边界受到干扰后,另一个可见边界可能保留足够的几何线索,使模型能够恢复原始道路几何形状。基于这一观察,我们提出了TwinIR,一种新型的机制引导物理攻击方法,旨在在线地图构建中优化攻击效果和点稀疏性,寻求抑制周围边界补偿几何线索所需的最小攻击点数量。为了降低多点攻击的可感知性,TwinIR对近红外照明下的相机响应进行建模,并将优化的攻击点映射到可行的物理位置,产生在可见光谱变化最小的情况下可见的干扰。在nuScenes数据集上对最先进的在线地图构建模型进行的实验表明,TwinIR在RSA下将mAP降低了8.18-8.96个百分点,在ETA下降低了2.84-5.62个百分点,同时使不可达目标率提高了25-28个百分点,安全规划轨迹率提高了19-20个百分点,相较于干净输入。这些攻击还在真实世界测试平台的自动驾驶车辆上得到了验证,TwinIR成功引发了道路拉直和早转变形,同时在全彩视图中保持不显眼。

仿真/数据

2
仿真/数据 / 1 / 2608.04811

StaticSegFormer: An Efficient High-Performance Semantic Segmentation Based on Static Structured Pruning

StaticSegFormer:一种基于静态结构剪枝的高效高性能语义分割方法
Timo Bartels, Danish Nazir, Jan Piewek, Thorsten Bagdonat, Tim Fingscheidt
cs.CV
Abstract
Structured pruning enhances the efficiency of deep neural networks (DNNs) by eliminating groups of parameters during inference. Previous methods mostly reduce computational complexity (FLOPs), while semantic segmentation performance (mIoU) slightly drops. Accordingly, recent dynamic structured pruning methods aim at reducing the performance drop, while lowering the FLOPs even more. However, on the ADE20K and Cityscapes benchmarks, our study reveals that on a GPU platform such dynamic methods exhibit a surprisingly low frame rate far below a simple static approach, while having comparable results in mIoU and FLOPs. To address this issue, we propose a static structured pruning method for attention layers, that achieves both, a lower FLOPs and a high frame rate [fps] of the SegFormer network, the latter increased by up to 34% relative on the Cityscapes dataset, while having no mIoU performance drop at all. Our so-called StaticSegFormer method is strongest for small encoders and large images.
Chinese Translation
结构剪枝通过在推理过程中消除参数组来提高深度神经网络(DNN)的效率。以往的方法主要降低计算复杂度(FLOPs),而语义分割性能(mIoU)略有下降。因此,最近的动态结构剪枝方法旨在减少性能下降,同时进一步降低FLOPs。然而,在ADE20K和Cityscapes基准测试中,我们的研究揭示,在GPU平台上,这些动态方法的帧率远低于简单的静态方法,而在mIoU和FLOPs方面却表现相当。为了解决这个问题,我们提出了一种针对注意力层的静态结构剪枝方法,既实现了更低的FLOPs,又提高了SegFormer网络的帧率,在Cityscapes数据集上相对提高了多达34%,同时没有造成mIoU性能下降。我们所称的StaticSegFormer方法在小型编码器和大图像上表现最为出色。
仿真/数据 / 2 / 2608.04776

NSF-HRPT: Neural Semantic Field meets Hierarchical Risk Perception Tree for Safety-Critical Scenario Assessment

NSF-HRPT:神经语义场与层次风险感知树相结合的安全关键场景评估
Yu Zhao, Jiangyu Pan, Tao Hu, Ming Yin, Fan Yang, Jiangfan Liu, Xiubo Liang
cs.AI
Abstract
The ability to accurately assess and anticipate risks in safety-critical scenarios is crucial for autonomous driving systems. While existing research has made progress in collision prediction, accurately quantifying risk levels from monocular vision inputs remains challenging due to the complex dynamics of multi-agent interactions and the inherent uncertainty in real-world environments. To address these challenges, we present NSF-HRPT, a novel framework that combines learning-based perception with structured reasoning for quantitative risk assessment. Our approach features a Neural Semantic Field (NSF) that learns to model scene semantics, trajectory predictions, and probabilistic Time-to-Collision (TTC) distributions from simulation data. During inference, the pre-trained NSF serves as a prior for our Hierarchical Risk Perception Tree (HRPT), which enables efficient parallel computation and spatial reasoning about multi-agent risks. Additionally, we introduce a Sim2Real enhancement strategy that improves real-world applicability without retraining by incorporating priors from foundation models. Extensive evaluations demonstrate that our framework achieves state-of-the-art performance on synthetic benchmarks and delivers competitive, near-state-of-the-art results on real-world datasets for both TTC estimation accuracy and risk localization precision. The proposed method provides an effective solution for real-time risk awareness from monocular camera inputs.
Chinese Translation
在安全关键场景中准确评估和预测风险的能力对自主驾驶系统至关重要。尽管现有研究在碰撞预测方面取得了一定进展,但由于多智能体交互的复杂动态以及现实环境中固有的不确定性,从单目视觉输入中准确量化风险水平仍然具有挑战性。为了解决这些问题,我们提出了NSF-HRPT,一个结合基于学习的感知与结构化推理的定量风险评估新框架。我们的方法采用神经语义场(Neural Semantic Field, NSF),该模型从仿真数据中学习场景语义、轨迹预测和概率时间到碰撞(Time-to-Collision, TTC)分布。在推理过程中,预训练的NSF作为我们层次风险感知树(Hierarchical Risk Perception Tree, HRPT)的先验信息,使得多智能体风险的高效并行计算和空间推理成为可能。此外,我们引入了一种Sim2Real增强策略,通过结合基础模型的先验信息,提升了实际应用的可行性,而无需重新训练。大量评估表明,我们的框架在合成基准测试中实现了最先进的性能,并在真实世界数据集上对于TTC估计准确性和风险定位精度均提供了具有竞争力的接近最先进水平的结果。所提出的方法为从单目摄像头输入实现实时风险感知提供了有效的解决方案。