Preloader

Teaching radar to read motion in traffic scenes

0

Teaching radar to read motion in traffic scenes

Lisa Lock

Scientific Editor

Andrew Zinin

Chief Editor

Teaching radar to read motion in traffic scenes
Comparison between existing self-supervised (SSF) and cross-modal supervised (CMS) radar scene flow estimation settings and our weakly supervised cross-modal learning setting. SF, FDS, and EM denote the predicted scene flow, foreground dynamic segmentation, and ego-motion, respectively. Lself is the self-supervised losses; Lopt, Lmot, Lseg, and Lego are cross-modal losses, with supervision from 2D optical flow, 3D LiDAR-based pseudo scene flow label and FDS ground-truth, and odometry-based ego-motion. Lic and Lis are our instance-aware losses, and Lstat is the rigid static loss. Credit: SUTD

Autonomous vehicles and robots need to understand not only what is around them, but how objects and people are moving. A cyclist crossing the road, a car slowing down, a pedestrian stepping off the curb, and a parked vehicle all create different motion cues that a machine must interpret quickly and accurately.

A research team led by assistant professor Zhao Na from the Singapore University of Technology and Design (SUTD) has developed IterFlow, a lightweight learning framework that helps 4D radar estimate the 3D motion of points in a traffic scene. Their study, posted to the arXiv preprint server, addresses a key challenge in autonomous perception: how to make radar-based motion understanding more accurate without depending on expensive LiDAR-based supervision.

Refining motion estimates from sparse radar

4D radar is attracting growing interest because it is more compact, more cost-effective and more robust in adverse environmental conditions than LiDAR, which is short for light detection and ranging. However, radar point clouds are also sparse and noisy, making it difficult for AI systems to estimate scene flow—the 3D motion of points between consecutive sensor frames.

IterFlow was developed to tackle this problem through a more focused design. The research team designed a task-specific network with a concise training strategy to improve radar scene flow performance. Rather than relying on increasingly complex multitask systems, IterFlow refines motion estimates step by step and uses targeted training signals to reduce errors in sparse radar data.

A central feature of IterFlow is that it does not require LiDAR-based pseudo scene flow labels during training. Instead, it uses RGB images and odometry—information about the vehicle’s own movement—as auxiliary supervision. At test time, the system needs only radar point clouds as input.

“IterFlow shows that better radar scene flow estimation does not have to depend on increasingly complex models or costly LiDAR supervision. By using images and odometry during training, we can make radar-based motion understanding lighter, more cost-effective and more applicable to real-world autonomous systems,” Zhao said.

Keeping moving objects distinct

In everyday terms, scene flow estimation helps a system work out how each point captured by a sensor is moving from one moment to the next. In a road scene, this could mean estimating the motion of points belonging to moving cars, cyclists or pedestrians while distinguishing them from static background points such as parked vehicles or roadside structures.

The team’s method uses camera images to provide object-level guidance. Through 2D tracking and segmentation, the system identifies object instances in images and projects this information into 3D radar space. This helps reduce a common source of error in radar scene flow learning, in which a model may mismatch moving foreground points with static background points.

For example, if radar points are sparse, a moving cyclist and nearby static background points may appear close together in the data. Existing methods that rely mainly on spatial distance may wrongly encourage these points to move in similar ways. IterFlow’s instance-aware losses reduce this problem by applying motion consistency within the same object instance, rather than across points that are merely nearby.

IterFlow also uses a ball query-based grouping method that is better suited to sparse radar data. Unlike k-nearest-neighbor methods, which always return a fixed number of neighbors even if some are far away, ball query first checks whether points fall within a defined spatial radius. This helps avoid false correspondences in sparse radar regions and improves robustness.

Efficiency gains and remaining limits

Experiments on the real-world View-of-Delft dataset showed that IterFlow outperformed the previous radar-based cross-modal scene flow method CMFlow while using only three losses, about 40 times fewer parameters and about 30 times fewer giga floating-point operations, or GFLOPs, a measure of computational cost. The results suggest that radar scene flow estimation can be improved without adding costly sensors or substantially increasing model complexity.

For now, this research remains at the experimental stage. The current method uses a PointNet++ point cloud feature extraction network, which supports only a fixed input point cloud size. Future work will focus on overcoming this limitation before the approach can be more broadly tested in practical vehicle or robotic systems.

By reducing reliance on costly LiDAR supervision and improving how radar learns motion from sparse data, IterFlow points toward more efficient radar-based perception for autonomous systems. Its broader significance lies not in replacing other sensors immediately, but in showing how lower-cost sensing can be made more capable through carefully designed machine learning.

Publication details

Jingyun Fu et al, Weakly Supervised Cross-Modal Learning for 4D Radar Scene Flow Estimation, arXiv (2026). DOI: 10.48550/arxiv.2605.18507

Journal information:
arXiv

Key concepts

Computational 3D vision

Who’s behind this story?

Lisa Lock

Lisa Lock

BA art history, MA material culture. Former museum editor, paramedic, and transplant coordinator. Editing for Science X since 2021.

Full profile →


Andrew Zinin

Andrew Zinin

Master’s in physics with research experience. Long-time science news enthusiast. Plays key role in Science X’s editorial success.

Full profile →

Citation:
Teaching radar to read motion in traffic scenes (2026, September 24)
retrieved 24 September 2026
from https://techxplore.com/news/2026-09-radar-motion-traffic-scenes.html
This document is subject to copyright. Apart from any fair dealing for the purpose of private study or research, no
part may be reproduced without the written permission. The content is provided for information purposes only.

Source: Tech Xplore

Choose your Reaction!
Leave a Comment