Teaching radar to read motion in traffic scenes
Lisa Lock
Scientific Editor
Andrew Zinin
Chief Editor

Autonomous vehicles and robots need to understand not only what is around them, but how objects and people are moving. A cyclist crossing the road, a car slowing down, a pedestrian stepping off the curb, and a parked vehicle all create different motion cues that a machine must interpret quickly and accurately.
A research team led by assistant professor Zhao Na from the Singapore University of Technology and Design (SUTD) has developed IterFlow, a lightweight learning framework that helps 4D radar estimate the 3D motion of points in a traffic scene. Their study, posted to the arXiv preprint server, addresses a key challenge in autonomous perception: how to make radar-based motion understanding more accurate without depending on expensive LiDAR-based supervision.
Refining motion estimates from sparse radar
4D radar is attracting growing interest because it is more compact, more cost-effective and more robust in adverse environmental conditions than LiDAR, which is short for light detection and ranging. However, radar point clouds are also sparse and noisy, making it difficult for AI systems to estimate scene flow—the 3D motion of points between consecutive sensor frames.
IterFlow was developed to tackle this problem through a more focused design. The research team designed a task-specific network with a concise training strategy to improve radar scene flow performance. Rather than relying on increasingly complex multitask systems, IterFlow refines motion estimates step by step and uses targeted training signals to reduce errors in sparse radar data.
A central feature of IterFlow is that it does not require LiDAR-based pseudo scene flow labels during training. Instead, it uses RGB images and odometry—information about the vehicle’s own movement—as auxiliary supervision. At test time, the system needs only radar point clouds as input.
“IterFlow shows that better radar scene flow estimation does not have to depend on increasingly complex models or costly LiDAR supervision. By using images and odometry during training, we can make radar-based motion understanding lighter, more cost-effective and more applicable to real-world autonomous systems,” Zhao said.
Keeping moving objects distinct
In everyday terms, scene flow estimation helps a system work out how each point captured by a sensor is moving from one moment to the next. In a road scene, this could mean estimating the motion of points belonging to moving cars, cyclists or pedestrians while distinguishing them from static background points such as parked vehicles or roadside structures.
The team’s method uses camera images to provide object-level guidance. Through 2D tracking and segmentation, the system identifies object instances in images and projects this information into 3D radar space. This helps reduce a common source of error in radar scene flow learning, in which a model may mismatch moving foreground points with static background points.
For example, if radar points are sparse, a moving cyclist and nearby static background points may appear close together in the data. Existing methods that rely mainly on spatial distance may wrongly encourage these points to move in similar ways. IterFlow’s instance-aware losses reduce this problem by applying motion consistency within the same object instance, rather than across points that are merely nearby.
IterFlow also uses a ball query-based grouping method that is better suited to sparse radar data. Unlike k-nearest-neighbor methods, which always return a fixed number of neighbors even if some are far away, ball query first checks whether points fall within a defined spatial radius. This helps avoid false correspondences in sparse radar regions and improves robustness.
Efficiency gains and remaining limits
Experiments on the real-world View-of-Delft dataset showed that IterFlow outperformed the previous radar-based cross-modal scene flow method CMFlow while using only three losses, about 40 times fewer parameters and about 30 times fewer giga floating-point operations, or GFLOPs, a measure of computational cost. The results suggest that radar scene flow estimation can be improved without adding costly sensors or substantially increasing model complexity.
For now, this research remains at the experimental stage. The current method uses a PointNet++ point cloud feature extraction network, which supports only a fixed input point cloud size. Future work will focus on overcoming this limitation before the approach can be more broadly tested in practical vehicle or robotic systems.
By reducing reliance on costly LiDAR supervision and improving how radar learns motion from sparse data, IterFlow points toward more efficient radar-based perception for autonomous systems. Its broader significance lies not in replacing other sensors immediately, but in showing how lower-cost sensing can be made more capable through carefully designed machine learning.
Publication details
Jingyun Fu et al, Weakly Supervised Cross-Modal Learning for 4D Radar Scene Flow Estimation, arXiv (2026). DOI: 10.48550/arxiv.2605.18507
Journal information:
arXiv
Key concepts
Teaching radar to read motion in traffic scenes (2026, September 24)
retrieved 24 September 2026
from https://techxplore.com/news/2026-09-radar-motion-traffic-scenes.html
part may be reproduced without the written permission. The content is provided for information purposes only.
Source: Tech Xplore


You must be logged in to post a comment.