Compositional Motion Generation From Demonstration With Object-Centric Neural Fields
Artikel i vetenskaplig tidskrift, 2026

Compositionality, by organizing complex behavior as combinations of simpler elements, enables robot learning that is scalable and data efficient. Leveraging this principle, we propose a generative learning-from-demonstration framework that enables compositional modeling of robotic behavior by connecting perception and motion through shared object-level representations. We render scenes from object-centric neural representations that integrate canonical neural fields with latent-conditioned deformations, capturing positional and geometric variations in a smooth, consistent, and interpretable way. For motion generation, a temporal mixture-of-experts (MoE) employs a gating mechanism to combine object-conditioned movement primitives over time, producing complete trajectories. This spatial–temporal compositionality maintains the data efficiency of movement primitives while grounding motion in visual structure, enabling systematic generalization across diverse scene configurations. In simulation, long-horizon manipulation tasks are successfully completed using the proposed model, which requires significantly less training data than other image-based baselines. Real-world experiments further demonstrate the method's robustness to noise, its ability to generalize at the category level through language-based segmentation models, and its capacity to operate directly on 3D scene representations.

Deep Learning in Grasping and Manipulation

Learning from Demonstration

Författare

Ahmet Ercan Tekden

Chalmers, Elektroteknik, System- och reglerteknik

Yasemin Bekiroglu

Chalmers, Elektroteknik, System- och reglerteknik

University College London (UCL)

IEEE Robotics and Automation Letters

23773766 (eISSN)

Vol. 11 9

Ämneskategorier (SSIF 2025)

Robotik och automation

Datorgrafik och datorseende

Datavetenskap (datalogi)

DOI

10.1109/LRA.2026.3713713

Mer information

Senast uppdaterat

2026-07-28