Multisensory Pretraining for Contact-Rich Manipulation

Pretraining a multimodal encoder on RGB, depth, segmentation and force-torque data to improve robot manipulation in contact-rich tasks.

2023–2025 · CSCI 5980 Deep Learning for Robotics, Robotics Perception and Manipulation Lab, University of Minnesota · advised by Karthik Desingh · created with Ryan Diaz, Hanchen Cui and Shirley Su

Final project presentation covering pretraining, masked reconstruction results and the three evaluation tasks.

A research project on pretraining a multimodal encoder for contact-rich robot manipulation. Vision alone is a poor guide for tasks defined by contact. The encoder fuses RGB images, depth maps and gripper force-torque readings into a single pretrained representation, and a policy is fine-tuned on top of it. Masked reconstruction across all three modalities during pretraining teaches the encoder both visual and contact information. The work uses PyTorch, synthetic scenes rendered in Blender, and a UR5e arm. On peg-in-hole the RGBD+F/T pretrained policy completed 10/20 rollouts against 7/20 for RGBD-only and 0/20 from scratch.

My role

Built the synthetic data generation pipeline and the multimodal masked autoencoder. Ran the peg-in-hole, hammer-cleanup and coffee-preparation evaluations.

Media

The four input streams, mapped to an end-effector pose and gripper action. They are RGB, depth, semantic segmentation, and force-torque with proprioception.
A rollout with the force-torque traces plotted alongside the visual modalities.
Proposed architecture from the project proposal.
Proposed architecture from the project proposal.
Blender render of the arm used to generate synthetic training scenes.
Synthetic tabletop scene with paired RGB, depth and segmentation.
Synthetic scene sweep over primitive shapes.
Single-object capture used for object-level supervision.
Bin-picking scene from the generated dataset.