Multisensory Pretraining for Contact-Rich Manipulation
Pretraining a multimodal encoder on RGB, depth, segmentation and force-torque data to improve robot manipulation in contact-rich tasks.
2023–2025 · CSCI 5980 Deep Learning for Robotics, Robotics Perception and Manipulation Lab, University of Minnesota · advised by Karthik Desingh · created with Ryan Diaz, Hanchen Cui and Shirley Su
A research project on pretraining a multimodal encoder for contact-rich robot manipulation. Vision alone is a poor guide for tasks defined by contact. The encoder fuses RGB images, depth maps and gripper force-torque readings into a single pretrained representation, and a policy is fine-tuned on top of it. Masked reconstruction across all three modalities during pretraining teaches the encoder both visual and contact information. The work uses PyTorch, synthetic scenes rendered in Blender, and a UR5e arm. On peg-in-hole the RGBD+F/T pretrained policy completed 10/20 rollouts against 7/20 for RGBD-only and 0/20 from scratch.
My role
Built the synthetic data generation pipeline and the multimodal masked autoencoder. Ran the peg-in-hole, hammer-cleanup and coffee-preparation evaluations.
Media