Multimodal Perception
多模态感知CommonUnderstanding a scene and task using vision together with touch, force, sound, and other senses at once.
Multimodal perception means a robot understands a scene and its task progress using not just vision but simultaneously touch, force and torque, sound, proprioception (sensing of its own state, such as joint angle and motor current), and even language instructions. Different modalities are good at different things: a 2022 paper, See, Hear, and Feel, summarizes it as vision seeing the global picture but often getting occluded, sound catching key moments that aren't visible, and touch providing precise local geometry. The difficulty is that modalities differ hugely in frequency, dimensionality, and noise characteristics, requiring a separate encoder for each one before they can be combined through concatenation, attention, or similar mechanisms, and data collection is also more involved. Multiple studies show that contact-rich manipulation tasks like peg insertion and pouring are more stable when touch and force are added rather than relying on vision alone.
ExampleSee, Hear, and Feel (CoRL 2022) has an arm use a camera, a contact microphone, and a vision-based tactile sensor together to do dense packing and pouring, fusing the three signals with self-attention, and outperforming setups that use only one or two modalities.
- Also called
- Multi-Sensory Perception
- Related
- Multi-Sensor Fusion · Visuo-Tactile Fusion · Proprioception · Tactile Sensor · Contact-rich Manipulation · Vision-Tactile-Language-Action Model
- Sources
- See, Hear, and Feel: Smart Sensory Fusion for Robotic Manipulation (CoRL 2022)
Making Sense of Vision and Touch: Self-Supervised Learning of Multimodal Representations for Contact-Rich Tasks (ICRA 2019)