Embodied AI Glossary中文

06Perception & Sensors

How robots see and feel: cameras, depth, IMUs, force and touch, plus point clouds, calibration, and SLAM. · 307 terms

  1. 6.1Perception overview and cameras27
  2. 6.2Depth cameras and lidar40
  3. 6.3Proprioception and force sensing14
  4. 6.4Tactile sensing30
  5. 6.5Calibration and spatiotemporal alignment13
  6. 6.6Detection, segmentation, and tracking35
  7. 6.7Recovering 3D from images34
  8. 6.8Point clouds and 3D representations24
  9. 6.9Object pose, grasping, and affordance17
  10. 6.10Human body, hands, and interaction14
  11. 6.11State estimation and SLAM36
  12. 6.12Maps, semantics, and spatial intelligence23

6.1Perception overview and cameras

Perception splits into internal and external sensing; start with the everyday RGB camera — placement, how it images, and its specs.

6.2Depth cameras and lidar

Adding distance on top of ordinary imaging: stereo, structured-light, and ToF depth cameras, plus lidar.

6.3Proprioception and force sensing

From sensing the outside world to sensing itself: encoders and IMUs measure motion, force/torque sensors measure force and contact.

6.4Tactile sensing

A finer-grained sense of touch than force sensors: tactile arrays, e-skin, vision-based tactile sensors, and what tactile data is used for.

6.5Calibration and spatiotemporal alignment

Now that you know the sensors, align them: camera and hand-eye calibration, multi-sensor extrinsics, and time synchronization.

6.6Detection, segmentation, and tracking

Entering vision algorithms: boxing, cutting out, and tracking objects in images, including targets specified by text.

6.7Recovering 3D from images

Without a depth sensor: estimating depth from ordinary images, multi-view geometry, and networks that output 3D in one step.

6.8Point clouds and 3D representations

Once you have 3D data: downsampling, registering, and running networks on point clouds, then meshes, NeRF, and Gaussian splatting.

6.9Object pose, grasping, and affordance

Putting 3D perception to work in manipulation: estimating 6D object pose, detecting grasps, and finding affordances and articulated structure.

6.10Human body, hands, and interaction

Shifting from objects to people: body and hand pose, 3D meshes, plus gesture, gaze, and speech.

6.11State estimation and SLAM

Answering ‘where am I’: from odometry and multi-sensor fusion to SLAM and localizing within a known map.

6.12Maps, semantics, and spatial intelligence

Beyond localization, understanding the whole scene: geometric and semantic maps, scene graphs, and LLM-based spatial reasoning and active perception.

See it in the full glossary →