Embodied AI Glossary中文

Camera Extrinsics

相机外参Essential

The rotation-plus-translation parameters describing where a camera sits in space and which way it points.

Camera extrinsics consist of a 3×3 rotation matrix R and a translation vector t, often combined into a 4×4 homogeneous transformation matrix, describing where the camera sits and which way it points relative to a world frame or a robot's base frame. OpenCV's convention transforms a point in the world frame into the camera frame; many robotics codebases instead store the opposite direction — a ‘camera pose’ that goes from camera to world — so it's worth confirming the direction before reusing someone else's data. Extrinsics are typically found using a calibration board together with PnP, or through hand-eye calibration, and need to be redone if the camera gets bumped. Once known, they let point clouds from multiple cameras be merged into one common frame, or let an object's position as seen by the camera be converted into coordinates a robot arm can act on.

ExampleThe DROID dataset was collected with two external ZED 2 cameras plus one wrist-mounted ZED Mini, covering 1,417 camera viewpoints in total, and ships each viewpoint's intrinsic and extrinsic calibration alongside the data.

Also called
Extrinsic Parameters, Camera Pose, Extrinsic Matrix
Related
Camera Intrinsics · Hand-Eye Calibration · Homogeneous Transformation Matrix · Coordinate Transformation · Perspective-n-Point · Camera Calibration
Sources
OpenCV calib3d.hpp:针孔相机模型与外参 R、t 的定义 (Chinese)
Wikipedia: Camera resectioning
DROID: A Large-Scale In-the-Wild Robot Manipulation Dataset

See it in the full glossary →