Embodied AI Glossary中文

Coordinate Transformation

坐标变换Essential

Converting a point's or pose's numeric value from one coordinate frame into another.

A coordinate transformation answers the question: given this point's coordinates in frame A, what are its coordinates in frame B? Two frames differ by a rotation plus a translation: p_B = R·p_A + t, where R is a 3×3 rotation matrix (the difference in orientation) and t is a translation vector (the offset between origins). In practice, R and t are usually packed into a single 4×4 homogeneous transformation matrix T, so translation also becomes a matrix multiplication, and chaining several transforms together is just multiplying matrices, as in T_base_obj = T_base_cam · T_cam_obj. This comes up everywhere in robotics: an object a camera sees has to be converted into the base frame before the arm can reach for it, and forward kinematics is nothing more than chaining transformation matrices link by link along the arm. ROS maintains and looks up these transforms through its TF library.

ExampleHand-eye calibration gives the camera's pose in the base frame, T_base_cam; a vision model gives the cup's position in the camera frame. Multiplying the two together gives the cup's position in the base frame, which the arm uses to plan its grasp.

Also called
Frame Transformation, Rigid Transformation, Pose Transformation
Related
Homogeneous Transformation Matrix · Rotation Matrix · TF / tf2 Transform Tree · Hand-Eye Calibration · Camera Extrinsics · Forward Kinematics (FK)
Sources
Wikipedia: Transformation matrix(Affine transformations / homogeneous coordinates)
ROS REP 105: Coordinate Frames for Mobile Platforms
Modern Robotics(Lynch & Park)Ch.3 Rigid-Body Motions

See it in the full glossary →