Embodied AI Glossary中文

Vision-Only Approach

纯视觉方案Common

Perceives the environment using only cameras, without lidar, millimeter-wave radar, or other ranging sensors.

A vision-only approach means perception relies solely on cameras — usually ordinary RGB cameras — with distance and 3D structure both inferred from images by an algorithm, without lidar, millimeter-wave radar, ultrasonic sensors, or other ranging hardware. The best-known example is Tesla: starting in May 2021, North American–built Model 3/Y dropped millimeter-wave radar under the name “Tesla Vision,” and in October 2022 Tesla announced it was removing ultrasonic sensors too. The advantages are cheap hardware and data that’s easy to scale; the cost is that depth can only be estimated, so errors are more likely in low light, backlight, and with transparent or reflective objects. In embodied AI, “vision-only” also describes a policy that looks only at images, without tactile or force input — for example, OpenVLA takes just a single RGB image and a language instruction. Whether a vision-only approach or multi-sensor fusion is better remains debated.

ExampleOpenVLA receives only a single RGB image and one instruction, with no depth, force, or proprioceptive input, and outputs the robot arm’s end-effector action directly.

Also called
Camera-Only
Related
RGB Camera · LiDAR · Depth Camera · Multi-Sensor Fusion · Visuo-Tactile Fusion · Autonomous Driving
Sources
Wikipedia: Tesla Autopilot hardware(Tesla Vision:2021 年去毫米波雷达、2022 年去超声波雷达) (Chinese)
OpenVLA: An Open-Source Vision-Language-Action Model(局限:仅支持单图输入) (Chinese)
As of
2022-10

See it in the full glossary →