Embodied AI Glossary中文

Room-to-Room

R2R / VLN-CE 视觉语言导航基准R2R / VLN-CECommon

The most widely used vision-language navigation benchmark: follow a human-written route description to reach a destination in an unfamiliar house.

R2R (Room-to-Room) was released in 2018 by teams from the Australian National University, the University of Adelaide, and others, alongside the Matterport3D simulator. It is built on scans of 90 real buildings and contains about 22,000 human-written route instructions, with the agent restricted to jumping between nodes of a pre-built navigation graph. In 2020, Oregon State University, Georgia Tech, and Facebook AI proposed VLN-CE, which moved R2R into Habitat's continuous environments: the agent instead has to walk using low-level actions such as “move forward 0.25 meters,” “turn left 15 degrees,” and “stop,” which is much closer to a real robot and causes scores to drop noticeably. Google's RxR (2020) contains 126,000 instructions in English, Hindi, and Telugu. Common metrics include success rate, SPL (success weighted by path length), and nDTW (normalized dynamic time warping).

ExampleThe official VLN-CE cross-modal attention baseline scores 0.28 success rate and 0.25 SPL on the R2R continuous-environment test set, showing that the continuous-environment version is far harder than the navigation-graph version.

Also called
R2R, VLN-CE, RxR, R2R-CE, RxR-CE, VLN in Continuous Environments
Related
Vision-and-Language Navigation · Habitat · Matterport3D · Success weighted by Path Length · normalized Dynamic Time Warping · NaVILA
Sources
Room-to-Room (R2R) 官网 (Chinese)
VLN-CE 项目页 (Chinese)
jacobkrantz/VLN-CE GitHub(含 RxR-Habitat) (Chinese)

See it in the full glossary →