Room-to-Room
R2R / VLN-CE 视觉语言导航基准R2R / VLN-CECommonThe most widely used vision-language navigation benchmark: follow a human-written route description to reach a destination in an unfamiliar house.
R2R (Room-to-Room) was released in 2018 by teams from the Australian National University, the University of Adelaide, and others, alongside the Matterport3D simulator. It is built on scans of 90 real buildings and contains about 22,000 human-written route instructions, with the agent restricted to jumping between nodes of a pre-built navigation graph. In 2020, Oregon State University, Georgia Tech, and Facebook AI proposed VLN-CE, which moved R2R into Habitat's continuous environments: the agent instead has to walk using low-level actions such as “move forward 0.25 meters,” “turn left 15 degrees,” and “stop,” which is much closer to a real robot and causes scores to drop noticeably. Google's RxR (2020) contains 126,000 instructions in English, Hindi, and Telugu. Common metrics include success rate, SPL (success weighted by path length), and nDTW (normalized dynamic time warping).
ExampleThe official VLN-CE cross-modal attention baseline scores 0.28 success rate and 0.25 SPL on the R2R continuous-environment test set, showing that the continuous-environment version is far harder than the navigation-graph version.
- Also called
- R2R, VLN-CE, RxR, R2R-CE, RxR-CE, VLN in Continuous Environments
- Related
- Vision-and-Language Navigation · Habitat · Matterport3D · Success weighted by Path Length · normalized Dynamic Time Warping · NaVILA
- Sources
- Room-to-Room (R2R) 官网 (Chinese)
VLN-CE 项目页 (Chinese)
jacobkrantz/VLN-CE GitHub(含 RxR-Habitat) (Chinese)