Embodied AI Glossary中文

Rhoda AI

Advanced

A US embodied-AI startup that pre-trains robot policies on internet video, aimed at industrial settings.

Rhoda AI is a US embodied-AI startup building robot foundation models for industrial sites. Its website describes a core capability called FutureVision and a “Direct Video Action” model: it first pre-trains on large-scale internet video so the model learns to predict how a scene will change, then fine-tunes on just 1 to 10 hours of robot trajectory data to get a policy that outputs actions directly — which the company says generalizes better than a conventional vision-language-action (VLA) pipeline. It also builds its own wheeled robot platform, with demonstrations of tasks such as returns processing, bearing kitting, and unboxing across automotive, manufacturing, logistics, and e-commerce settings. Investors listed on its website include Khosla Ventures, John Doerr, Temasek, and Samsung NEXT; the founding team, founding year, and funding amount are not disclosed on the site and are omitted here.

Related
Video Generation Model · World Action Model · Vision-Language-Action Model · Internet Video Data · Post-training
Sources
Rhoda AI 官网 (Chinese)
As of
2026-09

See it in the full glossary →