Embodied AI Glossary中文

Skild S1

S1Advanced

Skild AI's robot foundation model that can perform a task it never trained on after watching a single demonstration video.

Skild S1 was released in August 2026 by American robotics company Skild AI, described by the company as its flagship robot foundation model. Its core idea is 'in-context learning': at deployment, the robot is shown a demonstration video, and the model infers the demonstrator's intent, the correspondence between objects, and task progress from it, then executes the task directly without updating any weights, even if that exact task never appeared during pretraining. Traditionally, a VLA facing a new task would need fresh teleoperated data and fine-tuning; S1 aims to skip that step. Its pretraining data mixes robot teleoperation, UMI handheld data collection, first-person human video, and simulation data. According to the company's blog, at a pretraining scale of 100,000 hours, it reaches 96% success on tasks seen during training and 66% on unseen tasks (versus 9% for a language-instruction baseline); one demonstration is said to be worth roughly 380 post-training data points, and it can handle long-horizon tasks lasting up to about 10 minutes. S1 is already used by commercial customers and is offered externally through early access.

ExampleA staff member demonstrates repotting a plant once, and after watching the video, S1 goes on to complete this never-trained-on task; the company states that from setting up the scene to autonomous execution took just 11 minutes.

Also called
Skild AI S1
Related
In-Context Learning · Skild AI · Skild Brain · Human Video Data · Universal Manipulation Interface · Long-horizon Task
Sources
Introducing S1: In-Context Learning for Robotics (Skild AI blog)
Skild AI blog index
As of
2026-08

See it in the full glossary →