Scaling Law
缩放定律EssentialThe empirical pattern that model performance improves smoothly, as a power law, with more parameters, data, and compute.
Scaling laws quantify the intuition that bigger models, more data, and more compute produce better results. In 2020, OpenAI's Kaplan and colleagues found that a language model's test loss follows power-law relationships with parameter count, dataset size, and training compute, with some of these trends holding across seven or more orders of magnitude. In 2022, DeepMind's Chinchilla study refined this, showing that for a fixed compute budget, model size and the number of training tokens should be scaled up in roughly equal proportion. Scaling laws matter because they let researchers run small, cheap experiments and extrapolate the payoff of a much larger training run, which guides where to spend a limited budget. The embodied-AI field is now testing whether robot data obeys a similar law, and this is part of the motivation behind the industry's large-scale data-collection efforts.
ExampleTsinghua's Yang Gao lab published “Data Scaling Laws in Imitation Learning for Robotic Manipulation” (2024), collecting over 40,000 demonstrations and running more than 15,000 real-robot trials. They found that a policy's ability to generalize to new environments and objects scales roughly as a power law with the number of training environments and objects, while adding more demonstrations within a single environment gives diminishing returns past a certain point.
- Also called
- Neural Scaling Law, Robot Scaling Law, Data Scaling Law
- Related
- Data Scaling Laws in Imitation Learning (Robotic Manipulation) · Parameter Count (Model Size) · Large Language Model · Data Diversity · The Bitter Lesson · Emergent Abilities
- Sources
- Scaling Laws for Neural Language Models (arXiv 2001.08361)
Training Compute-Optimal Large Language Models (arXiv 2203.15556)
Data Scaling Laws in Imitation Learning for Robotic Manipulation (arXiv 2410.18647) - As of
- 2024-10