Hyperparameter
超参数EssentialA training setting chosen by a person beforehand rather than learned by the model itself, like the learning rate or batch size.
The numbers inside a model are called parameters (or weights), and training learns them automatically. Hyperparameters, by contrast, are settings chosen by a person before training starts and usually left unchanged during it — they determine how the learning happens. Common ones include the learning rate, batch size, number of training steps or epochs, the network's depth and width, and the strength of regularization; reinforcement learning adds things like the discount factor and PPO's clipping coefficient, while embodied models add things like the action-chunk length or the number of denoising steps in a diffusion model. Poorly chosen hyperparameters can keep a model from converging, cause overfitting, or just leave performance well below what's possible. The process of searching for a good combination is called hyperparameter tuning, commonly done with grid search, random search, or Bayesian optimization; in practice, people also often just start from a paper's default values and adjust a little — informally called “tuning knobs,” or in Chinese ML slang, “alchemy” (炼丹).
ExampleWhen training ACT, the action-chunk length k, the learning rate, and β (the weight on the KL term in the CVAE loss) are all hyperparameters that have to be set by hand.
- Related
- Learning Rate · Batch Size · Epoch · Overfitting · Alchemy (Deep-Learning Slang) · Population-Based Training
- Sources
- Google Machine Learning Glossary: hyperparameter
Wikipedia: Hyperparameter (machine learning)