Definition
A hyperparameter is a choice that controls a model’s design or training process. The training run does not learn it in the same way it learns model parameters such as weights. Learning rate and batch size are training hyperparameters. The number of layers is an architectural hyperparameter.
These choices can affect which parameter values training produces, even though they are not themselves those learned values.
Simple example
Suppose you train a classifier with a batch size of 32 and an initial learning rate of 0.001. You choose those settings for this run. The optimizer updates the model’s weights.
If you repeat the training run with a different learning rate, you have changed a hyperparameter. The resulting weights and validation results may change too.
Why it matters
Hyperparameters affect training cost and behavior. A larger batch can require more memory. A learning rate that is too high can make optimization unstable. One that is too low can make progress slow.
When comparing training runs, record the settings alongside the data split and resulting checkpoint. A validation score alone cannot tell you which training choice produced it.
One important nuance
“Chosen” does not mean “fixed forever” or “picked by hand.” A search process can select a learning rate using validation results. A schedule can also change it during training. In both cases, the learning rate remains a training setting rather than a weight learned from examples by the ordinary parameter update step.