Definition

A model parameter is a numerical value in a model that helps determine its computations. Training usually learns its value. A newly initialized weight is already a parameter, even before its first update. In a neural network, weights and biases in layers, along with numbers in embedding tables, are parameters. Together with the model’s architecture, they govern its computations for a given input.

Training updates trainable parameters to reduce loss. At inference, a request uses the stored values without updating them. Deploying a checkpoint with different weights or activating an adapter changes the parameters used at inference.

Parameters are distinct from hyperparameters such as the learning rate, which is chosen to control training. They are also distinct from request settings such as decoding temperature, which an application can change when it calls a language model.

Simple example

A language model has an embedding table that maps token IDs to vectors. Training adjusts the numbers in those vectors. Those numbers are model parameters.

The training job might use a learning rate of 0.001. The trainer chooses that value to control how far each update moves the parameters. Later, an application might request a response with temperature 0.2. Changing temperature can change the generated text, but it leaves the embedding table as it was.

Why it matters

A prompt or decoding setting affects a request. Fine-tuning updates learned values that subsequent requests can use once the resulting model or adapter is deployed. A model rollout can change behavior even when the application code stays the same.

When comparing model versions, record the checkpoint or adapter alongside the request settings. Otherwise, a behavior change can be hard to trace to its source.

One important nuance

A parameter does not have to be trainable in the current run. Fine-tuning can freeze the base model while training an adapter. The frozen weights remain model parameters, and the adapter introduces its own learned parameters. The number of parameters updated during a training run can therefore be smaller than the model’s total parameter count.