Definition
Model size usually means the total number of parameters in a model. These are numerical values, such as weights and biases, that training learns. A label such as “7B” typically means roughly seven billion parameters. It doesn’t describe how many tokens the model was trained on or how much text fits in the context window.
Size can also mean the storage the model’s weight files take up or the memory needed to run it. In those cases, the measures depend on how the parameters are stored and how the model runs. They are not interchangeable with the parameter count.
Sparse mixture-of-experts models distinguish total parameters from active parameters used for each token. The active count helps describe computation, while the total count still matters for storing the weights.
Simple example
Suppose a model has exactly seven billion parameters. Storing each in a 16-bit format requires about 14 GB for the parameter values alone: seven billion parameters multiplied by two bytes.
Packing those values into four bits each gives a raw estimate of 3.5 GB. Quantized formats commonly add scales and other encoding data, and some tensors may retain higher precision. The model still has seven billion parameters.
Neither estimate is a complete inference memory budget. Running the model also needs working memory. Typical autoregressive transformers also use a KV cache to store key and value representations for previously processed tokens in active sequences.
Why it matters
The number of parameters and storage precision help estimate whether a model fits on your hardware. For comparable dense architectures, more parameters generally mean more per-token computation, which affects deployment planning.
Use size to shortlist deployment candidates. Then measure memory and latency with realistic input lengths and concurrency. Hardware, batching, and the serving runtime affect results, so parameter count alone cannot predict request speed or deployment fit.
One important nuance
A larger model can perform better under comparable training conditions, but parameter count alone cannot determine quality. Training data quality and quantity also matter, along with architecture and post-training. A smaller model can outperform a larger one on a particular task.
Choose between candidates using representative task evaluations alongside resource measurements. An increase from 7B to 14B describes twice as many parameters, but it does not promise twice the accuracy.