Definition

A vocabulary is a model’s finite inventory of token entries, each assigned a numeric ID. Entries can be word fragments, byte sequences, or special tokens such as an end-of-sequence marker. The tokenizer converts text into IDs from this inventory. During text generation, a language model scores candidate IDs from its output vocabulary.

The vocabulary defines which IDs exist. It does not, by itself, define where the tokenizer splits text or how it normalizes input. Those rules and the ID mapping must match the model checkpoint.

Simple example

Imagine a tiny vocabulary with cat at ID 0, s at ID 1, " sat" at ID 2, and an end marker at ID 3. A tokenizer using it might encode cats sat as [0, 1, 2]. The model receives those numbers, not the spelling of the entries.

Why it matters

The model learns parameters associated with token IDs during training. If an application swaps in a tokenizer with a different ID mapping, ID 0 might now mean something other than cat. The request can still run while the model receives the wrong input. Generated IDs may decode to the wrong text as well.

Keep the tokenizer files with the model checkpoint when deploying it. Record the tokenizer version if you store token IDs or token counts for later use.

One important nuance

“Fixed” applies to a particular tokenizer and model version. Extending a vocabulary may require more input embedding rows. If the model should generate the new tokens, its output layer must also be able to score them. Any new weights need suitable initialization and usually further training before they are useful. Adding a string to the tokenizer alone does not give the model a learned representation for it.