Definition

Self-supervised learning trains a model on targets derived from the training data itself. A training procedure might withhold part of an example and ask the model to predict it. The withheld part supplies the target used to calculate loss. Nobody has to add a separate answer label to each example.

For language models, the target might be the next token in a sequence or a token hidden from the input. The term describes where the training signal comes from, not a particular model architecture.

Simple example

Suppose the training text contains The certificate expired yesterday. With a next-token objective, the model receives the tokens up to expired and predicts what follows. The target is read from the rest of that sentence. Its exact token boundary depends on the tokenizer.

The trainer computes a loss from the model’s predicted token probabilities and the target token, then updates the model parameters. It repeats this across many positions and documents.

Why it matters

Text is much easier to collect than carefully reviewed question-and-answer examples. Self-supervised objectives let model builders use large text collections for pretraining before adapting a model to specific tasks.

This helps explain why a base language model may complete text fluently without reliably following a user’s instructions. Its training task was to predict text. Producing a useful answer to a request is a different test.

One important nuance

Deriving targets automatically does not make them trustworthy. If a source sentence makes a false claim, predicting its next token still counts as a correct training prediction. The loss measures agreement with the source text, not factual accuracy. Model builders still need to examine their data, and application teams need to test the behavior they rely on.