Definition
TF-IDF stands for term frequency times inverse document frequency. It assigns a numerical weight to each term in a document. A term gets more weight when it occurs often in that document and appears in fewer documents across the collection, or corpus.
Term frequency (TF) measures occurrences within one document. In its simplest form, it is the raw count. Document frequency (DF) counts how many documents contain the term. Ten occurrences in one document still contribute only one to DF. IDF uses that count to reduce the contribution of terms shared by many documents.
One classic formulation is weight = TF × log10(N / DF), where N is the total number of documents and DF is the number containing the term. Implementations may smooth IDF, scale TF, or normalize the resulting document vector.
Simple example
Suppose a collection has 100 incident reports. One report contains “service” three times and “deadlock” three times. Across the collection, 80 reports contain “service” and five contain “deadlock”.
Using the raw counts and formula above, before vector normalization:
| Term | TF | DF | IDF | TF-IDF weight |
|---|---|---|---|---|
| service | 3 | 80 | 0.097 | 0.29 |
| deadlock | 3 | 5 | 1.301 | 3.90 |
The terms occur equally often in this report. “Deadlock” gets more weight because it appears in fewer reports across the collection.
Why it matters
TF-IDF turns text into numerical features for document comparison, clustering, or classification. Each dimension corresponds to a vocabulary term. Most dimensions are zero because a document contains only part of the vocabulary.
It also provides a term-weighting approach for sparse retrieval. You can inspect which terms contributed weight without calling a neural model.
One important nuance
A high weight indicates statistical distinctiveness within this corpus. A rare typo can receive a high weight too. TF-IDF alone does not recognize that “deadlock” and “circular wait” describe related ideas when they share no terms.
Use TF-IDF when vocabulary overlap is a useful signal and you need inspectable text features. Do not rely on it alone when the task requires matching meaning across different wording.