Definition
Lexical similarity evaluates how similar two texts are based on their surface features: the words, characters, or tokens present. A comparison may count shared words, compare short sequences such as n-grams, or use edit distance to measure the number of character changes needed to transform one string into another.
There is no universal formula for lexical similarity. Some methods ignore word order, while others compare sequences. How the text is split into tokens, and whether capitalization or punctuation is normalized, also influence the outcome. The resulting score reflects resemblance under those specific rules.
Simple example
Compare Restart the service with Restart the payment service.
Lowercase both texts, split them on spaces, and treat each as a set of unique words. The sets share three words: restart, the, and service. Together, they have four distinct words, including payment.
Jaccard similarity divides the number of shared words by the total number of distinct words in both sets: 3 / 4 = 0.75. This approach ignores word order and repeated words. The score measures word overlap, not a 75% chance that the instructions have the same meaning.
Why it matters
Lexical comparison helps identify nearly duplicated passages before they fill a RAG context window. It can also measure wording overlap between generated text and a reference answer without invoking another model.
Use lexical similarity when shared wording is a useful signal and you need to examine what matched. For comparisons across runs, keep the method, preprocessing rules, and any corpus-dependent parameters consistent.
One important nuance
Lexical similarity does not establish agreement in meaning. Adding not can reverse the intent of an instruction while leaving almost all words the same. A valid paraphrase may have little word overlap with its reference.
Do not rely on lexical similarity alone to judge answer correctness or to determine whether a summary preserves the source’s claims. These tasks require methods that assess the actual meaning of the text.