301 chain301 chain
Permanent redirects stacked one after another.
Found by Redirect checker →
302302
A temporary redirect on a page that moved for good.
Found by Redirect checker →
ContentUpdated

Lemmatisation

In short

Lemmatisation reduces inflected word forms to a lemma so language analysis can group related forms without treating each spelling as a separate concept.

Use language-aware normalization

Lemmatisation differs from crude stemming: it relies on vocabulary and grammar to return a valid base form. The result depends on language, part of speech and tokenisation, so a multilingual keyword list needs language-aware handling. It is an analysis aid, not a reason to repeat every inflection on a page.

How to inspect it

Keep the raw query alongside the lemma and language label. Sample ambiguous terms manually, especially brand-like words, abbreviations and mixed scripts. Compare clusters using the original queries as well as normalized forms; users search in surface language, while the lemma helps reveal overlap.

The diagram shows how query wording, intent and location can lead to different result interpretations.

How to implement or improve it

Use lemmas to deduplicate research, discover topic variants and review query groups. Build pages around the intent and the wording people understand, then verify the resulting page against the search results. Keep an exception list rather than silently collapsing useful distinctions.

Limits and common mistakes

The mistake is assuming a shared lemma proves identical intent. Grammar similarity cannot replace SERP review, product coverage or editorial judgement.

For a keyword export, begin with a raw column, a detected-language column and a normalised column. Keep punctuation, quotes and query operators in the raw value: they may signal a different task. A Russian inflection, an English plural and a transliterated brand should not be sent through one undifferentiated rule. If the parser needs part-of-speech context, treat the result as a candidate and sample it against the original list. For example, “running shoes” and “run a business” share an English token but belong in different editorial groups. The verification step is therefore a small manual sample from every high-volume or ambiguous cluster, plus a count of records that could not be confidently normalised. Do not use a lemma as visible copy automatically; a page should use the language that answers the searcher’s question naturally.

Configuration example: run a language detector first, keep its confidence, then pass only confident records to the appropriate lemmatiser. Store the model version so a later rerun can explain a changed cluster. Review a sample that includes the most frequent queries, the longest queries and every group whose raw forms differ in more than grammar. Mark false mergers as exclusions and false splits as review candidates. When two phrases share a lemma but produce different result pages, retain both in the research model. The final verification is editorial: the target page should answer the exact original query family, not merely contain a matching stem.

A repeatable review can be run in five steps. First, remove duplicate rows without destroying their source counts. Second, identify language and preserve the raw phrase. Third, generate a lemma only for tokens the language model recognises. Fourth, group candidates by both lemma and observed intent. Fifth, open a sample of result pages before choosing a target URL. Consider a query list containing “best trail runs”, “trail running shoes” and “run trail events”: surface similarity is useful for research, but the likely page types differ. The check is passed only when the cluster’s target page has a clear job and the excluded forms are documented. This prevents a linguistic convenience from silently becoming a content decision.

Keep the transformation reversible: raw phrase, language, lemma, exception flag and chosen target should remain separate fields. If the model changes, rerun from raw data rather than editing the lemma column by hand.

FAQ

Lemmatisation reduces inflected word forms to a lemma so language analysis can group related forms without treating each spelling as a separate concept.
Lemmatisation differs from crude stemming: it relies on vocabulary and grammar to return a valid base form. The result depends on language, part of speech and tokenisation, so a multilingual keyword list needs language-aware handling. It is an analysis aid, not a reason to repeat every inflection on a page.