Lexical Diversity

Paste your draft, get a variety score, and click any finding to highlight it in your text.

Your text

0 words · 0 sentences

Vocabulary variety

Needs 30+ words

Unique words

Avg sentence

Sentences

Paste at least 30 words and the analysis appears here: your variety score, overused words, repeated transitions, and monotonous sentence patterns - click any finding to highlight it in your text.

What is lexical diversity?

Lexical diversity measures how varied your vocabulary is - the ratio of unique words to total words. Writing with low diversity leans on the same words, transitions, and sentence shapes again and again, which makes essays and reports feel flat and repetitive. Grammar checkers rarely flag this: the sentences are correct, they just sound monotonous.

How to use the checker

  1. Paste your draft into the editor - everything runs in your browser.
  2. Check your variety score, then open the findings on the right.
  3. Click any overused word, transition, or sentence pattern to highlight every instance in your text, and revise using the suggested alternatives.

Frequently asked questions

Is the lexical diversity checker free?

Yes - it's completely free with no sign-up. The analysis runs on your device, so there are no word limits and your text is never uploaded to a server.

How is the variety score calculated?

It's based on the moving-average type-token ratio (MATTR), a standard linguistic measure of vocabulary variety that stays fair for both short paragraphs and long essays. Related word forms like “analyze” and “analyzing” are grouped, so the repetition findings reflect real word roots.

Why do some repeated words not get flagged?

Common function words (“the”, “and”, “of”) repeat naturally in all writing, so they are excluded. A meaningful word is flagged when it repeats often for the length of your text or when several uses land close together in one passage.

Academic register check

Everything above runs in your browser. This one does not — it compares your writing against published academic prose using a model on our server, so your text is sent to us when you press the button. It is not stored, and nothing runs until you ask.

Needs at least 80 words

What the score actually measures

The number at the top is derived from MATTR - a moving-average type-token ratio. The plain version of that idea is: take a window of 100 consecutive words, count how many of them are distinct, then slide the window one word along and do it again. The score is the average across every window in your text.

The averaging is the part that matters. The obvious measure - unique words divided by total words - falls apart on anything long, because every text repeats "the" and "of" more as it grows. A 2,000-word essay scores worse than a 200-word one written by the same person with the same vocabulary. MATTR does not have that problem: a fixed window means a long text and a short one are measured on the same terms.

A raw MATTR of about 0.55 is very repetitive prose and about 0.85 is genuinely wide-ranging; the 0-100 figure here is stretched across that range so the ends are usable rather than crowded into a narrow band.

It measures variety of word forms, and nothing else. It cannot tell whether the words are the right ones, whether the argument holds, or whether the writing is any good. A thesaurus rampage will raise the score and make the prose worse.

Why a low score is a symptom, not a verdict

Repetition is not automatically a fault. Academic and technical writing repeats its key terms deliberately, because swapping in a synonym for a defined term is how a reader loses the thread. If your dissertation says "chloroplast" ninety times, that is precision, not poverty.

What a low score usually points at is the connective tissue rather than the terminology: the same handful of verbs doing all the work, every paragraph opening the same way, one hedging phrase repeated until it stops meaning anything.

So read the score alongside the findings underneath it. The score tells you there is something to look at; the findings tell you where. If the repeated words are your subject matter, ignore it. If they are "important", "various" and "significantly", you have found your edit.

The three kinds of finding

Overused words groups related forms together before counting. "Analyse", "analysed", "analysing" and "analysis" are one item, not four, because a stemmer reduces them to a common root first - otherwise the list fills with the same idea wearing different endings.

Transitions are matched as whole phrases, so "on the other hand" is one hit rather than four common words. They get their own category because they are the single most reliable sign of prose that has been assembled rather than written: "moreover", "furthermore" and "in addition" stacked three paragraphs running.

Sentence flow looks at rhythm rather than vocabulary. It flags runs of four or more consecutive sentences whose lengths sit in a narrow band, and three or more in a row that open with the same word. Both are invisible while you are writing and obvious the moment they are pointed out.

Every finding is clickable and highlights each occurrence in your text, because a count on its own tells you a problem exists somewhere and a highlight tells you where to put the cursor.

Using it without wrecking your writing

Fix the rhythm findings first. Breaking up a run of eight sentences that are all nineteen words long improves a piece more than any individual word swap, and it costs nothing in precision.

Then look at the transitions. Most can simply be deleted: if the logical relationship between two sentences is already clear, "furthermore" is noise. Deleting is usually better than substituting.

Only then consider the word list, and only where the repetition is genuinely accidental. The synonym suggestions are a prompt, not an instruction - they come from a general-purpose English database that has no idea what your sentence is about, and the wrong synonym is worse than an honest repetition.

The vocabulary analysis is not uploaded. It runs entirely in your browser, which is why it is instant and why it works with the network off. Two things do leave your machine, both only when you ask: the optional synonym lookup, which fires when you click a single word, and the academic register check, which needs a model on our server and so sends the text you pressed the button on. Neither is stored.

Common questions

What is a good lexical diversity score?

There is no single target, and chasing one is a mistake. Most competent non-fiction lands somewhere in the middle of the scale. Technical and academic writing scores lower by nature because it repeats defined terms on purpose. The useful comparison is between two drafts of your own piece, not against a number.

Does a low score mean my writing is bad?

No. It means the same word forms recur often. That can be precision, a house style, or a genuinely narrow vocabulary - the score cannot tell which. Read the findings underneath it and decide for yourself which repetitions are doing work.

Why does my score change when I paste more text?

MATTR is far more stable across lengths than a plain type-token ratio, but it is not perfectly flat: adding a section on a new topic introduces new vocabulary and usually nudges the score up. Compare like with like - whole draft against whole draft.

Is my text uploaded anywhere?

No. The whole analysis - tokenising, stemming, scoring, and finding the repetitions - happens in your browser. The only network request is the optional synonym lookup for a single word, and only when you ask for it.

Why are "analyse" and "analysis" counted as one word?

Because they are one idea. A stemmer strips inflections so related forms group together, which stops the list filling up with the same root repeated in four tenses and gives you a truer count of how often you are reaching for that concept.

Can I use this on an essay in another language?

The score is language-agnostic - counting distinct word forms works on any text. The findings are not: the transition phrases, the stemmer and the overused-word lists are all English, so those will be quiet or wrong on other languages.

Related tools