Skip to tool
ecech.
💬 English Exams & Writing

Free Repeated Word and Lexical Variety Checker That Is Not Fooled by Length

Finds the word you used nine times without noticing — and reports a diversity score that survives comparison between texts of different lengths.

times or more

Words

0

Distinct words

0

MATTR — comparable

Raw TTR — not comparable

Words you repeated

Where each repeat sits

Your text, marked up

Advertisement

How the calculation works

Why a “vocabulary diversity” score usually measures length raw TTR MATTR 100 words 400 words 700 words high low Same writer, same style, one continuous text. Raw TTR falls purely because the token count grows.

How to Use This Tool

Everyone has a word they lean on without hearing themselves do it. This finds them, shows you every place they appear, and — the part that matters — gives you a variety score that still means something when you compare two pieces of different lengths.

The measurement problem nobody mentions

The standard measure of lexical variety is the type-token ratio: distinct words divided by total words. It is easy to compute and, used naively, close to useless.

The reason is arithmetic. As a text grows, the token count keeps rising but the supply of genuinely new words runs out — you have already used the, and, and most of your topic vocabulary. So TTR falls as length increases even when the writing has not changed at all. Compare a 250-word essay with a 400-word one and the shorter piece wins automatically.

Same writer, same vocabulary. One is just longer. Essay A — 250 words 140 distinct TTR 0.56 Looks more varied — but only because it stopped sooner. Essay B — 400 words 196 distinct TTR 0.49 Scores lower despite using 56 more distinct words.
Essay B has a richer vocabulary by any sensible reading and a worse TTR. This is the metric failing, not the writing.

What MATTR does instead

MATTR — moving-average type-token ratio — slides a fixed window across the text, measures TTR inside each window, and averages the results. Because every window is the same size, the length bias disappears: a 250-word text and a 4,000-word one are both scored on 100-word windows.

That is the number to compare across drafts, and it is the one shown first here. Raw TTR is shown too, greyed out, so you can see how far apart they are — but do not use it to compare two pieces.

Two toggles worth understanding

Ignore function words removes the, of, is, and and the rest of the closed grammatical class. Leave it on. Function words are not a style choice — you cannot write English without repeating them, and counting them buries the words you actually control.

Group word forms treats important and importantly, and technology and technologies, as one item each. This is suffix stripping, not a real lemmatiser, and it misses two specific classes:

  • Irregular forms. go / went, be / was, child / children stay separate, because nothing about their spelling connects them.
  • Noun/adjective pairs. important and importance stay separate, because the shared part is import and stripping that far also turns parent into par.
  • Short verbs. use / used are not grouped. The rule refuses to strip a suffix that would leave a stem of two or three letters, because loosening it turns need into ne — and a wrong grouping is worse than a missed one, since you would never think to check it.

So the grouped counts are a floor, not an exact figure. Read the list rather than trusting the number.

Reading the position strip

A word used five times spread evenly across 600 words is usually fine — it is your topic. The same word five times inside one paragraph is what a reader notices. The strip shows where each repeat falls, so clusters are visible at a glance, and words that cluster get flagged separately from words that are merely frequent.

What to do about a flagged word

Not necessarily replace it. Reaching for a thesaurus produces the other failure mode — elegant variation, where a paper calls the same thing a study, an investigation, an inquiry and an analysis and the reader starts wondering whether these are four different things. In technical and academic writing, repeating the precise term is correct.

Usually the better fix is structural: two sentences that both start with the same subject can often become one, and a paragraph that repeats a word four times is often making the same point four times.

Nothing leaves your browser

The analysis runs in JavaScript on your device and no request is made. That matters when the text is an unpublished paper or a graded assignment.

Advertisement

Frequently Asked Questions

What is a good type-token ratio for an essay?
There is no useful target, because TTR depends heavily on length — the same writer scores lower on a longer piece. If you want a number to compare across drafts, use MATTR, which measures diversity inside fixed-size windows and so is not distorted by length.
Why does my longer essay score worse on vocabulary variety?
Because raw type-token ratio falls as texts get longer. The token count keeps rising while genuinely new words run out, so a 400-word piece almost always scores below a 250-word one by the same author. This is the metric misbehaving, not your writing. MATTR corrects for it.
What counts as too many repeats?
Position matters more than count. A word appearing five times spread across 600 words is usually just your topic. The same five uses inside one paragraph is what a reader notices, so this tool flags clustering separately from raw frequency.
Should I replace every repeated word with a synonym?
No. Reaching for synonyms produces elegant variation, where the same thing is called a study, an investigation, an inquiry and an analysis until the reader wonders whether they are different things. In academic and technical writing, repeating the precise term is correct. The better fix is usually structural — merging sentences, or cutting a point made twice.
Why are 'the' and 'of' excluded?
Because they are function words, a closed grammatical class you cannot avoid repeating in English. Counting them tells you nothing about style and buries the content words you actually chose. You can switch the exclusion off, but the results become much less useful.
How does grouping word forms work?
By suffix stripping — important and importantly collapse to one item, as do technology and technologies. It is not a lemmatiser, so it misses irregular forms such as go and went, and it also skips short verbs like use and used. That second limit is deliberate: the rule refuses to strip a suffix that would leave a two or three letter stem, because loosening it turns need into ne. A wrong grouping is worse than a missed one, because you would never think to check it. Treat the grouped counts as a floor and read the list.
Is my text uploaded?
No. Everything runs in JavaScript in your browser and the page makes no network request while you use it. You can confirm it by going offline — the tool keeps working.

Related tools in English Exams & Writing

Browse all English Exams & Writing tools
The desk where ecech. tools get written: a laptop, a notebook of to-dos and a whiteboard listing the tools on the site.

Made by one person

ecech. is not a content farm. Every tool here is written and checked by hand, one at a time, by someone who wanted the tool to exist and could not find a version that showed its working.

No accounts and no sign-in, and nothing you type reaches a server — every calculation on this page runs inside your browser. The ads are served by Google and do set their own cookies, which is set out in full on the privacy page. More about the site.