Text Tools
Lexical diversity calculator.
Paste text to calculate type-token ratio (TTR), moving-average TTR (MATTR), Guiraud’s R and Herdan’s C. Every formula and limitation is shown, and your text stays in your browser.
Input
Results
Type or paste text to see the results.
About this lexical diversity calculator
Lexical diversity measures how varied the vocabulary in a text is. A token is one running word; a type is one distinct word. With N tokens and V types:
- TTR = V / N. Ranges from near 0 (heavy repetition) to 1 (no word repeated).
- Guiraud’s R (root TTR) = V / √N.
- Herdan’s C (log TTR) = ln V / ln N. Undefined when N = 1.
- MATTR = the mean of the TTR of every window of W consecutive tokens, sliding one token at a time (Covington & McFall, 2010). It needs N ≥ W.
Plain TTR depends on text length
The longer a text gets, the more its words repeat, so plain TTR falls as N grows even when the writer’s vocabulary has not changed. Do not compare the TTR of a 200-word text with the TTR of a 2,000-word one. Guiraud’s R and Herdan’s C were proposed to reduce this effect but do not remove it (Tweedie & Baayen, 1998). MATTR uses a fixed window, so texts of different lengths are compared on equal footing as long as you use the same window size.
How words are counted
Text is split into words with your browser’s Intl.Segmenter (word granularity, Unicode-aware), keeping only word-like segments, so punctuation is ignored and numbers count as words. “don’t” is one token. There is no lemmatisation: “run” and “running” are different types. With case ignored, text is lower-cased before types are counted. If your browser lacks Intl.Segmenter, a simpler letter-and-digit splitter is used instead.
Sources
- Covington, M. A., & McFall, J. D. (2010). Cutting the Gordian Knot: The Moving-Average Type–Token Ratio (MATTR). Journal of Quantitative Linguistics, 17(2), 94–100. doi:10.1080/09296171003643098
- Tweedie, F. J., & Baayen, R. H. (1998). How Variable May a Constant Be? Measures of Lexical Richness in Perspective. Computers and the Humanities, 32, 323–352. doi:10.1023/A:1001749303137 (presents R and C)
- Guiraud, P. (1954). Les caractères statistiques du vocabulaire. Presses Universitaires de France (root TTR).
- Herdan, G. (1960). Type-Token Mathematics. Mouton, The Hague (log TTR).
The formulas above are the standard definitions as they appear in the Tweedie & Baayen review and in the documentation of the quanteda R package (textstat_lexdiv). The two books are print-only, so this page cites them as they are cited there. For a simpler baseline, count total words, characters, sentences and paragraphs with the Word Count Checker, or browse all text analysis and writing tools.
Lexical diversity calculator FAQ
What does lexical diversity measure?
Lexical diversity describes how varied a text’s vocabulary is. This calculator reports distinct-word measures including TTR, MATTR, Guiraud R and Herdan C; it does not grade the quality of the writing.
Why does TTR usually fall for longer texts?
Words repeat as a text grows, so the share of distinct words usually falls. Compare plain TTR only across similarly sized texts, or use the same MATTR window for each text.
Is my text uploaded for lexical analysis?
No. Tokenization and every lexical-diversity calculation run locally in your browser. This tool does not upload, store or submit the text you enter.