The exact number you get depends heavily on how you count words, so the same text can score differently depending on your method.
TL;DR:
- Knowing 98% coverage generally requires understanding about 8,000 to 9,000 word families for reading; lower coverage levels significantly impair comprehension.
- Coverage percentages are best used as planning tools; high scores do not guarantee understanding, especially if key words or expressions are unfamiliar.
- Accurate measurement depends on clear rules for counting by lemmas, word families, or forms, and should be validated with quick comprehension checks.
- Pre-teaching five to eight high-impact words per text optimizes learning without overloading the learner, especially for complex, specialized, or academic passages.
- Tools like WordByWord automate coverage tracking and vocabulary review, turning raw percentages into actionable study and reading decisions.
Table of Contents
- What “known-words percentage” actually measures
- What the research says about coverage thresholds
- How to measure known-words percentage for a text
- Using coverage targets to choose texts and design activities
- Limitations and common pitfalls of a known-words percentage
- A step-by-step workflow to measure and act on coverage
- How WordByWord turns comprehension measurement into study practice
- Treating coverage numbers as a planning aid, not a verdict
- Try WordByWord to measure your own reading comprehension level
- Sources
- FAQ
What “known-words percentage” actually measures
Known-words percentage sounds like a single, stable number. It isn’t. The figure changes depending on whether you count running tokens or unique word types, and whether you treat related forms as one word or many.
Running-token coverage is the standard for reading research: you count every word occurrence in a passage, including repeats, and calculate what share of those occurrences the reader already knows. A 500-word article that repeats “climate” fifteen times only asks the reader to know that word once, but each occurrence counts toward the total. This matters because comprehension depends on how often you stumble while reading, not on how many distinct words a text contains. A text with a small vocabulary of difficult words repeated often can feel easier than a text with more variety, even if the type count looks scarier on paper.
Type counting, by contrast, tracks unique words rather than occurrences. It’s useful for describing a text’s lexical range but a poor proxy for how comprehension actually unfolds while reading, since a reader hits repeated tokens far more often than the type count suggests.
Then there’s the question of what counts as “one word.” A lemma groups inflected forms of the same base (“run,” “runs,” “running”) together. A word family goes further, grouping derived forms too (“run,” “runner,” “runaway”) under a single headword. Choosing lemmas versus word families changes your denominator and, with it, your final percentage. Paul Nation’s vocabulary size and coverage research uses word families as its counting unit, which is one reason its coverage figures don’t map neatly onto studies that count lemmas or raw word forms. If you switch counting units midway through a project, your numbers stop being comparable, even on the same text.
There’s a third distinction worth holding onto: population-level prevalence versus an individual’s known-word percentage. Prevalence describes how many people in a large sample recognize a given word, a property of the word itself. Known-words percentage, in contrast, describes one learner’s relationship to one specific text. A word can have high prevalence across English speakers generally and still be unfamiliar to a particular learner reading a particular passage. The two measures answer different questions: one tells you how “common” a word is across a population, the other tells you whether this reader can handle this text right now.

What the research says about coverage thresholds
The 95% and 98% figures that circulate in language-teaching circles trace back to a specific line of research on the relationship between vocabulary coverage and comprehension.
The most cited study here comes from work associated with Alderson, later revisited by Schmitt and colleagues on the Lextutor platform. Their research, drawing on 661 participants across eight countries, found a fairly linear relationship between the percentage of words a reader knows in a text and how well they comprehend it, tested across two separate reading passages. Alderson’s study on vocabulary coverage and reading comprehension recommended 98% as a more defensible target for academic reading than the older 95% benchmark, though the researchers were careful to note that neither number functions as a hard cutoff. Comprehension rose steadily with coverage rather than jumping at some magic threshold, which means 95% and 98% work better as planning heuristics than as pass or fail lines.
Paul Nation’s vocabulary-size research translates these coverage percentages into concrete word counts. His work indicates that English learners typically need roughly 8,000 to 9,000 word families for 98% coverage of written text, and about 6,000 to 7,000 word families for the same coverage of spoken text, since spoken English draws on a narrower, more frequent vocabulary.
Brysbaert and colleagues approached the question from the vocabulary-size side rather than the coverage side. Their large-scale estimate found that a median 20-year-old native English speaker knows about 42,000 lemmas, equivalent to roughly 11,100 word families. Brysbaert’s team was explicit that these counts shift depending on definitional choices, whether you count lemmas or families, and how strictly you define “knowing” a word. Two studies using different definitions can report very different totals while both being correct within their own terms.

How to measure known-words percentage for a text
Measuring coverage for a specific text is a reproducible process, whether you’re a teacher prepping a reading assignment or a learner checking whether an article is worth the effort.
Start by fixing your counting rules before you count anything, since changing them halfway through invalidates comparisons across texts.
- Decide whether to count by lemma, word family or raw word form, and stick with that choice for every text you measure.
- Decide how to treat proper names, numbers and multiword expressions such as “in spite of” or “give up,” since these often get skipped or miscounted.
- Decide whether case and punctuation matter (usually they don’t for this purpose).
- Choose a sample of adequate length: a few hundred running tokens gives a much more stable estimate than a 40-word paragraph.
- Run your chosen measurement method and record the resulting percentage alongside the rules you used.
For short texts, a manual checklist works fine. Paste the passage into a spreadsheet, list each unique word, mark known or unknown, then weight by how often each word appears to get a token-based percentage rather than a type-based one.
For longer texts, frequency-list matching scales better. Tools built on the RANGE approach compare your text against tiered frequency lists (the first 1,000 most common word families, the second 1,000, and so on) and report what percentage of tokens falls within each band. This method, described in detail in vocabulary-load research, is fast and consistent, though it assumes frequency rank is a reasonable stand-in for what any individual learner actually knows, which won’t always hold.
A third option leans on prevalence and familiarity data rather than pure frequency. The word-prevalence norms for 62,000 English lemmas dataset, built from crowdsourced judgments across more than 220,000 respondents, reports how widely recognized specific words are across a large population. Cross-referencing a text against prevalence data can catch cases where a word is technically low-frequency but still broadly familiar, something raw frequency counts miss. An integrated tool that checks a learner’s own saved vocabulary against a text, such as WordByWord’s Comprehension level feature, shortcuts this whole process by comparing the text directly to what that individual already knows rather than to population averages.
Whichever method you use, validate it. Coverage percentage predicts comprehension trends well, but it doesn’t guarantee understanding of any single passage.
Pro Tip: After running a coverage check, ask one or two quick comprehension questions about the main idea; if the reader answers correctly despite a low score, or struggles despite a high one, treat the questions as the tiebreaker.
- A 98% score with poor comprehension answers usually points to a few critical unknown terms, not a general vocabulary gap.
- A 95% score with strong comprehension answers suggests the text is workable with light support.
Using coverage targets to choose texts and design activities
Coverage percentages become useful once you attach them to a decision: which text to assign, whether to pre-teach vocabulary, and what kind of reading task fits.
For independent study, where a learner reads alone with no dictionary or teacher support, 98% coverage is the safer target, matching the recommendation from Alderson’s coverage research. Below that, unknown words start interrupting the reading process often enough to break comprehension flow. For guided, in-class reading with a teacher present to field questions, coverage can drop lower still, because scaffolding replaces some of what missing vocabulary would otherwise cost. Readers weighing extensive versus intensive reading approaches will find the coverage target shifts meaningfully between the two.
Pre-teaching vocabulary makes sense when a small number of unknown words carry outsized importance for the text’s central idea. Rather than pre-teaching every unfamiliar word, a more efficient approach targets five to eight high-impact terms that a reader will hit repeatedly or that carry the argument. Beyond that number, pre-teaching starts eating into class time without proportional payoff, and incidental learning from context takes over for the rest.
Genre changes the math too:
- News writing tends to recycle a fairly narrow set of high-frequency words, so shorter samples give stable coverage estimates.
- Fiction introduces more dialogue-driven, informal vocabulary, which behaves differently from expository prose.
- Academic text draws on specialized, lower-frequency vocabulary concentrated in specific subfields, so a general coverage score can mask difficulty in a narrow but critical set of terms.
For any genre, a sample under 100 to 150 running tokens produces a noisy estimate. Longer samples, several hundred tokens or more, smooth out the effect of any single unusual word.
Limitations and common pitfalls of a known-words percentage
A coverage number is a useful planning tool, not a verdict on whether a text will make sense. Several blind spots are worth knowing before you lean on the figure too heavily.
Recognition is not the same as usable understanding. A learner might mark a word “known” because they’ve seen it before, without being able to retrieve its meaning quickly enough to keep pace with the sentence around it. Coverage tools count recognition, not retrieval speed or depth of understanding, so a high score can still hide gaps.
Multiword expressions complicate counting further. Phrases like “put up with” or “on the other hand” carry meaning as a unit, but most counting methods score each word separately, which can inflate the apparent coverage of a passage whose real difficulty lies in phrasal meaning rather than individual words.
Short texts produce unstable percentages. A 60-word paragraph with two unusual words already sits at a coverage level that a slightly different 60-word paragraph might miss by ten percentage points. Treat any measurement on a short excerpt as a rough estimate rather than a precise reading, a point echoed in vocabulary-load evaluation research.
Frequency and importance don’t always align. A single rare but central term, the name of a chemical process in a science article, say, or a key legal concept in a policy piece, can block comprehension of an entire paragraph even when it’s the only unknown word in it.
Finally, resist treating one number as the whole picture. Educators should validate coverage figures against comprehension performance rather than trusting the percentage alone.
A step-by-step workflow to measure and act on coverage
A practical routine turns coverage measurement from a one-off exercise into something teachers and learners can repeat across every new text.
- Choose the text and sample. Pick a passage of at least a few hundred running words when possible, matching the genre students will actually encounter.
- Run the coverage check. Use a frequency-list tool, a prevalence dataset, or an integrated comprehension-level feature to get a percentage against the reader’s known vocabulary.
- Validate with comprehension. Ask one or two questions about the main idea or a key detail; a mismatch between the score and the answers signals a problem word or a measurement quirk.
- Decide on intervention. A score near 98% with good comprehension answers means assign as is. A lower score, or a mismatch, means pre-teach a handful of terms or scaffold with a glossary.
- Schedule a follow-up check. Re-measure coverage on the next text in a few weeks to see whether the learner’s vocabulary is closing the gap.
A single coverage check plus two comprehension questions takes most teachers under ten minutes per text once the workflow is familiar, which fits comfortably into lesson planning rather than competing with it.
Pro Tip: Batch pre-teaching across an upcoming unit rather than one text at a time; many high-impact words recur across related readings, so teaching them once covers several assignments at once.
Following measurement with action is what separates a coverage percentage from a decision. A practical classroom protocol that pairs an automated coverage check with a couple of comprehension questions gives teachers a fast, repeatable basis for choosing between assigning as is, scaffolding, or swapping the text entirely.
How WordByWord turns comprehension measurement into study practice
WordByWord’s Comprehension level feature applies the coverage logic described above directly to whatever a learner is about to read or watch.
Spotlight mode extends that measurement onto the page itself. Instead of a single summary percentage, it highlights individual words by how well the learner knows them, from completely unknown to fully mastered, so unfamiliar vocabulary stands out at a glance while reading. As a learner’s knowledge grows, the highlights fade, which gives a visual, ongoing version of the coverage tracking a teacher might otherwise do with a spreadsheet.
Once a reader spots the words behind a low comprehension score, WordByWord turns them into study material with one click, translating any word directly on the page, in a PDF or in YouTube subtitles across more than 50 languages. Those words get saved into collections, which can be built manually, generated by AI from a prompt, or imported from Anki, Quizlet or a spreadsheet via CSV for anyone migrating an existing word list. From there, spaced repetition flashcards, offered across seven training modes including typing, sentence builder and listening, turn a one-time vocabulary gap into long-term knowledge.
Students then read with Spotlight mode active, saving unfamiliar words as they go, and review that day’s new vocabulary through flashcards afterward. The teacher gets a fast, per-student read on text fit without running a manual coverage check for every learner individually.
Treating coverage numbers as a planning aid, not a verdict
Coverage percentages are useful precisely because they’re simple, and that simplicity is also their biggest risk. A single number invites teachers and learners to treat it as a final answer, when it works far better as a starting point for a judgment call.
The research is consistent on this: coverage predicts comprehension trends well across groups, but any individual reading experience depends on which specific words are missing, how central they are to the passage, and how much support is available. Treat a coverage score the way you’d treat a fitness tracker’s step count: informative, worth tracking over time, and not something to argue with when your own body, or in this case your own comprehension, tells you something different.
The most useful habit is repetition rather than precision. Measuring coverage once on one text tells you about that text. Measuring it across several texts over a few weeks tells you whether a learner’s vocabulary is actually growing, which is the number that matters for planning. Tools that make coverage visible without extra manual work earn their place in a routine; the judgment about what to do with the number still belongs to the teacher or the learner reading the room, not the percentage itself.
— WordByWord Team
Try WordByWord to measure your own reading comprehension level
Everything in this guide, running-token coverage, the 95% and 98% targets, the ten-minute measurement workflow, works better with a tool that checks a real text against your real vocabulary instead of a generic list.
WordByWord’s Comprehension level feature does exactly that: open any article, PDF or YouTube video, and see what share of its words you already know before committing time to it. Spotlight mode then marks the unfamiliar words directly on the page, from unknown to mastered, so you can decide in seconds whether a text fits your level or needs a different approach. Save what you don’t know into a collection, built by hand, generated by AI from a prompt, or imported from Anki, Quizlet or a CSV file, and review it through flashcards across seven modes including typing, listening and sentence building, scheduled by spaced repetition so the words stick.
The Chrome extension and web app work together across more than 50 languages, and everything happens inside the browsing and viewing you’re already doing rather than in a separate study session. The Free Forever plan covers the core comprehension and vocabulary features with usage limits, while Premium, at $5.99 per month, removes them for anyone measuring and studying regularly.
Install the extension, open a text you’re considering for your next reading session, and check its comprehension level against your own saved words to see the workflow in this guide in action.
Sources
- The Percentage of Words Known in a Text and Reading Comprehension (Alderson / Schmitt)
- Paul Nation: vocabulary size and coverage (RANGE / vocabulary load resources)
- How Many Words Do We Know? Practical Estimates of Vocabulary Size (Brysbaert et al., 2016)
- Word prevalence norms for 62,000 English lemmas (2018)
FAQ
Is knowing 10,000 words fluent?
Fluency doesn’t hinge on a single vocabulary count, since comprehension depends on which words a text uses, not just how many words a learner has stored overall. Vocabulary-size research from Brysbaert et al. suggests a median adult native speaker knows roughly 42,000 lemmas, far beyond a few thousand, though Nation’s coverage research indicates around 8,000 to 9,000 word families can already support 98% coverage of general written text.
What are the 100 most commonly used words?
The most frequent English words are overwhelmingly short function words: articles like “the” and “a,” pronouns like “I” and “you,” prepositions like “of” and “in,” and common verbs like “is” and “have.” These high-frequency words make up a large share of any English text’s running tokens, which is why frequency-based coverage tools weight the first thousand word families so heavily.
Is 60% of English Latin?
Estimates of Latin’s contribution to English vocabulary vary widely depending on whether you count direct borrowings, French-mediated Latin words, or scientific and academic terms, and no single authoritative figure applies across all registers of English. This article does not have a sourced figure for that specific claim, so it’s best treated as a commonly repeated estimate rather than a settled number.
What is the average number of words a person knows?
The figure depends heavily on age and how “know” and “word” are defined. Brysbaert et al. (2016) estimate a median 20-year-old native English speaker knows about 42,000 lemmas, equivalent to roughly 11,100 word families, and note that vocabulary size keeps growing with age.
How is known-words percentage different from a standard vocabulary test score?
A vocabulary test score usually measures how many words a learner knows in general, independent of any specific text. Known-words percentage measures coverage of one particular passage, so the same learner can score differently on different texts depending on genre and topic, as described in Alderson’s coverage research.




