logo

Comprehensible Input: How to Find Content at Your Level

·by WordByWord Team
19 min read
Comprehensible Input: How to Find Content at Your Level

Comprehensible input is language material you understand almost all of — and "almost all" turns out to be a specific number. Research on lexical coverage puts comfortable, unassisted comprehension at around 95–98% of known words, with the productive learning zone just below that. The catch: the number depends on your vocabulary, not on the content's official difficulty level, and you cannot estimate it by eye.

Every language learner has been told to consume content they can mostly understand. Almost nobody is told how to recognise it. So you open a video, struggle for three minutes, close it, and quietly conclude the problem is you.

It isn't. The problem is that "mostly understand" has a measurable threshold, and until recently there was no way to check where a given video, article or PDF fell relative to your own vocabulary. This article covers both halves: what the research actually established, and how WordByWord's Comprehension Level turns it into a number you see before you commit your evening to a video.

What is comprehensible input?

Comprehensible input is language you are exposed to that you can understand most of, even though it contains elements you have not learned yet. That is the whole definition of comprehensible input, and the word carrying it is "most": comprehensible input means material where understanding comes first and the unfamiliar parts ride along on top of it, close enough to the familiar ones to be worked out.

Krashen's input hypothesis and i+1

The term comes from Stephen Krashen's input hypothesis — the best-known component of his theory of second language acquisition. Its claim is that language is acquired by understanding messages slightly beyond your current level, not by studying rules about the language. Comprehension is the mechanism; the grammar arrives as a by-product. What gets loosely called comprehensible input theory is really this single claim plus four companion hypotheses about how acquisition works.

Krashen described that "slightly beyond" as i+1 — where i is your current competence and +1 is the next step. The formulation is deliberately loose, and that looseness is where the theory stops being useful in practice. It says input should stretch you a little, but it never says how much a little is, and it gives you no way to tell whether a specific episode of a specific show is your i+1 or someone else's. Everything below is an attempt to put a number on exactly that gap.

Comprehensible input examples

The same material is comprehensible input for one learner and noise for another. A rough illustration for a learner at an intermediate level:

  • Comprehensible: a cooking video in your target language — repetitive vocabulary, visible actions, predictable structure. You miss words and still follow everything.

  • Borderline: a podcast interview on a familiar topic. You follow the thread, lose the jokes, and have to re-listen to some sentences.

  • Not comprehensible: a legal drama with courtroom vocabulary and overlapping dialogue. You catch fragments. Nothing accumulates.

Notice what determines the category: not the format, not a label on the content, but the overlap between the words in it and the words in your head.

Why "comprehensible" is a number, not a feeling

Linguists measure that overlap and call it lexical coverage: the percentage of running words in a text that the reader or listener already knows. "Running words" means every occurrence — if the appears forty times, it counts forty times, because that is how reading and listening actually work.

Coverage is the quantitative version of Krashen's qualitative idea. And the research on it produced something surprisingly useful: hard thresholds.

What the research says about the thresholds

Where the numbers come from: extensive reading research

The thresholds below were not invented by the language-learning industry. They come out of research into extensive reading — the practice of reading large amounts of material you find easy, instead of grinding through short difficult texts with a dictionary. That tradition had a practical question to answer: how easy does "easy" actually have to be before reading in quantity does its work? Answering it meant measuring comprehension against the share of known words, and the numbers everyone now quotes are what came out of those experiments.

Reading: 98% for comfort, 95% as the floor

In the classic study, Hu and Nation manipulated a fiction text so that learners knew 80%, 90%, 95% or 100% of the words, then measured comprehension. The finding that became a reference point in the field: around 98% coverage is needed for comfortable, unassisted reading, and 95% is roughly the minimum at which adequate comprehension becomes achievable for most readers. A later replication by Kremmel and colleagues supported the same pattern.

The counterintuitive part is not the numbers themselves — it is their shape.

Listening and video are more forgiving

Van Zeeland and Schmitt ran the same logic on spoken narratives and found listening behaves differently from reading: comprehension at 90% and at 95% coverage was close enough to be practically similar, with 95% the safer target. Work on viewing comprehension — video rather than audio alone — points the same direction: images, gesture and situation carry meaning that text has to spell out, so visuals partly compensate for words you don't know.

Which matches ordinary experience. You can follow a cooking video in a language you barely speak. You cannot follow a page of prose the same way.

Why comprehension collapses instead of fading

The important property of these thresholds is that going below them does not cost you a proportional amount of understanding. Comprehension does not decline gracefully — it falls apart. Drop from 95% to 90% and you are not losing 5% of the meaning; you are losing the thread, the inferences, and the ability to guess the unknown words from context, because guessing from context itself requires that the context be clear.

This is why "just push through harder content" fails so reliably, and why it feels like a personal failure rather than a mismatch of numbers.

Why 2% unknown is not 2 words per page

The thresholds sound lax and feel strict, and the gap between those two impressions is worth closing. 98% coverage means roughly one unknown word in every fifty — about one per two lines of text, not one per page. At 90%, the number people casually describe as "understanding most of it", you are stopping every ten words.

What that looks like as a number

image 90.jpg

That is the same idea with a number attached. WordByWord counts the words in whatever you have open — the subtitles of a video, the text of an article, the pages of a PDF — matches them against the words you have actually learned, and shows the share you already know before you start. 88% on that video means ten points below the 98% at which unassisted reading is comfortable, inside the range where visuals and subtitles carry you, with the specific words responsible listed underneath.

The rest of this article is what sits behind that number: where it comes from, how it is computed, and where we deliberately made it stricter than it needed to be.

The real problem: finding input at your level

None of the above is controversial in language learning circles. The bottleneck is practical: the thresholds are personal, and nothing you open tells you where it sits relative to your vocabulary.

The usual workarounds all approximate:

  • Graded content and level labels (A2, B1, "beginner-friendly") describe an average learner. Your vocabulary is not average — it is shaped by whatever you happen to have read and watched.

  • Readability scores (Flesch and friends) measure sentence and word length in the abstract. They know nothing about you.

  • "Try it and see" works, but the cost of a wrong guess is a wasted evening and a dent in motivation.

  • Curated level systems from input-focused projects are genuinely useful, and they exist precisely because the selection problem is real — but they can only cover their own library.

What is missing is a measurement: this specific page, this specific video, against this specific vocabulary.

A comprehension score for anything you open

That measurement is what WordByWord's Comprehension Level does. WordByWord already tracks which words you know — every word you save carries a learning status as you read and watch. The comprehension score matches the full word list of what you are looking at against that personal vocabulary, counts repetitions the way coverage is measured in the research, and shows you the result before you start.

Any YouTube video

On YouTube a button appears next to the player with the score in its tooltip — Comprehension: 88% — click for details. The score is computed from the video's subtitles in the language you are learning, so it reflects what is actually said rather than the title or the topic. If subtitles in that language do not exist, WordByWord says so plainly instead of inventing a number.

Any web article

On ordinary web pages the same engine runs over the article text and shows a small chip with the count of new words. One glance tells you whether this is a five-minute read or a forty-minute slog with a dictionary.

image 96.jpg

Any PDF

PDFs are scored progressively: text is fed to the engine page by page and the score updates as more of the document is read, so a long document produces a running estimate rather than nothing at all.

image 108.jpg

Any language you are learning

The score has nothing English-specific about it. It compares tokens to your saved vocabulary, whatever language that vocabulary is in — Spanish, French, Japanese, German, Korean, Italian, Chinese and dozens more. Learners of Spanish looking for input at their level get exactly the same number for a Spanish video that learners of Japanese get for a Japanese one.

image 99.jpg

What the four zones mean

The raw percentage is turned into four zones, so you can decide without doing arithmetic:

ZoneCoverageWhat it means Easy95% and upAn easy watch or read — good for consolidating what you already know Your level85–94%The growth zone: you follow the content and still meet new words Challenging70–84%Doable with effort — interactive subtitles, pauses and highlighting will carry you Too hardbelow 70%Save it. Start with its most frequent words instead

These four zones are WordByWord's guidance, not academic thresholds. Only the 95% line comes straight from the coverage research, and that research measured comprehension without any support. The zones assume the opposite: that you have interactive subtitles, highlighting and click-to-translate at hand. The 85% and 70% boundaries are our own calibration for those conditions — no study produced either number.

The wording adapts to what you opened — a video is "an easy watch", an article "an easy read", and the too-hard hint points you at the word list rather than just telling you no.

How the number is computed — and where we made it stricter on purpose

A score like this is only worth having if you know what it counts. Ours makes several deliberately conservative choices, and they are worth stating openly.

  • "Known" means genuinely known. A word counts towards your coverage only from the Learned level upwards. Words you have saved but are still working through — New, Recognised, Familiar — are reported separately as "learning" and do not inflate the percentage.

  • The denominator excludes noise, not difficulty. Anything the highlighting engine would not treat as a word to learn is left out entirely: your personal ignore list, text in a different writing system, URLs, e-mail addresses, hashtags, and tokens without letters. The score measures vocabulary, not page furniture.

  • Too little text, no score. Below 20 counted tokens nothing is shown. A percentage over a handful of words is noise dressed up as data.

  • 100% does not lie. 99.5% rounds down to 99. You only see 100% when coverage is genuinely complete.

  • The growth zone sits below the academic threshold, and that is intentional. The studies measured comprehension without support. WordByWord gives you click-to-translate, interactive subtitles and highlighting, and the listening and viewing research shows spoken content with visuals is more forgiving than prose. Support raises effective coverage, so 85–94% is a realistic growth zone here even though 95% is the unassisted-reading benchmark.

The words that buy the most understanding

Why a handful of words moves the number

Word frequency is extremely lopsided — the pattern known as Zipf's law. A small number of words does most of the work in any text, and a long tail of words appears once or twice. Because frequency is lopsided, the unknown words in a specific video are also lopsided: a handful of them account for a large share of the gaps. Learning the ten most frequent unknown words of one video can move your coverage of that video by several percentage points at once — a small, targeted investment with a disproportionate return.

Because the score counts repetitions, it can also answer the more interesting question: which words are worth learning for this content. The panel shows the most frequent new words with their occurrence counts, and above them a line like:

These 11 words add +6% understanding of this video

That is not a marketing estimate. It is arithmetic over the transcript and your vocabulary: if those words moved into the "known" column, the coverage percentage would rise by that much.

The list is not a fixed "top 10". Its length adapts to the shape of the text — between 5 and 15 words, chosen to cover the head of the frequency distribution rather than a round number of entries. A dense technical article where four terms repeat constantly needs a shorter list than a documentary where new words are spread thin; a fixed size would be wrong in both directions. Words with identical frequency are never split across the boundary, because cutting "five of the six words that appear ten times" would be arbitrary.

image 104.jpg

Each word can be saved at a level, ignored, heard aloud, or read with its translation right there. Learn them, reopen the video, and the score has moved — usually enough to shift it from Challenging into Your level.

Do you have to avoid translation?

This deserves a straight answer, because part of the input-focused community — automatic language growth, and projects built on it — deliberately avoids translation, dictionaries and word lists, on the argument that they interrupt acquisition. WordByWord offers all three. So it is fair to ask whether a tool like this belongs anywhere near comprehensible input.

Two things are worth separating.

The score is a measurement, not a method. It tells you what share of the words in front of you you already know. That is useful whether your next move is to click a word for its translation or to watch the video twice and let the meaning settle on its own. Nothing about the percentage assumes translation.

The selection problem is method-neutral. Whatever school you belong to, you still have to choose what to watch tonight, and choosing badly costs you either boredom or frustration. That is exactly why curated level systems exist in translation-free communities. If you practice input without translation, use the score and the zones and leave the rest of the extension switched off — you will still be answering the question those level systems were built to answer, and for content nobody has curated for you.

We are not claiming to be an implementation of Krashen's method. We are claiming to measure the "comprehensible" part.

How to use it in practice

  1. Choosing a series. Open the first episodes of three candidates and compare scores before you commit. One will be in your growth zone; the others are for later or for consolidation.

  2. Filtering a reading list. Scan your saved articles and let the new-words chip sort them. Read the ones in your zone now; keep the too-hard ones and revisit them in a month, when the number will have changed on its own.

  3. Getting through a book or PDF. Score it, and if it lands in Challenging, learn the word list first instead of starting and abandoning it. Ten to fifteen words is an evening; the difference it makes to a 300-page document is not.

FAQ

What is a comprehensible input example?

A video, podcast or text where you understand nearly everything and still meet a few new words — a cooking video with repetitive vocabulary and visible actions is a common example for early learners. The same material stops being comprehensible input once your level moves past it, and starts being it once your vocabulary catches up. It is defined by the match to the learner, not by the material alone.

What is the comprehensible input method?

Broadly, an approach to language learning that prioritises understanding large amounts of material at or slightly above your level over studying grammar rules or drilling isolated words. Variants differ sharply on details — especially whether translation is allowed — but they share the premise that acquisition comes from understood input.

What is Stephen Krashen's theory of comprehensible input?

Krashen's input hypothesis holds that language is acquired by understanding messages containing structures slightly beyond your current level, which he summarised as i+1. It is one of five hypotheses in his broader model of second language acquisition, and it is descriptive rather than quantitative — it does not specify how far beyond your level input should be.

Does comprehensible input work?

The evidence that understood input drives vocabulary and grammar acquisition is strong, and the coverage research above is part of it: comprehension depends on how much you already know, in measurable proportions. The debated questions are narrower — how much explicit study helps alongside input, and whether translation interferes. The practical failure mode is not the principle but the execution: input that is too hard is not comprehensible input, and most people overestimate how much they understand.

How to get comprehensible input as a beginner?

At the start, coverage against native content is low enough that unassisted comprehension is not realistic, so lean on material where meaning comes from outside the words: visuals, gesture, familiar situations, slow speech made for learners. Use the score to avoid the trap of picking something that merely looks easy — a short video with dense vocabulary can be harder than a long, repetitive one.

How much comprehensible input per day do you need?

There is no established number, and anyone quoting one precisely is quoting a habit rather than a finding. What the coverage research does imply is that the hours are not interchangeable: an hour spent on material where you know 90% of the words does far more than an hour spent on material where you know 50%, because below the threshold you stop being able to infer anything and the time stops compounding. Consistency and match to your level both matter more than the size of the daily block.

How many words do I know?

Vocabulary size tests give a rough estimate of the total, and it is a satisfying number to have. But it is not the number that decides what you can watch tonight: what matters is the overlap between your words and the words in a specific piece of content, and two texts with the same official difficulty can sit far apart on that measure. WordByWord counts the words you have saved and learned, then reports the overlap for whatever you have open — a total vocabulary figure cannot tell you that.

How do I know if a video is comprehensible input for me?

Check the share of words in it you already know. Above roughly 95% it will feel comfortable; 85–94% is the productive zone where you follow along and still pick things up; below about 70% you are pattern-matching rather than understanding. WordByWord computes that percentage for a video, article or PDF against your own saved vocabulary before you start.

Is 95% coverage the same as 95% comprehension?

No, and this trips people up. Coverage is the proportion of words you know; comprehension is how much of the meaning you get. They are related but not equal — which is the whole reason the research had to measure comprehension separately at fixed coverage levels, and why the thresholds are as high as they are.

Is there a comprehensible input app?

Several kinds exist, and they solve different halves of the problem. Graded-reader apps and curated video libraries hand you material already sorted by level, which works well until you want something outside their catalogue. Browser-based tools instead work on the content you already watch and read, and that is the category WordByWord belongs to: rather than supplying a library, it measures whatever you open against your own vocabulary, in any language you are learning.

What is the 15 30 15 method?

A study-session structure occasionally recommended for input practice: a short warm-up, a longer block of focused input, then a short review — the exact splits vary by whoever is recommending it. It is a scheduling technique, not a claim about acquisition, and it is unrelated to how coverage is measured.

Try it on the next thing you open

The shortest version of everything above: comprehensible input is not a genre you can shop for, it is a relationship between content and your vocabulary — and that relationship is now a number you can see before you spend an evening on the wrong video.

WordByWord is a language learning Chrome extension that does the measuring for you. Install it, open something in the language you are learning, and look at the number before you commit your evening to it.

References

  • Hu, M., & Nation, P. (2000). Unknown vocabulary density and reading comprehension. Reading in a Foreign Language, 13(1), 403–430. Article page

  • van Zeeland, H., & Schmitt, N. (2013). Lexical coverage in L1 and L2 listening comprehension: The same or different from reading comprehension? Applied Linguistics, 34(4), 457–479. Publisher page

  • Kremmel, B., Indrarathne, B., Kormos, J., & Suzuki, S. (2023). Unknown vocabulary density and reading comprehension: Replicating Hu and Nation (2000). Language Learning. Publisher page (open access)

  • Durbahn, M., Rodgers, M., & Peters, E. Lexical coverage in L1 and L2 viewing comprehension. Studies in Second Language Acquisition. Publisher page