logo

Vocabulary Over Memory: L2 Reading Comprehension for Researchers and Teachers

·by WordByWord Team
23 min read
Vocabulary Over Memory: L2 Reading Comprehension for Researchers and Teachers

The strongest predictors of reading comprehension in L2 are language knowledge variables, not general cognitive ones: vocabulary, syntactic knowledge, and L2 listening comprehension consistently outweigh working memory and metacognition in meta-analytic reviews. Jeon and Yamashita’s landmark synthesis found L2 knowledge explains more variance than language-general skills, and a 2024 secondary meta-analysis confirms comprehension skills matter more than decoding. Multi-strategy instruction, meanwhile, carries one of the largest effect sizes in applied linguistics. The sections below unpack the evidence, the component processes, the intervention data, and what all of it means for measurement and teaching.


TL;DR:

  • Vocabulary, syntactic knowledge, and listening comprehension explain more variance in L2 reading than working memory or metacognition, with comprehension skills outweighing decoding.
  • Multi-strategy instruction has a large average effect size of 0.91, but outcomes vary depending on the learner level, instructor type, and intervention duration.
  • Vocabulary breadth strongly predicts overall comprehension, while depth, syntax, and semantic networks influence inferential and literal understanding separately.
  • Word-to-text integration components independently predict literal and inferential comprehension, making separate assessment important for targeted instruction.
  • Effective strategies include activating prior knowledge, self-questioning, and matching text topics to learners’ background, with digital tools like WordByWord supporting vocabulary retention and difficulty assessment.

WordByWord
Build Vocabulary From Real Reading
WordByWord helps you translate unfamiliar words, save them to collections, and retain them through spaced repetition while reading.
Explore WordByWord

Table of Contents

Meta-Analytic Evidence: What Explains Variance in L2 Reading

Ask a room of second language reading researchers what drives comprehension, and you’ll get an argument about decoding versus knowledge. The data has largely settled it in favor of knowledge. Jeon and Yamashita’s 2014 meta-analysis, still the most cited synthesis in the field, aggregated correlations across dozens of studies and found that L2-specific knowledge variables, vocabulary, grammar, and morphological awareness, correlate with reading comprehension at a noticeably higher rate than language-general variables like working memory or metacognitive awareness.

That distinction matters for how researchers design studies and how educators allocate instructional time. If working memory only weakly predicts comprehension once vocabulary is accounted for, then interventions built around memory training are chasing the wrong lever. The 2024 replication of this work, published as a secondary meta-analysis in Studies in Second Language Acquisition, pushed the finding further by directly testing the Simple View of Reading in an L2 context. It found L2 comprehension skills contribute substantially more to reading outcomes than decoding skills, reinforcing that comprehension, not word recognition, is where the larger share of variance lives for most L2 samples.

Effect sizes from the intervention side tell a complementary story. Where correlational meta-analyses describe what predicts comprehension, the 46-study meta-analysis of strategy interventions describes what changes it.

Statistic Callout: Multi-strategy reading instruction produces a mean effect size of g = 0.91 across 46 studies, a large effect by any conventional benchmark in educational research.

That figure deserves context. An effect size near 0.91 sits well above what Cohen’s conventions typically label “large” (0.8), and it dwarfs effects reported for many other classroom interventions in second language acquisition. But heterogeneity across those 46 studies is real: samples ranged from university EFL learners to secondary students, and moderators like instructor type and intervention length shifted outcomes considerably (more on that in the strategies section below).

What ties these findings together:

  • L2 knowledge variables (vocabulary, grammar, morphology) outpredict general cognitive variables in correlational research.
  • Comprehension-level skills explain more variance than decoding in most adult and university L2 samples.
  • Multi-strategy instruction produces large average gains, though effect sizes vary by context.
  • Working memory and metacognition remain measurable but secondary correlates once knowledge variables are controlled.

For researchers designing new studies, the practical takeaway is straightforward: knowledge variables deserve primary billing in any model of L2 reading comprehension, and general cognitive measures belong as covariates, not headline predictors.

The Linguistic and Cognitive Building Blocks Behind Comprehension

Vocabulary does the heaviest lifting, but not all vocabulary knowledge lifts equally. Breadth (how many words a reader knows) and depth (how well they know each word’s nuances, collocations, and polysemy) predict different outcomes. Breadth measures, typically vocabulary size tests, tend to correlate strongly with overall comprehension scores. Depth measures track more specifically with inferential comprehension, the kind of reading that requires connecting a word’s connotation to an implied meaning rather than its dictionary definition.

Syntactic knowledge works differently. Regression models consistently show that syntactic parsing ability, the capacity to correctly assign grammatical roles within a sentence, predicts literal comprehension even after vocabulary and working memory are statistically controlled. The effect is modest in raw size but statistically reliable, which tells you syntax is doing real, independent work rather than riding on vocabulary’s coattails.

Semantic network knowledge, how densely and accurately a reader’s mental lexicon connects related concepts, plays a different role again. It shows up most clearly in inference generation and in how well a reader builds a coherent mental model across sentences rather than parsing any single one correctly.

Decoding and word recognition, the skills most associated with beginning reading instruction, show smaller incremental contributions in many L2 samples, particularly among adult and university learners who have already automatized basic word recognition in their L2. Working memory and metacognition round out the list as real but comparatively modest correlates.

A quick map of what each component tends to predict:

  • Vocabulary breadth: overall comprehension scores across proficiency levels.
  • Vocabulary depth: inferential comprehension specifically.
  • Syntactic parsing: literal comprehension, independent of vocabulary.
  • Semantic network knowledge: inference generation and cross-sentence coherence.
  • Decoding/word recognition: a necessary base skill with diminishing marginal returns once automatized.
  • Working memory and metacognition: secondary correlates, more useful as moderators than as primary predictors.

Word-to-Text Integration: Why Literal and Inferential Comprehension Need Different Skills

Word-to-text integration, often abbreviated WTT integration in the research literature, describes the process by which a reader connects an individual word’s meaning to the broader sentence and passage context. It has two measurable components: syntactic parsing (structural integration) and semantic network knowledge (associative integration). These aren’t just theoretical categories. They predict genuinely different comprehension outcomes, which is the part most instructional guides skip.

Frontiers in Education’s 2022 study tested both components directly and found syntactic parsing predicts literal comprehension, while semantic network knowledge uniquely predicts inferential comprehension, even after controlling for vocabulary and working memory. Neither component was redundant with the other; each carried independent predictive weight.

The magnitude is worth stating honestly rather than inflating it. Syntactic knowledge tends to explain a small but statistically significant slice of variance in literal comprehension, often in the 1 to 3 percent range after vocabulary and working memory are accounted for, while semantic network measures add a similarly modest but unique contribution to inferential comprehension. Small numbers, but they’re independent contributions on top of everything else already in the model, which is exactly what makes them theoretically interesting even when they’re practically modest.

Pro Tip: If you’re designing an intervention study, decide up front whether your outcome measure targets literal recall, inferential reasoning, or both. A test that blends the two into a single composite score will mask which skill your intervention actually moved.

This has direct implications for test design and classroom instruction:

  • A test heavy on detail-recall items is mostly measuring syntactic parsing and vocabulary, not deeper comprehension.
  • A test built around “what does the author imply” items is leaning on semantic network knowledge and inference skills.
  • Instruction aimed at literal comprehension gains (skimming for facts, answering “who did what”) should emphasize sentence-level grammar work.
  • Instruction aimed at inferential gains benefits more from semantic mapping, discussion of connotation, and building associative vocabulary networks.

Researchers reporting a single composite comprehension score without breaking out literal and inferential subscales are, in effect, averaging over two different cognitive processes with two different sets of predictors. That’s a measurement choice with real consequences for what a study can and cannot conclude.

Strategy Interventions: What Works, By How Much, and For Whom

Not all reading strategies earn their place in the curriculum. The 46-study meta-analysis that anchors this discussion reported a mean effect size of g = 0.91 for multi-strategy interventions, but that average conceals considerable variation in which specific strategies did the heavy lifting.

Statistic Callout: Among the strategies examined, those tied to connecting new information with prior knowledge and generating self-questions produced relatively stronger gains, while visualization techniques showed weaker or inconsistent effects in several L2 contexts.

Three findings stand out for anyone designing an intervention or a curriculum:

  1. Prior-knowledge activation wins consistently. Strategies that explicitly link new text content to what readers already know outperformed strategies taught in isolation, likely because they reduce the cognitive load of building a mental model from scratch.
  2. Self-questioning generates durable gains. Asking readers to pose their own questions about a text, rather than only answering questions someone else wrote, appears repeatedly among the higher-effect interventions, plausibly because it forces active monitoring rather than passive scanning.
  3. Visualization underperforms in several L2 samples. Strategies asking readers to form mental images of text content showed weaker or inconsistent effects, possibly because visualization demands cognitive resources that L2 readers are already spending on basic linguistic processing.

Moderators matter as much as the strategies themselves. Instructor type made a measurable difference: non-standard or specialist instructors sometimes produced larger effect sizes than standard classroom teachers, which points to implementation fidelity, not just strategy choice, as a driver of outcomes. Intervention duration and pedagogical approach (explicit, teacher-led instruction versus more discovery-based approaches) also shifted results across studies, and few of the 46 studies tracked retention beyond the immediate posttest, a gap, covered further in the measurement section below.

Collaborative Strategic Reading (CSR), which combines several of these high-performing strategies into a structured, peer-supported routine, has its own empirical support for improving comprehension, motivation, and metacognitive awareness in quasi-experimental EFL studies, making it one of the more evidence-backed packaged approaches available to classroom teachers.

Measurement Choices That Make or Break L2 Reading Research

Get the measurement wrong and even a well-designed intervention will produce misleading results. The single most common design flaw in L2 reading research is collapsing literal and inferential comprehension into one score, which erases the very distinction the WTT integration research shows matters most.

Researchers should report literal and inferential subscales separately, along with the correlation between them, rather than defaulting to a single composite. Doing so lets readers of the study see whether an intervention moved surface-level recall, deeper inference, or both, instead of guessing from an averaged number.

Delayed posttests are the field’s biggest blind spot. Most strategy-intervention studies measure gains immediately after training, when novelty and short-term memory inflate scores, and comparatively few return weeks later to check whether the gains held. That scarcity matters because an effect size like g = 0.91 measured immediately after instruction says little about whether readers still use the strategy a month later, and instructional decisions built on immediate-posttest data risk overstating durability.

A few recurring validity threats deserve routine checking in any new study design:

  • Proficiency mismatches: pooling beginner and advanced L2 readers in one analysis can mask moderator effects that only show up within a narrower proficiency band.
  • Ceiling and floor effects: passage-level tests that are too easy or too hard for a given sample compress variance and understate true relationships.
  • Unreported item characteristics: studies that don’t specify whether items test literal recall or inference make replication and comparison across studies nearly impossible.
  • Missing retention data: without delayed posttests, an intervention’s practical value to real classrooms remains an open question.

Turning the Evidence Into Classroom and Self-Study Practice

Evidence without application is just an interesting paper. For educators and advanced learners, the research above points toward a fairly specific set of priorities rather than a vague call to “read more.”

Build L2 knowledge first, and teach strategies explicitly on top of it, rather than assuming strategies alone will compensate for thin vocabulary or weak listening skills. Vocabulary and L2 listening comprehension are the strongest correlates in the research, so instructional time spent there tends to pay off faster than time spent on general strategy drills. Multi-strategy instruction, and structured approaches like Collaborative Strategic Reading, should sit alongside that knowledge-building, not replace it.

Scaffolding needs to shift with proficiency. Lower-proficiency readers benefit more from prior-knowledge activation and explicit vocabulary support before a text; higher-proficiency readers get more mileage from inference-focused strategies and self-questioning, since their word-to-text integration skills are already more developed.

  • Pair explicit strategy instruction with vocabulary and listening practice rather than treating them as separate tracks.
  • Use extensive and intensive reading approaches deliberately, matching each to the comprehension outcome you’re targeting.
  • Track literal and inferential gains separately across an instructional cycle, not just a single composite score.
  • Retest after a delay to check whether strategy gains persist, rather than relying only on immediate posttests.

Pro Tip: When choosing which vocabulary to prioritize, focus on words that recur across multiple texts a learner is likely to encounter rather than rare items from a single passage. Frequency and relevance beat novelty for retention.

How WordByWord’s Tools Align With the Research

Digital tools rarely make it into a reading comprehension meta-analysis, but the mechanisms behind them map onto the same predictors the research keeps surfacing. Vocabulary knowledge is the strongest correlate in the literature, and any tool that helps a reader see and retain unfamiliar words in the flow of authentic reading is addressing the highest-leverage variable available.

WordByWord’s Spotlight mode highlights vocabulary directly on the page by mastery level, which mirrors the instructional principle of making unfamiliar-versus-known distinctions visible rather than leaving learners to guess. Its comprehension-level estimate, showing what share of a text’s words a reader already knows before starting, gives a rough, practical proxy for the kind of vocabulary-breadth measure researchers use to predict comprehension difficulty. And its spaced repetition flashcards, built to move vocabulary from unknown to mastered over time, operationalize the retention question the delayed-posttest research keeps raising.

A few resources worth reading alongside this article:

Background Knowledge and Cultural Familiarity

A reader who already knows the subject matter of a text will out-comprehend a linguistically stronger reader who doesn’t, and this shows up repeatedly across receptive skills. Prior-knowledge activation strategies rank among the most effective interventions in the strategy meta-analysis discussed earlier, which is itself indirect evidence for how much background knowledge matters once a reader engages a text.

The effect isn’t limited to reading. A 2026 study on listening comprehension found topic familiarity often explains more variance in comprehension outcomes than strategy use itself, a pattern that plausibly extends to reading given how closely the two receptive skills share underlying comprehension mechanisms.

Cultural familiarity compounds the effect. A text referencing unfamiliar customs, institutions, or historical events demands that a reader build schema from scratch while simultaneously parsing unfamiliar language, a double cognitive load that vocabulary tests alone don’t capture. This is why comprehension scores on culturally distant passages tend to underestimate a learner’s “true” linguistic comprehension ability, and why researchers comparing L2 reading across different text topics should treat topic familiarity as a variable to control, not an incidental detail.

For instructional design, the implication is direct: matching text topics to a learner’s existing knowledge, at least early in a course, isolates linguistic comprehension gains from background-knowledge effects. Later, deliberately introducing culturally unfamiliar material becomes a way to build the schema-building skills that real-world reading eventually demands.

Alphabetic Versus Non-Alphabetic Scripts: A Different Kind of Challenge

Not every L2 reading challenge is linguistic. Learners moving from an alphabetic L1 (English, Spanish, German) into a non-alphabetic L2 script (Chinese characters, Japanese kanji) face a decoding hurdle that has no real equivalent for same-script language pairs. Word recognition in a logographic system depends on visual pattern recognition and character memorization rather than phoneme-to-grapheme mapping, which means decoding automaticity often takes measurably longer to develop.

Comparison of alphabetic and logographic decoding

This matters for how the vocabulary-versus-decoding balance discussed earlier plays out differently across script types. In alphabetic-to-alphabetic L2 learning, decoding tends to automatize relatively quickly, leaving vocabulary and syntax as the dominant long-term predictors. In alphabetic-to-logographic learning, decoding remains a meaningful bottleneck for longer, and comprehension research on these learner populations needs to account for character knowledge as a variable in its own right, not fold it into a generic “vocabulary” measure.

The reverse direction carries its own friction. Learners moving from a logographic L1 into an alphabetic L2 sometimes show underdeveloped phonological awareness relative to same-script peers, since their L1 reading system never required the same grapheme-phoneme mapping skills. Reading fluency researchers increasingly argue that script distance between L1 and L2 deserves explicit treatment as a moderator variable in comprehension studies, rather than being averaged away across mixed-script samples.

Practically, this means comprehension benchmarks and normative comparisons should rarely cross script families without adjustment. A vocabulary-size test calibrated on alphabetic L2 learners may not transfer cleanly to logographic-script learners, whose developmental timeline for basic word recognition looks fundamentally different.

Motivation and Affective Factors in L2 Reading Development

Comprehension research tends to focus on cognition, but affect shapes how much cognitive effort a reader is willing to spend. Motivation influences how long a reader persists through a difficult passage, how often they reread confusing sections, and whether they engage strategies like self-questioning at all, which connects it directly to the strategy-effectiveness findings covered earlier.

Collaborative Strategic Reading studies found measurable gains not just in comprehension scores but in reading motivation and metacognitive awareness together, suggesting the affective and cognitive gains move as a package rather than independently. That’s a meaningfully different claim than saying motivation simply correlates with comprehension. It suggests well-designed strategy instruction builds motivation as a byproduct, not just as a precondition.

Anxiety works in the opposite direction. L2 reading anxiety, often tied to fear of misunderstanding culturally unfamiliar content or unfamiliar scripts, tends to suppress the exact strategic behaviors, rereading, self-questioning, inference-checking, that the intervention research identifies as highest-value. A learner too anxious to reread a confusing sentence loses access to one of the cheapest, most effective comprehension repair strategies available.

For instructional design, this argues for choosing texts and tasks that build early success experiences before introducing higher-difficulty or culturally distant material, since early failure experiences can suppress strategic engagement well beyond the specific text that caused them. Framing reading tasks as low-stakes practice, rather than as assessment, appears to preserve the willingness to use effortful strategies like self-questioning and rereading.

How L2 Reading Comprehension Develops Over Time

L2 reading comprehension doesn’t develop on a fixed timeline the way some curricula imply, and the predictors that matter shift as proficiency grows. Early-stage L2 readers, still automatizing decoding and building a base vocabulary, show comprehension scores driven heavily by word recognition speed and basic vocabulary breadth. Syntax and semantic network effects, the WTT integration components discussed earlier, become more detectable once decoding stops consuming most of a reader’s cognitive bandwidth.

This has a direct implication for longitudinal study design: a predictor model that fits an intermediate-proficiency sample well may fit a beginner sample poorly, because the beginner group hasn’t yet automatized the skills the model assumes are already in place. Researchers pooling learners across a wide proficiency range without testing for interaction effects risk averaging over genuinely different developmental stages.

Development also isn’t strictly linear. Plateaus are common, particularly around the point where basic decoding has automatized but vocabulary depth and syntactic complexity in target texts start outpacing a learner’s existing knowledge. Progress at this stage tends to depend more on deliberate vocabulary expansion and exposure to syntactically varied input than on additional decoding practice, which is often where instruction mistakenly continues to focus.

Retention adds another timeline dimension entirely. Because so few strategy-intervention studies include delayed posttests, the field has a limited empirical picture of how long comprehension gains from any given intervention actually last, an open question worth flagging again for anyone designing a new longitudinal study.

L1 Reading Skills and Transfer Effects on L2 Comprehension

A learner’s L1 reading ability doesn’t disappear when they start reading in a second language. It transfers, sometimes helpfully, sometimes in ways that complicate L2 comprehension development. Strong L1 reading skills, particularly strategic reading habits like self-questioning and monitoring, tend to transfer to L2 reading once the learner has enough L2 linguistic knowledge to apply them, a pattern often described as the threshold hypothesis in second language reading research.

L1-mediated scaffolding can be leveraged deliberately rather than left to chance. An experimental study on L1-assisted reading found that the timing and type of L1 pre- or post-reading questions differentially affected reading speed and integrative comprehension, showing that L1 support isn’t uniformly helpful. Pre-reading detail questions in the L1 sped up reading but didn’t necessarily deepen integrative comprehension, while other prompt types shifted the balance differently, which argues for treating L1 scaffolding as a design variable with real trade offs rather than a blanket good practice.

Script distance between L1 and L2, discussed earlier in the context of alphabetic versus non-alphabetic systems, is really a special case of this broader transfer question. Where L1 and L2 share an alphabetic base, phonological awareness and print-concept knowledge transfer relatively cleanly. Where they don’t, transfer is more limited, and L2 decoding development proceeds on a more independent track from L1 reading ability. Either way, a learner’s L1 reading profile is a variable worth measuring, not assuming away, in any serious model of L2 comprehension.

Research Perspective: Where the Field Should Go Next

The strategy-intervention literature has an effect-size problem hiding behind its impressive average. A g of 0.91 sounds definitive until you notice how few of the 46 underlying studies checked whether the gains survived past the final training session. That gap deserves more attention than it gets: the field needs more delayed-posttest designs, replicated across genuinely diverse samples, not just more immediate-posttest studies confirming what’s already well established.

Standardized reporting would help more than another single-sample study. Literal and inferential subscales, vocabulary depth (not just breadth) measures, and word-to-text integration indices should become default reporting practice, not an occasional addition. Right now, comparing findings across studies means guessing whether “comprehension” meant the same thing in each one.

The most promising direction may be hybrid: trials testing explicit strategy instruction combined with digital vocabulary tools, measuring whether blended pedagogies outperform either approach alone.

— WordByWord Team

Practice What the Research Recommends With WordByWord

WordByWord is built around the exact predictors this research keeps pointing to: vocabulary, in-context reading, and retention over time, rather than a fixed curriculum disconnected from what you actually read. The Chrome extension lets you translate unknown words in one click across websites, PDFs, and YouTube subtitles in over 50 languages, and Spotlight mode highlights vocabulary on the page by mastery level, exactly the kind of visible unknown-versus-known distinction the strategy research shows supports comprehension.

WordByWord

The comprehension-level estimate tells you what share of a text’s words you already know before you start reading it, a quick way to judge whether a passage matches your current level instead of guessing. Vocabulary you collect gets reinforced through the spaced repetition flashcard trainer, with seven training modes covering typing, listening, and sentence building, which addresses the retention gap the intervention literature keeps flagging as underreported. The Free Forever plan costs $0 per month with usage limits, and Premium removes them for $5.99 per month. Install the extension and start building your comprehension-level baseline on the next article you read.

Sources

FAQ

What Is L2 Reading in Language Research?

L2 reading refers to reading comprehension in a language learned after a person’s first language, typically studied in contexts like university EFL programs or immersion settings. Research treats it as a distinct skill set from L1 reading because it depends heavily on L2-specific knowledge such as vocabulary and syntax, more than on general cognitive ability.

What Are the Signs of Poor Reading Comprehension in L2?

Common signs include heavy reliance on word-by-word translation, difficulty answering inference-based questions even when literal recall is fine, and slow reading speed that doesn’t improve with familiar vocabulary. A learner who understands individual sentences but can’t summarize a passage’s overall argument is often showing a word-to-text integration gap rather than a vocabulary gap, since semantic network knowledge, not vocabulary alone, drives inferential comprehension.

What Do L1, L2, and L3 Mean in Language Education?

L1 refers to a person’s first or native language, L2 to a second language learned later, often in a classroom or immersion setting, and L3 to a third language, which sometimes shows transfer effects from both L1 and L2 rather than from L1 alone. These labels matter in research because a learner’s L1 reading skills can transfer to L2 reading once enough L2 linguistic knowledge is in place, a pattern documented in transfer and threshold research.

How Can Tools Like WordByWord Support L2 Reading Comprehension Practice?

Vocabulary knowledge is the single strongest correlate of L2 reading comprehension in the meta-analytic evidence, so tools that surface unfamiliar words during real reading target the highest-leverage variable available. WordByWord’s Spotlight mode and comprehension-level estimate let learners see unfamiliar vocabulary and gauge text difficulty before reading, while the spaced repetition flashcard trainer supports the retention question that strategy-intervention research often leaves unmeasured.

Does Improving Vocabulary Alone Fix L2 Reading Comprehension Problems?

No. Vocabulary is the strongest single predictor, but syntactic parsing and semantic network knowledge each add independent predictive power for literal and inferential comprehension respectively, even after vocabulary is statistically controlled. A learner can know every word in a passage and still struggle with inference if their semantic network knowledge or syntactic parsing skills lag behind their vocabulary size.

Request a feature

What should we add or fix? We read every message and send a little gift for the idea.

0 / 4000