A frequency-based word list is a corpus-derived ranking of words that shows which vocabulary gives you the most text coverage for the least effort. Instead of guessing which words matter, you study the ones that show up most often in real writing and speech. Canonical lists like NGSL and the Academic Word List turn that ranking into something you can actually download and use, and the right one depends on your level, your goals, and the kind of text you read.
TL;DR:
- The NGSL claims coverage of over 92% of general English texts, while the AWL and NAWL should supplement rather than replace a general list.
- Before downloading, confirm a list names its corpus, counting unit, tested coverage, revision date, and license; older or region mismatched sources can mislead.
- Test candidate lists against a passage from your target material, checking coverage among the top 1,000, 2,000, or 3,000 entries before committing.
- Frequency rankings miss idioms, phrasal verbs, word senses, and register, so verify meanings and common combinations through a dictionary or real reading.
- Research found that combining frequency with semantic links and teacher judgment produced shorter lists with 11% to 18% more coverage on comparable tests.
Table of Contents
- How frequency-based word lists work
- Canonical public lists and quick use cases
- How lists are built, validated, and the trust signals to look for
- How to choose the right list and use it in teaching or study
- Where to download lists and simple tools to analyze coverage
- Where frequency lists fall short
- Frequency lists versus thematic and semantic approaches
- Combining frequency lists with personalized study
- Lists set priorities, reading builds fluency
- Study frequency-priority words inside what you’re already reading
- FAQ
- Sources
How frequency-based word lists work
A frequency list ranks words by how often they appear in a corpus, a large collection of real text or speech. Two related ideas sharpen that ranking: range, which measures how many different texts or genres a word appears in, and dispersion, which checks whether a word’s occurrences cluster in one source or spread evenly. A word that appears often but only in one novel is less useful than one that shows up moderately often across hundreds of sources.
List makers also have to decide what counts as “one word.” A token is a single occurrence in the text; a type is a unique spelling. A lemma groups inflected forms (run, runs, running) under one headword, while a word family goes further, grouping derived forms (run, runner, runaway) together too.
That choice changes what makes the final list:
- Lemma-based counting treats “runs” and “running” as the same item as “run,” raising its combined frequency.
- Family-based counting pulls in derived words like “runner,” which can push the family’s rank even higher.
- Surface-form counting keeps every inflection separate, which tends to push common grammatical variants down the list even when the root word is extremely frequent.
Two lists built from the same corpus can rank words differently depending on which unit they use, so checking a list’s documentation before you trust its rankings matters.
Canonical public lists and quick use cases
Several public lists cover most learning and teaching situations:
- New General Service List (NGSL): a modern update to West’s original service list, built to provide broad coverage of everyday English, with the New General Service List stated to cover over 92% of most general English texts.
- General Service List (GSL): Michael West’s original 1953 list. It’s historically important and still shows up in older textbooks, but its source corpus predates modern usage, so treat it as background rather than a primary study tool.
- Academic Word List (AWL) and New Academic Word List (NAWL): both target the vocabulary that shows up across academic writing regardless of subject. Neither is a standalone study list: pair either one with a general list like NGSL so you cover everyday words alongside academic ones.
- Dolch sight words and Fry instant words: short, classic lists aimed at very young or beginning readers in English-medium classrooms. They’re built around high-frequency function words and common nouns rather than corpus coverage statistics, which makes them a poor fit for adult or academic learners.
- Wiktionary frequency lists and related corpus projects: community-maintained frequency data drawn from sources like film subtitles and web text, useful when you need frequency counts for a language or domain the major academic lists don’t cover.
The NGSL project also hosts companion lists built the same way: the New Academic Word List for academic vocabulary, a Business Service List for workplace English, and a TOEFL-oriented list for test preparation. Having all of them built on comparable methodology makes it easier to combine them without worrying about mismatched counting units.
How lists are built, validated, and the trust signals to look for
A trustworthy frequency list documents three things: the corpus it came from, the counting unit it uses, and the coverage it claims. Corpus composition matters because a list built only from news text will rank words differently than one built from spoken conversation or fiction. Strong lists mix written and spoken sources and specify which established corpus they drew from, such as the British National Corpus (BNC), the Corpus of Contemporary American English (COCA), or the Cambridge English Corpus.
A documented coverage threshold tells you how far a list goes toward comprehension. Nation’s research on vocabulary and coverage ties specific coverage percentages to specific reading and listening demands, and notes that corpus choice and counting unit both shift the resulting size estimate.
Before trusting a list, check for these signals:
- A named, described corpus, not just “millions of words” with no source.
- A stated counting unit (lemma, family, or surface form) and consistent use of it throughout.
- A coverage percentage tied to a defined test corpus, not a vague claim.
- A publication or revision date, since corpora more than a couple of decades old miss newer common usage.
Common pitfalls include lists built from British sources applied uncritically to American learners, outdated corpora that miss current vocabulary, and downloads with no methodology notes at all. A Cambridge Core replication study flags exactly this kind of corpus mismatch as a recurring problem in vocabulary research.
How to choose the right list and use it in teaching or study
Picking a list works best as a short, repeatable process rather than a one-time download.
- Define your target text and goal. Reading comprehension, academic writing, and spoken fluency draw on different vocabulary, so a list optimized for one may underperform for another.
- Run a quick coverage check. Take a sample of the material you actually want to understand and see how many of its words appear in the top 1,000, 2,000, or 3,000 entries of your candidate list. Tools built for this kind of comparison, such as the Range program referenced in frequency-based text studies, can automate the count, or you can tally manually for a short passage.
- Trim the list to a usable size. Cut entries irrelevant to your goal and add domain-specific terms the general list misses, especially technical or field-specific vocabulary.
For independent study, pair your trimmed list with extensive reading in your target language so the words appear again in natural context, and schedule review with spaced repetition rather than a single cram session.
For classroom use, break a trimmed list into lesson-sized chunks, recycle the same words across several lessons and assessments, and track which words each student has actually mastered rather than just introduced.
Pro Tip: Test a candidate list against a real sample of your target material before committing to it; a list that claims high general coverage can still underperform on a specific genre or domain.
Where to download lists and simple tools to analyze coverage
The NGSL project pages host direct downloads for NGSL and its companion lists, including NAWL, alongside documentation of the source corpus. Academic Word List downloads and Dolch and Fry lists circulate on various education sites, so check each download against the trust signals above before relying on it.
For broader or multilingual needs:
- Wiktionary’s frequency list pages host surface-form, lemma, and word-family counts across many languages.
- Leipzig Corpora and OpenSubtitles-based frequency collections offer downloadable counts outside the major English-teaching lists.
- Lightweight tools like Range or the open-source
wordfreqlibrary let you compare a text against a baseword list without specialized software; a spreadsheet works fine for short passages.
Before using any downloaded list, confirm it states its corpus, its license, and its counting unit.
Where frequency lists fall short
Frequency counts tell you how often a word appears, not what it means in context, and that gap matters most with idioms and multiword expressions. “Kick,” “bucket,” and “the” are all common words individually, but “kick the bucket” means something a word-by-word frequency count can’t capture. The same problem shows up with phrasal verbs, collocations, and set phrases: a list built on single tokens will rank “take,” “break,” and “down” highly without ever flagging “take a break” or “break down” as units learners need to recognize together.
Polysemy creates a related issue. “Bank” ranks high in a general corpus, but a learner studying finance needs the money sense, while one reading travel writing needs the riverbank sense. A frequency list gives you the word without telling you which meaning dominates in your material, so you still have to confirm the sense fits before you spend study time on it.
Register and tone are invisible to frequency counts too. A word can be common overall but inappropriate in formal writing, or vice versa, and no raw frequency number distinguishes between them.
None of this makes frequency lists useless, but it does mean they work best as a starting filter rather than a final answer. Treat a frequency list as a way to decide what to look up and learn first, then confirm meaning, register, and common collocations through a dictionary or real reading before you consider a word mastered.
Frequency lists versus thematic and semantic approaches
Frequency-based lists optimize for one thing: statistical coverage of running text. Thematic vocabulary lists, by contrast, group words by topic, like travel, food, or business, regardless of how often each word appears in a general corpus. Semantic cluster approaches group words by meaning relationships rather than frequency or topic, which can help learners build conceptual networks around a word rather than memorizing it in isolation.
Each approach solves a different problem. A frequency list gets you the highest statistical return per word studied, which makes it efficient for general comprehension goals. A thematic list is weaker on raw efficiency but stronger for situational readiness: if you’re traveling, “boarding pass” matters more than its frequency rank suggests, because you’ll need it in one specific, high-stakes moment. Semantic clustering supports retention by connecting new words to ones you already know, which can make recall easier even for words that wouldn’t rank highly on a pure frequency count.
The research on corpus-informed list design suggests the strongest lists blend these approaches rather than picking one. Combining raw frequency with connectedness, how a word links to others a learner already knows, and teacher judgment about classroom relevance produced lists that were shorter than older frequency-only lists while still delivering 11 to 18% more coverage for comparable test corpora. In practice, that means a frequency list makes a strong backbone, but layering in thematic or semantic groupings around your specific goals can often improve results compared to using either method alone.

Combining frequency lists with personalized study
A frequency list tells you what’s statistically worth learning across English in general, but it says nothing about what you personally already know or need. The strongest study approach starts with a frequency list to set priorities, then personalizes from there based on your own reading, interests, and gaps.
A practical version of this looks like a layered system. Start with a trimmed frequency list as your baseline target. Add words you actually encounter in material you choose to read or watch, since those carry built-in context and motivation that a generic list entry doesn’t. Remove words you already know well, even if they rank highly on the source list, so your remaining study time goes to genuine gaps rather than review of mastered vocabulary.

Spaced repetition scheduling works well on top of this kind of personalized set because it adjusts review timing to how well you know each word rather than treating every entry the same way. A word you learned a year ago and still recall instantly needs far less review time than one you keep forgetting, and a system that tracks that difference saves significant study time over flat, unscheduled review.
The combination matters because pure frequency study can feel abstract: a list of 2,000 words with no connection to anything you’re reading is hard to sustain. Personalizing around real content keeps the words tied to something you actually care about finishing, which tends to matter more for long-term retention than theoretical coverage percentages ever do.
Lists set priorities, reading builds fluency
Frequency lists are a prioritization tool, not a curriculum. They tell you which 2,000 or 3,000 words give you the best return, but a list sitting in a spreadsheet never taught anyone a language. The real work happens when those words show up again in something you’re actually reading or watching, because repeated, meaningful exposure is what turns a ranked entry into a word you reach for without thinking.
The lists in this guide work best as a filter you apply before you read, not a substitute for reading itself. Combine a trimmed frequency list with content you’d read anyway, and let the words you already prioritized surface naturally as you go.
— WordByWord Team
Study frequency-priority words inside what you’re already reading
We built our Chrome extension and web app to put frequency-list thinking directly into the browsing you already do. Spotlight mode highlights unknown words on any page, PDF, or YouTube video, so prioritized words from a frequency list stand out the moment they appear in real content, and a comprehension level check tells you upfront what share of a text you already know before you commit to reading it.
- Collect the words that matter into your own set, built by hand, generated by AI from a prompt, or imported as a CSV file from Anki, Quizlet, a spreadsheet, or a frequency list you already trust.
- Review those words with flashcards and eight training modes, including typing and listening, scheduled by spaced repetition so mastered words fade from the highlights automatically.
Start on our Free Forever plan and upgrade to Premium for $5.99 a month whenever you want the monthly caps on AI trainings and AI generation removed.
FAQ
What are the most common English words by frequency?
The most frequent English words are overwhelmingly function words: “the,” “be,” “to,” “of,” and “and” dominate the top of nearly every corpus-based list regardless of genre. Content words like “time,” “person,” and “year” typically follow close behind in general-purpose corpora such as COCA and the BNC.
What is the second most common word in English?
Across most major English corpora, “be” (in its various forms) ranks as the second most frequent word, right behind “the.” Exact ranking can shift slightly between corpora depending on whether spoken or written text dominates the sample.
Can you give me a list of low-frequency words?
Low-frequency words are typically specialized, technical, or rare terms that fall far down a corpus-based ranking, such as niche scientific vocabulary or archaic words. Rather than a single fixed list, you can generate one by taking any frequency list and looking at ranks beyond the top 5,000 to 10,000 entries, since what counts as “low frequency” depends on which corpus and cutoff you use.
Is there an English frequency dictionary available?
Yes, several frequency-based resources function like dictionaries, including the New General Service List and its companion lists, along with open frequency data hosted on Wiktionary. These resources rank words by corpus frequency rather than defining them alphabetically, so they work best alongside a standard dictionary rather than in place of one.




