The concept explainer
How many English words do I know? What a vocabulary estimate really means
Learn what vocabulary-size tests actually count, why different tests give different totals, and how to interpret an English vocabulary estimate without overclaiming.
An idea, brought into focus.
08 sources, open to exploreIn this article 16 sections
An English vocabulary estimate is an inference about how many lexical items you would meet a particular knowledge threshold for, based on a sample of test questions. It is not a literal inventory of every word stored in your memory. The number only becomes meaningful when you know two things: what the test counts as a word and what the test accepts as knowing one.
Two vocabulary tests can give you different totals on the same afternoon without either test being obviously broken. One may count word families, another lemmas. One may ask you to recognize a meaning from four options, another may make you retrieve the meaning with no choices at all. A score that looks like one simple number is really the end of several measurement decisions.
If you want to try VocaNum before reading the details, take the free English vocabulary test. Then come back to the result with a better question than “Is this number exact?” Ask: What evidence does this number summarize?
First decide what counts as a word
“English has X words” sounds precise until you try to decide where one word ends and another begins.
Consider run, runs, ran, and running. A dictionary may group those under one headword. A corpus can count them as separate written forms. Vocabulary research may collapse them into one lemma. A broader word family may also group transparently related derived forms, depending on the counting rules being used.
Now add runner, rerun, run-down, and an expression such as run out of. The total changes again depending on the unit.
That is why vocabulary-size figures should never be compared before checking their counting unit.
The same vocabulary can produce different totals
| Counting unit | Rough idea | What happens to the total? |
|---|---|---|
| Word form | Counts surface forms such as run, runs, ran, running separately | Usually produces a larger number |
| Lemma | Groups inflected forms that belong to the same basic lexical item | Produces a smaller, more consolidated total |
| Word family | Groups a base word with some related derived forms according to defined rules | Can produce a much smaller total than counting forms or lemmas |
| Multiword expression | Treats expressions such as by the way or run out of as lexical units when the method includes them | Adds units that single-word counts may miss |
These are not interchangeable labels for the same thing. They are different ways of defining the population being counted.
One study shows how large the definition effect can be
A useful illustration comes from Brysbaert, Stevens, Mandera, and Keuleers (2016). Their estimates for a typical 20-year-old native speaker of American English were about 42,000 lemmas plus about 4,200 non-transparent multiword expressions, derived from roughly 11,100 word families.
Those figures are not a target that English learners should compare themselves against. The study concerned native speakers, used its own operational definitions, and explicitly noted that “knowledge” could be shallow — as little as knowing that a word exists in some parts of the analysis.
The important lesson is the gap between 42,000 lemmas and 11,100 word families. The speaker did not suddenly lose most of their vocabulary when the unit changed. The ruler changed.
Then decide what “knowing” means
Even after the counting unit is fixed, vocabulary knowledge is not a single switch marked known / unknown.
Paul Nation's widely used framework describes knowledge of a word across form, meaning, and use, including spoken and written form, word parts, form–meaning connections, concepts and associations, grammatical behavior, collocations, and constraints on use. Each aspect can also be stronger receptively than productively. Nation's discussion of word knowledge makes the central problem clear: there are many things to know about one word, and many degrees of knowing them.
Vocabulary size therefore answers a narrower question than vocabulary depth. In his review of the research, Norbert Schmitt (2014) argues that breadth and depth are related but not identical constructs; measures of depth can add information beyond a count of how many words meet a threshold.
A vocabulary-size test has to choose a slice of that larger construct. Otherwise one word could require a miniature exam of its own.
One word, several levels of evidence
Take reluctant.
You might show four different kinds of knowledge:
- Familiarity: reluctant looks like a real English word you have seen before.
- Recognition: when shown several meanings, you choose “unwilling or hesitant.”
- Contextual understanding: you understand She was reluctant to agree without more evidence without translating the sentence word by word.
- Productive use: when you need the idea yourself, you retrieve reluctant and use it naturally in speech or writing.
A test that verifies level 2 has learned something useful about your vocabulary. It has not automatically proved level 4.
That distinction is the reason a recognition estimate can be informative without pretending to measure every aspect of lexical knowledge. For a deeper look, read recognition vs recall: why recognizing a word is not the same as retrieving it.
Why vocabulary tests use samples
Testing every item in a large lexical database would be exhausting and, for a learner, mostly pointless. Vocabulary-size tests therefore sample from a defined vocabulary population and use performance on that sample to estimate knowledge beyond the exact items shown.
A classic example is Nation and Beglar's Vocabulary Size Test. Its original 140-item form sampled ten items from each of fourteen 1,000-word-family frequency levels. In that design, each correct item contributed to an estimate of knowledge in a much larger band rather than simply adding “one known word” to a literal count.
That kind of sampling is what makes a short vocabulary-size test possible. It is also what makes the result an estimate.
A sampled item does not mean “100 words entered your brain”
When a test scales a sample to a larger vocabulary population, the multiplier belongs to the measurement model, not to the individual question.
Answering one sampled word correctly does not prove that you know every neighboring word in its frequency band. Likewise, missing one item does not prove that an entire group of words is unknown. The usefulness of the final estimate depends on how representative the sample is, how the items behave, and how uncertainty is handled.
This matters most when people treat a vocabulary score as an exact inventory rather than as an inference from evidence.
Item format changes what a correct answer means
Vocabulary assessment research has repeatedly shown that item format affects score interpretation.
The original Vocabulary Size Test was designed to measure written receptive vocabulary knowledge, and a Rasch-based validation found strong reliability and generally good item behavior in the studied sample (Beglar, 2010). That does not mean every possible interpretation of a multiple-choice score is equally justified.
A later review by Stoeckel, McLean, and Nation highlights several limitations of written receptive size and levels tests. Recognition formats can allow correct answers through guessing or test-taking strategies, and recognizing a meaning among choices generally requires a lower threshold of knowledge than recalling that meaning without options. The review also questions the assumption that knowing one tested member necessarily proves knowledge of every member of a word family.
Research comparing formats reaches the same practical conclusion: the task is part of the score. Kremmel and Schmitt (2016) examined how different item formats relate to learners' ability to employ words, while Stewart (2014) focused specifically on whether multiple-choice options can inflate Vocabulary Size Test estimates through guessing.
This does not make multiple-choice recognition useless. It means the correct label is something like recognition-based estimate, not “proof that every counted word is instantly available in conversation.”
The question changes the evidence
| Task | What the learner sees | What a correct answer supports |
|---|---|---|
| Meaning recognition | Target word plus possible meanings | The learner can identify an appropriate meaning with cues available |
| Meaning recall | Target word, but no answer choices | The learner can retrieve a meaning with fewer external cues |
| Productive recall | Meaning, situation, or communicative need without the target word | The learner can retrieve the word form itself |
| Natural use | A real speaking or writing context | The learner can retrieve and use the word with appropriate grammar, collocation, and register |
These tasks overlap, but they are not equivalent. A useful test says which one it is using.
Why two good vocabulary tests can disagree
If one test says 4,800 and another says 6,200, the first question should not be “Which one is lying?” Check whether they are estimating the same construct.
Differences can come from:
- Counting unit: forms, lemmas, flemmas, or word families.
- Reference list: different corpora and frequency lists define a different vocabulary population.
- Frequency range: one test may focus on high-frequency vocabulary while another samples much further into low-frequency words.
- Sampling rate: a test with relatively few items is making a larger inference from each response.
- Item format: recognition, recall, matching, yes/no, translation, or production place different demands on the learner.
- Knowledge threshold: one test may accept partial meaning recognition where another expects stronger recall.
- Language and modality: written receptive vocabulary is not identical to spoken, productive, or context-rich vocabulary knowledge.
- Scoring model: guessing correction, “I don't know” behavior, confidence, and adaptive selection can change how raw evidence becomes a final estimate.
The number is inseparable from those design choices.
The most tempting interpretation is the wrong one
Myth: “My score is the exact number of English words in my memory.”
No vocabulary test can open memory and count lexical entries one by one. A practical test samples behavior and generalizes from it.
Reality: “My score is an estimate under a stated definition of word knowledge.”
That is less dramatic, but more useful. Once the definition is clear, you can compare later results using the same method, locate uncertainty, and decide what deserves practice.
Myth: “A larger number automatically means better English in every situation.”
Vocabulary breadth matters, but speaking, listening, reading, writing, grammar, fluency, pronunciation, strategic competence, and depth of word knowledge do not collapse into one vocabulary total.
Reality: “Vocabulary size is one measurement inside a larger language profile.”
Treat it as one strong signal, not the whole dashboard.
Do not convert a vocabulary total directly into a CEFR certificate
The CEFR describes language ability across a much broader system of communicative activities and competences. Its Companion Volume includes vocabulary range, but also grammatical accuracy, phonological control, reception, production, interaction, mediation, and other dimensions. The Council of Europe's CEFR descriptors are built around what users can do with language, not around one universal table that turns a vocabulary count into A1, B2, or C1.
A study may find relationships between vocabulary measures and proficiency in a particular population. That is different from saying “X words equals B2” for everyone, across every test.
How to read your own estimate
Instead of asking whether the final number is perfectly exact, ask five more useful questions.
1. What is the counting unit?
If the result counts word families, do not compare it directly with a result reported in lemmas or surface forms.
2. What kind of knowledge was tested?
Was the task recognition, recall, productive use, or a mixture? A recognition score should be interpreted as recognition evidence.
3. What vocabulary population was sampled?
A test built from frequency-ranked general English vocabulary answers a different question from a specialist academic or technical vocabulary test.
4. How much uncertainty sits behind the number?
Guessing, hesitation, inconsistent answers, and sparse sampling all matter. A result should not become more precise in your head than the method deserves.
5. What decision will the estimate improve?
The most useful score changes what you do next: which words to review, which words are stable, and where more evidence is needed.
In VocaNum, the estimate is intended as a learning signal rather than an official proficiency certificate. Definition and sentence questions provide evidence about recognition in context; uncertainty and later performance can change what deserves attention. The point is not to defend one impressive number forever. The point is to build a more accurate map as more evidence arrives.
Use the number as a map, not a verdict
A vocabulary estimate becomes useful when you stop treating it like a trophy count.
Keep four ideas attached to the number:
- The unit: what exactly was counted?
- The threshold: what did “known” mean on this test?
- The evidence: how much sampling, guessing, and uncertainty sits underneath the estimate?
- The next action: which words should receive attention now?
For practice, that last question matters most. Separate words that are stable from words that are developing, forgotten, or still unmeasured. Try to retrieve weak words before revealing the answer, meet them in clear context, and return after a delay.
If you want to turn measurement into a repeatable learning loop, continue with how to improve your English vocabulary without random word lists. If the difference between “I know it when I see it” and “I can actually retrieve it” is the part that interests you, read recognition vs recall.
The best vocabulary estimate is not the one with the most impressive number. It is the one whose meaning you understand well enough to make a better learning decision.
Sources
- Brysbaert, M., Stevens, M., Mandera, P., & Keuleers, E. (2016). How Many Words Do We Know? Practical Estimates of Vocabulary Size Dependent on Word Definition, the Degree of Language Input and the Participant's Age. Frontiers in Psychology.
- Nation, P., & Beglar, D. (2007). A vocabulary size test. The Language Teacher, 31(7), 9–13.
- Beglar, D. (2010). A Rasch-based validation of the Vocabulary Size Test. Language Testing, 27(1), 101–118.
- Schmitt, N. (2014). Size and Depth of Vocabulary Knowledge: What the Research Shows. Language Learning, 64(4), 913–951.
- Kremmel, B., & Schmitt, N. (2016). Interpreting vocabulary test scores: What do various item formats tell us about learners' ability to employ words? Language Assessment Quarterly, 13, 377–392.
- Stoeckel, T., McLean, S., & Nation, P. (2020). Limitations of size and levels tests of written receptive vocabulary knowledge. Studies in Second Language Acquisition.
- Stewart, J. (2014). Do Multiple-Choice Options Inflate Estimates of Vocabulary Size on the VST? Language Assessment Quarterly, 11(3), 271–282.
- Council of Europe. (2020). Common European Framework of Reference for Languages: Companion Volume.