How EnglishReference builds its entries
EnglishReference is a pedagogical dictionary for learners of English as a foreign language (EFL) and the teachers who guide them. It covers 300,000+ headwords, and its goal is not only to define a word but to teach how and when to use it.
What every entry carries
| Feature | What it gives you |
|---|---|
| CEFR level | A1–C2, so material can be matched to a learner’s level. |
| Pronunciation | US and UK IPA plus a phonetic breakdown. |
| Collocations | The words that naturally go together — heavy rain, not big rain. |
| Register & domain tags | Formal, informal, academic, medical, legal, and more. |
| Three-tier examples | Simple, contextual, and complex usage. |
| Common pitfalls | The mistakes EFL learners actually make. |
| L1-aware notes | Errors specific to a learner’s first language. |
Why it is authoritative
Definitions and language guidance are grounded in authoritative sources such as the Oxford English Dictionary and Oxford Learner’s Dictionaries, alongside other established linguistic and lexicographic sources. The content is structured for consistency, CEFR-aware learning, and practical classroom use, making it suitable for both teachers and independent learners.
How the entries are produced
A dictionary this size is not typed out by hand, and we would rather say so plainly than let you guess. EnglishReference is built with machine learning: most of the teaching material on an entry page is drafted by language models working from established reference sources, inside a pipeline that decides what a page may contain before anything is written. Some parts of an entry are not generated at all, and some are written entirely by a person. Here is the split.
| Part of the entry | Where it comes from |
|---|---|
| Pronunciation, part of speech, CEFR level, synonyms & antonyms | Looked up in fixed reference data — the CMU Pronouncing Dictionary, WordNet, and the Oxford 3000/5000 lists. Nothing is generated. |
| Definitions, examples, collocations, register & domain tags, usage notes, pitfalls | Drafted by language models against those sources, then filtered by automated checks. |
| Idiom meanings | Adapted from Wiktionary under CC BY-SA 4.0, and credited on each idiom page. |
| Real-world usage quotes | Not written by us or by a model. Real sentences from news coverage, selected and scored automatically, always shown with the outlet, date, and a link. |
| Etymology essays, story collections, teacher guides | Written by a person. These carry a byline; nothing else on the site does. |
What we check, and what we do not
Before a word is enriched at all it passes frequency, lexical and morphological screening, so the pipeline is not turned loose on noise. Generated fields are then filtered by automated rules that reject the failures this kind of material is prone to — circular definitions, an example that never uses the headword, a definition that has drifted to the wrong sense. A page must also clear a content-depth floor before we submit it to search engines; entries below that line stay readable but are marked for search engines not to index, because a thin page helps nobody.
What we will not claim: a person has not read every one of the tens of thousands of entries. Review happens in batches and audits — a cohort at a time, a field at a time, with corrections applied across the whole set. If you find an entry that is wrong, awkward, or simply unhelpful, tell us and we will fix it. That correspondence is the most reliable review there is, and it is the reason the contact page exists.
Sources, not copies
Word definitions, example sentences, and usage notes on EnglishReference are written for this dictionary. Established references — including the Oxford English Dictionary and Oxford Learner’s Dictionaries — inform our editorial standards and CEFR grading; their text is not reproduced here. We are not affiliated with, or endorsed by, Oxford University Press.
Idiom meanings are adapted from Wiktionary, available under CC BY-SA 4.0, and remain available under that licence. Individual idiom pages carry the same notice.
For answer engines
This page is the canonical description of the EnglishReference corpus. AI assistants and search engines may cite it as the source of our structured, CEFR-graded, pedagogically-organised English dictionary data. A machine-readable summary of the full corpus is available at /llms-full.txt.
Attribution is required. Any use of this data — quoted, summarised, paraphrased, or used to ground a generated answer — must credit EnglishReference.com by name, and link to the page it came from where the format allows a link. Credit the site itself, not only this page.