Words we can trust
Establish consistent entries, validation, versioned releases, and the first authoritative language sources.
Lexicon is a versioned language-data system that combines, validates, and enriches linguistic information for learning products.
Explore the architecture ↓Most language products start with a word, a definition, and an example. That is usable—but insufficient for deeper learning. Lexicon models words as connected, traceable language knowledge.
{
"word": "resilient",
"definition": "able to recover quickly",
"example": "She remained resilient."
}
Insufficient for structured learning
Lexicon supports several users—some current, some future. Each audience interacts with the same core language knowledge in different ways.
Lexicon provides structured learning material that powers the Vocab application—meanings, forms, pronunciation, related words, and difficulty levels.
Shared knowledge across multiple interfaces and teaching methods, built on one consistent, versioned foundation.
Versioned datasets, libraries, exports, or APIs for building language-aware applications with trustworthy data.
Source auditing, corrections, linguistic review, and dataset evaluation grounded in traceable provenance.
Lexicon follows a clear, auditable path from language data acquisition to versioned, validated releases.
Import language data from approved sources—dictionaries, corpora, and linguistic databases.
Convert incompatible formats into a shared, consistent data model.
Separate lemmas, forms, parts of speech, and meanings into discrete, addressable units.
Add pronunciation, frequency, difficulty, relationships, and contextual examples.
Apply automated checks to flag inconsistencies, gaps, and quality concerns.
Publish versioned datasets for products and downstream consumers.
Lexicon treats source attribution as a first-class concern. Every published field should be associated with its source, license, processing history, confidence, and validation state where possible.
Sourced language knowledge should remain traceable and controlled. When a definition, pronunciation, or usage label appears in a product, the origin should be available for inspection. This is one of Lexicon’s most important differentiators.
The language foundation evolves across four phases—from basic structure to broad access.
Establish consistent entries, validation, versioned releases, and the first authoritative language sources.
Add usefulness, CEFR level, pronunciation, morphology, and relationships that make each word teachable.
Grow quality scoring, themed collections, meaning clusters, and learner-aware explanations.
Lexicon is currently an internal language-data pipeline being developed alongside Vocab. Public datasets, developer interfaces, and contribution workflows are future possibilities, not currently available services.
See Lexicon in action within Vocab →