Product

Language knowledge,
built to be taught.

Lexicon is a versioned language-data system that combines, validates, and enriches linguistic information for learning products.

Explore the architecture

More than a dictionary entry.

Most language products start with a word, a definition, and an example. That is usable—but insufficient for deeper learning. Lexicon models words as connected, traceable language knowledge.

{
  "word": "resilient",
  "definition": "able to recover quickly",
  "example": "She remained resilient."
}

Insufficient for structured learning

Lexeme
  ├── lemma
  ├── language
  ├── pronunciation
  ├── grammatical category
  ├── inflected forms
  ├── senses
     ├── definition
     ├── examples
     ├── usage labels
     └── difficulty
  ├── word family
  ├── semantic relationships
  ├── frequency
  ├── source provenance
  └── quality status

Designed to serve multiple audiences.

Lexicon supports several users—some current, some future. Each audience interacts with the same core language knowledge in different ways.

Current

Vocab

Lexicon provides structured learning material that powers the Vocab application—meanings, forms, pronunciation, related words, and difficulty levels.

Future

Future Veriteum language products

Shared knowledge across multiple interfaces and teaching methods, built on one consistent, versioned foundation.

Future

Developers

Versioned datasets, libraries, exports, or APIs for building language-aware applications with trustworthy data.

Future

Researchers and contributors

Source auditing, corrections, linguistic review, and dataset evaluation grounded in traceable provenance.

From raw sources to trusted knowledge.

Lexicon follows a clear, auditable path from language data acquisition to versioned, validated releases.

01

Acquire

Import language data from approved sources—dictionaries, corpora, and linguistic databases.

02

Normalize

Convert incompatible formats into a shared, consistent data model.

03

Resolve

Separate lemmas, forms, parts of speech, and meanings into discrete, addressable units.

04

Enrich

Add pronunciation, frequency, difficulty, relationships, and contextual examples.

05

Validate

Apply automated checks to flag inconsistencies, gaps, and quality concerns.

06

Release

Publish versioned datasets for products and downstream consumers.

Every field should know where it came from.

Lexicon treats source attribution as a first-class concern. Every published field should be associated with its source, license, processing history, confidence, and validation state where possible.

Sourced language knowledge should remain traceable and controlled. When a definition, pronunciation, or usage label appears in a product, the origin should be available for inspection. This is one of Lexicon’s most important differentiators.

Lexicon direction.

The language foundation evolves across four phases—from basic structure to broad access.

NOW · FOUNDATION

Words we can trust

Establish consistent entries, validation, versioned releases, and the first authoritative language sources.

DefinitionsQuality checksPublished datasets
NEXT · RICHNESS

Words with dimension

Add usefulness, CEFR level, pronunciation, morphology, and relationships that make each word teachable.

FrequencyPronunciationWord families
LATER · ACCESS

Knowledge that connects

Grow quality scoring, themed collections, meaning clusters, and learner-aware explanations.

Knowledge graphDomain packsComplexity

Current maturity.

Lexicon is currently an internal language-data pipeline being developed alongside Vocab. Public datasets, developer interfaces, and contribution workflows are future possibilities, not currently available services.

See Lexicon in action within Vocab