# Estonian dictionary modifications

The EKI headwords were checked for exact schema and safe IDs, normalized to NFC, limited to already-lowercase single tokens of 2–30 written letters in the 27-letter Estonian alphabet, and intersected with exact-case ENC 2021 lemma frequency at a minimum count of 3,000. The tagged ENC source supplies a known common POS and excludes any key with H (name) or Y (abbreviation). Underscores are removed only from the documented POS join key. POS counts are never summed for rank.

Identical headwords were deduplicated into one runtime record while retaining every EKI wordId in the reproducible foundation provenance. Records are ranked by descending untagged lemma count, then raw lexical order. The selected records were packaged as 21 IAG2 core/display shard pairs and a content-addressed set manifest using the repository's dictionary-set generator. No inflections, phrase data, or guessed morphology were added.

Candidate policy: `et-ys2024-enc21-lemma-frequency3000-v1`. Runtime flags: standard, lemma, frequencyRanked (21). See [POS-MAPPING.md](POS-MAPPING.md), [source-identity-binding.json](source-identity-binding.json), and [publication-manifest.json](publication-manifest.json).
