# Modifications

1. Read the exact bgOffice 4.4 listed `bg.dic` rows as Windows-1251 with CRLF. Do not expand Hunspell affixes. Retain lowercase listed surfaces only.
2. Read the exact MULTEXT-East v4 Bulgarian UTF-8 wordform/lemma/MSD triples. Validate the entire MSD against `bulgarian-msd-map.json`, byte-bound to the publisher Bulgarian V4 MSD index; check map consistency before candidate output. Select only allowed nonverbal classes. Exclude entire ambiguous classes and vocatives.
3. Read the 2026-09-01 Bulgarian Wiktionary XML dump. For adjectives, adverbs, adpositions and conjunctions require an own bounded Bulgarian-section POS template and non-placeholder definition; unknown lexical classification fails closed. Reject conflicting templates, obsolete-only definitions and unqualified proper-name collisions.
4. Intersect the listed publisher and lexical source surfaces, apply the positive rules and empty curation ledger, deduplicate surfaces, assign uniform rank 1 and standard flag 1, then encode the selected surfaces as sequential IAG2 assets.
5. Split the unmodified Wiktionary compressed dump into three byte-identical 16 MiB-or-smaller publication parts. Reassemble and verify its source SHA-256 offline before parsing. This split changes delivery only, not source content.

These modifications are deterministic. The candidate contains 5,114 unique surfaces. See `candidate-final.tsv`, `generation-lock.json`, `wiki-parts.json` and `reproduce.py` for byte identities and offline reproduction.
