# Afrikaans Words from letters: source identity

The published 7,585 word forms derive from the pre-enumerated entries in LibreOffice Dictionaries `af_ZA.dic` at commit `32b006a2c22a4ac7e8ed3f03346f7b3d85a970a4`. The exact raw dictionary, affix file and upstream README are in this directory. The affix file is interpreted for eligibility flags; it does not generate any additional forms. The three raw files remain subject to LGPL 2.1 or later and the notices in `README_af_ZA.txt`.

The NCHLT Afrikaans Text Corpora, dated 2014-05-30, were developed for the South African Department of Arts and Culture by CTexT, North-West University. The source is the [SADiLaR item](https://repo.sadilar.org/items/289c350d-25a1-4c27-8285-c83b031318c7), licensed [CC BY 2.5 South Africa](https://creativecommons.org/licenses/by/2.5/za/). Its exact-surface frequency list supplied bulk admission and ranking. Its named-entity list screened case risks, and its clean corpus supplied context for individual curation. The corpus and original ZIP are not mirrored here; `source-matrix.json` binds the exact archive and member identities. The original README member is preserved byte-for-byte inside `data-sources/af-nchlt/Info.NCHLT.Corpora.Readme.raw.bin.gz` in the repository. This public directory includes a separately labeled Windows-1252 analytical view.

The files `candidate.tsv`, `policy.json`, `curation-v2.json`, and `generation-lock.json` preserve the editable selection and byte-identical IAG2 build inputs. Both upstream license boundaries remain distinct; NCHLT is not relicensed under LGPL. Neither source owner endorses Instant Anagram.
