EXP-0012 — Base rate of exact count equality: what an equality claim has to draw on
completed Deterministic computation
Question
In the Qur'an's vocabulary, how often do two distinct types share an exact occurrence count, and is that rate a property of this text or of word-frequency distributions in general?
Hypothesis
Zipf-distributed vocabularies produce many exact ties, so a large stock of ready-made 'X occurs as often as Y' pairs should exist in every text tested. Directional prediction fixed before running: comparison texts will show equal-pair rates of the same order as the Qur'an's.
Counting policy
Qur'an counted two ways: diacritized QAC lemma (the rule under which EXP-0008's surviving pair holds) and letters-basic surface types. Comparison texts have no morphology layer, so they are counted on letters-basic surface types only, and are compared against the Qur'an's SURFACE profile for an apples-to-apples reading; the lemma profile is reported separately and not compared cross-text. Frequency bands are declared thresholds, not tuned. focusCounts are the target numbers of claims already in the registry, listed before their bucket contents were examined.
Reproducibility record
- Corpus
- quran:hafs-kufan:v1
- Method version
- equality-background-v1
- Results sha256
- 34811850ba48a72692d9ac94954ed03f…
- Environment
- {"platform":"Linux-6.18.5-fc-v20-x86_64-with-glibc2.39","python":"3.11.15"}
- Ran
- 2026-08-15T05:44:38Z → 2026-08-15T05:44:52Z
Re-run with python3 services/research-worker/scripts/run_experiment.py data/research/specs/EXP-0012.json. Input checksums are recorded in the result file.
