← Laboratory

EXP-0012 — Base rate of exact count equality: what an equality claim has to draw on

completed Deterministic computation

Question

In the Qur'an's vocabulary, how often do two distinct types share an exact occurrence count, and is that rate a property of this text or of word-frequency distributions in general?

Hypothesis

Zipf-distributed vocabularies produce many exact ties, so a large stock of ready-made 'X occurs as often as Y' pairs should exist in every text tested. Directional prediction fixed before running: comparison texts will show equal-pair rates of the same order as the Qur'an's.

Counting policy

Qur'an counted two ways: diacritized QAC lemma (the rule under which EXP-0008's surviving pair holds) and letters-basic surface types. Comparison texts have no morphology layer, so they are counted on letters-basic surface types only, and are compared against the Qur'an's SURFACE profile for an apples-to-apples reading; the lemma profile is reported separately and not compared cross-text. Frequency bands are declared thresholds, not tuned. focusCounts are the target numbers of claims already in the registry, listed before their bucket contents were examined.

Reproducibility record

Corpus
quran:hafs-kufan:v1
Method version
equality-background-v1
Results sha256
34811850ba48a72692d9ac94954ed03f…
Environment
{"platform":"Linux-6.18.5-fc-v20-x86_64-with-glibc2.39","python":"3.11.15"}
Ran
2026-08-15T05:44:38Z → 2026-08-15T05:44:52Z

Re-run with python3 services/research-worker/scripts/run_experiment.py data/research/specs/EXP-0012.json. Input checksums are recorded in the result file.