Methodology & provenance
TDBBR is a reader and a research instrument, not an authority. It never generates interpretation, never calls anything “tafsir,” and never claims statistical proof of divine origin. Everything visible in this build is either canonical source text (attributed, checksummed, immutable) or a deterministic computation you can reproduce.
Content labels
Qur’anic text Translation Deterministic computation Personal reflection
These universes are kept technically separate end to end. AI-hypothesis and scholarly layers do not exist in this build; when they arrive they will carry their own labels and default to hidden until reviewed.
Sources
| Source | Version | License | Review status | SHA-256 |
|---|---|---|---|---|
| quran-json-3.1.2 | 3.1.2 | CC-BY-4.0 (package); underlying texts carry their own terms | pending | bae3acab517304dc… |
| openiti-0256Bukhari-Sahih | JK000110-ara | OpenITI corpus release: CC BY-NC-SA 4.0 (verify per release); underlying classical work public domain | pending | cdbf6ba76dfc918d… |
| openiti-0354Mutanabbi-Diwan | JK007610-ara | OpenITI corpus release: CC BY-NC-SA 4.0 (verify per release); underlying classical work public domain | pending | c5076264b4fcc0b6… |
| quran-morphology-8f38b39 | 8f38b3901682 | GNU GPL (inherited from QAC v0.4) | pending | c9028b44ea859ea2… |
| tafsir-muyassar-5bbe089 | 5bbe08955432 | repo MIT (code); al-Muyassar text: King Fahd Complex publication — distribution terms REVIEW PENDING | license-review-pending — private prototype only | d36fdf06d7469e05… |
| hadith-json-70b83d6 | 70b83d6d2199 | repo license UNSTATED; sunnah.com data terms REVIEW PENDING | license-review-pending — private prototype only | 9d2e4194786c275f… |
| athan-mp3-cf1a9dd | cf1a9dd4f1e5 | UNSTATED upstream — distribution terms REVIEW PENDING | license-review-pending — private prototype only | 73bffee51e6f2d24… |
| quran-wbw-f41d709 | f41d7090bb93 | Apache-2.0 (upstream repository) | license recorded — Apache-2.0; underlying en.transliteration originates from Tanzil (CC BY 3.0, verbatim) | — |
| tafsir-ar-tafsir-al-jalalayn-main | main | repo MIT (code); Tafsir al-Jalalayn text: Jalal al-Din al-Mahalli (d. 864 AH) and Jalal al-Din al-Suyuti (d. 911 AH). Work long out of copyright; terse verse-for-verse gloss. | cleared-as-public-domain-work (edition claim unreviewed) | 881b9324f80360be… |
| tafsir-ar-tafseer-al-saddi-main | main | repo MIT (code); Taysir al-Karim (al-Sa'di) text: Abd al-Rahman al-Sa'di (d. 1376 AH / 1957 CE). A copyright term may still be running depending on jurisdiction — REVIEW BEFORE COMMERCIAL USE. | license-review-pending — NOT CLEARED FOR COMMERCIAL USE | 2bf84dde930cefeb… |
| juz-boundaries-semarketir | 5726efc2887e | structural fact about the Qur'an's conventional divisions, not creative expression; verified against our own corpus rather than trusted | verified-by-tiling | 702defa398d83d38… |
The corpus baseline is Hafs ‘an ‘Asim with the documented Kufan numbering (6,236 ayahs). The Uthmani text arrives via the pinned quran-json 3.1.2 package (text from the Noble Qur’an Encyclopedia); byte-level cross-verification against official Tanzil is an open release gate, and until it completes this build is for private study use. The Saheeh International translation carries its own copyright and is likewise pending distribution review.
Transforms
Canonical text is never modified. Analysis runs on separate, declared representations:
- letters-basic-v1 — NFC; Arabic letters and spaces only; alif-wasla → alif; marks, symbols, and digits dropped; hamza carriers kept distinct (over-normalization would fabricate matches).
- search-normalized-v1 — letters-basic plus forgiving unifications (alif forms, teh marbuta, alif maqsura, hamza carriers). Retrieval only, never research.
Connection algorithms
| Edge type | Rule | Count |
|---|---|---|
| EXACT_PHRASE_MATCH | Maximal shared word sequences of ≥ 5 tokens on letters-basic-v1, whole corpus, all ayah pairs (exact-phrase-match-v1) | 3,121 |
| SHARES_RARE_ROOT | Ayah pairs sharing a root that appears in ≤ 12 ayahs corpus-wide, per the attributed morphology dataset (shares-rare-root-v1) | 10,682 |
| PRECEDES | Canonical reading order within each surah | 6,122 |
Thresholds (≥ 5 tokens, ≤ 12 ayahs) are declared parameters recorded in every edge’s provenance; changing them creates a new algorithm version. The morphology layer covers 6,236 ayahs with 1,651 distinct roots; 10 ayahs have tokenization differences between the morphology dataset and this corpus text — recorded, not hidden.
Reproducibility
pnpm data:pipeline rebuilds every derived file from the checksummed raw sources; pnpm data:verify:quran fails the build on any checksum drift. Corpus version quran:hafs-kufan:v1, pipeline pipeline-v1, layout layout-phyllotaxis-v1 (a rendering artifact, not a property of the text). Generated 2026-07-31T01:41:23Z.
What this build does not do
No model calls. No semantic “feels related” links. No miracle claims. No piety scores or public tracking. Notes, bookmarks, and reading position never leave your browser, and you can export or erase them at any time from the Read page.
Attributions
Qur’anic text: Noble Qur’an Encyclopedia (quranenc.com) via quran-json (Risan Bagja Pradana, CC-BY-4.0). Translation: Saheeh International (Umm Muhammad) via Tanzil.net. Morphology: Quranic Arabic Corpus v0.4 (Kais Dukes, University of Leeds, GPL) with documented corrections by mustafa0x. Type: Amiri Quran (Khaled Hosny, SIL OFL 1.1). Rendering: Sigma.js + Graphology. Full registry: docs/DATA_SOURCES.md and ATTRIBUTIONS.md in the repository.
