← Agents

Agent morphology

3 candidates · 0 approved · 0 dismissed · 3 awaiting review

1743 lemmas occur exactly once in the corpus (hapax legomena) unreviewed deterministic candidate
{
 "count": 1743,
 "sample": [
  {
   "ayah": "1:7",
   "form": "ٱلۡمَغۡضُوبِ",
   "lemma": "مغضوب",
   "word": 6
  },
  {
   "ayah": "2:16",
   "form": "رَبِحَت",
   "lemma": "ربحت",
   "word": 7
  },
  {
   "ayah": "2:17",
   "form": "ٱسۡتَوۡقَدَ",
   "lemma": "استوقد",
   "word": 4
  },
  {
   "ayah": "2:19",
   "form": "كَصَيِّبٖ",
   "lemma": "صيب",
   "word": 2
  },
  {
   "ayah": "2:26",
   "form": "بَعُوضَةٗ",
   "lemma": "بعوضة",
   "word": 9
  },
  {
   "ayah": "2:30",
   "form": "وَنُقَدِّسُ",
   "lemma": "نقدس",
   "word": 21
  },
  {
   "ayah": "2:36",
   "form": "فَأَزَلَّهُمَا",
   "lemma": "أزل",
   "word": 1
  },
  {
   "ayah": "2:54",
   "form": "بِٱتِّخَاذِكُمُ",
   "lemma": "اتخاذ",
   "word": 9
  },
  {
   "ayah": "2:60",
   "form": "فَٱنفَجَرَتۡ",
   "lemma": "انفجرت",
   "word": 9
  },
  {
   "ayah": "2:61",
   "form": "بَقۡلِهَا",
   "lemma": "بقل",
   "word": 18
  },
  {
   "ayah": "2:61",
   "form": "وَقِثَّآئِهَا",
   "lemma": "قثائ",
   "word": 19
  },
  {
   "ayah": "2:61",
   "form": "وَفُومِهَا",
   "lemma": "فوم",
   "word": 20
  },
  {
   "ayah": "2:61",
   "form": "وَعَدَسِهَا",
   "lemma": "عدس",
   "word": 21
  },
  {
   "ayah": "2:61",
   "form": "وَبَصَلِهَاۖ",
   "lemma": "بصل",
   "word": 22
  },
  {
   "ayah": "2:68",
   "form": "فَارِضٞ",
   "lemma": "فارض",
   "word": 15
  },
  {
   "ayah": "2:68",
   "form": "عَوَانُۢ",
   "lemma": "عوان",
   "word": 18
  },
  {
   "ayah": "2:69",
   "form": "صَفۡرَآءُ",
   "lemma": "صفراء",
   "word": 14
  },
  {
   "ayah": "2:69",
   "form": "فَاقِعٞ",
   "lemma": "فاقع",
   "word": 15
  },
  {
   "ayah": "2:69",
   "form": "تَسُرُّ",
   "lemma": "تسر",
   "word": 17
  },
  {
   "ayah": "2:71",
   "form": "شِيَةَ",
   "lemma": "شية",
   "word": 15
  },
  {
   "ayah": "2:72",
   "form": "فَٱدَّٰرَْٰٔتُمْ",
   "lemma": "ادارأ",
   "word": 4
  },
  {
   "ayah": "2:74",
   "form": "قَسۡوَةٗۚ",
   "lemma": "قسوة",
   "word": 11
  },
  {
   "ayah": "2:74",
   "form": "يَتَفَجَّرُ",
   "lemma": "يتفجر",
   "word": 16
  },
  {
   "ayah": "2:85",
   "form": "تُفَٰدُوهُمۡ",
   "lemma": "تفاد",
   "word": 18
  },
  {
   "ayah": "2:93",
   "form": "وَأُشۡرِبُواْ",
   "lemma": "أشرب",
   "word": 15
  },
  {
   "ayah": "2:96",
   "form": "أَحۡرَصَ",
   "lemma": "أحرص",
   "word": 2
  },
  {
   "ayah": "2:96",
   "form": "بِمُزَحۡزِحِهِۦ",
   "lemma": "مزحزح",
   "word": 17
  },
  {
   "ayah": "2:98",
   "form": "وَمِيكَىٰلَ",
   "lemma": "ميكال",
   "word": 8
  },
  {
   "ayah": "2:102",
   "form": "بِبَابِلَ",
   "lemma": "بابل",
   "word": 21
  },
  {
   "ayah": "2:102",
   "form": "هَٰرُوتَ",
   "lemma": "هاروت",
   "word": 22
  },
  {
   "ayah": "2:102",
   "form": "وَمَٰرُوتَۚ",
   "lemma": "ماروت",
   "word": 23
  },
  {
   "ayah": "2:114",
   "form": "خَرَابِهَآۚ",
   "lemma": "خراب",
   "word": 13
  },
  {
   "ayah": "2:121",
   "form": "تِلَاوَتِهِۦٓ",
   "lemma": "تلاوت",
   "word": 6
  },
  {
   "ayah": "2:125",
   "form": "مَثَابَةٗ",
   "lemma": "مثابة",
   "word": 4
  },
  {
   "ayah": "2:125",
   "form": "مُصَلّٗىۖ",
   "lemma": "مصلى",
   "word": 11
  },
  {
   "ayah": "2:148",
   "form": "وِجۡهَةٌ",
   "lemma": "وجهة",
   "word": 2
  },
  {
   "ayah": "2:148",
   "form": "مُوَلِّيهَاۖ",
   "lemma": "مولي",
   "word": 4
  },
  {
   "ayah": "2:158",
   "form": "ٱلصَّفَا",
   "lemma": "صفا",
   "word": 2
  },
  {
   "ayah": "2:158",
   "form": "وَٱلۡمَرۡوَةَ",
   "lemma": "مروة",
   "word": 3
  },
  {
   "ayah": "2:158",
   "form": "ٱعۡتَمَرَ",
   "lemma": "اعتمر",
   "word": 11
  },
  {
   "ayah": "2:159",
   "form": "ٱللَّـٰعِنُونَ",
   "lemma": "لاعن",
   "word": 20
  },
  {
   "ayah": "2:164",
   "form": "ٱلۡمُسَخَّرِ",
   "lemma": "مسخر",
   "word": 37
  },
  {
   "ayah": "2:171",
   "form": "يَنۡعِقُ",
   "lemma": "ينعق",
   "word": 6
  },
  {
   "ayah": "2:177",
   "form": "وَٱلۡمُوفُونَ",
   "lemma": "موفي",
   "word": 36
  },
  {
   "ayah": "2:178",
   "form": "ٱلۡقَتۡلَىۖ",
   "lemma": "قتلى",
   "word": 8
  },
  {
   "ayah": "2:178",

Skeptic checks

  • [neutral] Per morphology dataset quran-morphology-8f38b39 (QAC 0.4 derivative); rarity is annotation-dependent. Lemma normalization (letters-basic) merges spelling variants; counts shift under other normalizations.
  • [weakens] Large hapax counts are a statistical property of essentially every natural-language corpus (Zipf tail); the list is useful for study, not remarkable in itself.
Your decision is recorded in data/research/reviews.json
420 roots appear in exactly one ayah unreviewed deterministic candidate
{
 "count": 420,
 "sample": [
  {
   "ayah": "2:16",
   "root": "ربح",
   "words": 1
  },
  {
   "ayah": "2:61",
   "root": "بصل",
   "words": 1
  },
  {
   "ayah": "2:61",
   "root": "بقل",
   "words": 1
  },
  {
   "ayah": "2:61",
   "root": "عدس",
   "words": 1
  },
  {
   "ayah": "2:61",
   "root": "فوم",
   "words": 1
  },
  {
   "ayah": "2:61",
   "root": "قثأ",
   "words": 1
  },
  {
   "ayah": "2:69",
   "root": "فقع",
   "words": 1
  },
  {
   "ayah": "2:71",
   "root": "وشي",
   "words": 1
  },
  {
   "ayah": "2:158",
   "root": "مرو",
   "words": 1
  },
  {
   "ayah": "2:171",
   "root": "نعق",
   "words": 1
  },
  {
   "ayah": "2:185",
   "root": "رمض",
   "words": 1
  },
  {
   "ayah": "2:197",
   "root": "زود",
   "words": 2
  },
  {
   "ayah": "2:255",
   "root": "أود",
   "words": 1
  },
  {
   "ayah": "2:255",
   "root": "وسن",
   "words": 1
  },
  {
   "ayah": "2:256",
   "root": "فصم",
   "words": 1
  },
  {
   "ayah": "2:259",
   "root": "سنه",
   "words": 1
  },
  {
   "ayah": "2:264",
   "root": "صلد",
   "words": 1
  },
  {
   "ayah": "2:265",
   "root": "طلل",
   "words": 1
  },
  {
   "ayah": "2:267",
   "root": "غمض",
   "words": 1
  },
  {
   "ayah": "2:273",
   "root": "لحف",
   "words": 1
  },
  {
   "ayah": "2:275",
   "root": "خبط",
   "words": 1
  },
  {
   "ayah": "3:41",
   "root": "رمز",
   "words": 1
  },
  {
   "ayah": "3:49",
   "root": "ذخر",
   "words": 1
  },
  {
   "ayah": "3:61",
   "root": "بهل",
   "words": 1
  },
  {
   "ayah": "3:75",
   "root": "دنر",
   "words": 1
  },
  {
   "ayah": "3:156",
   "root": "غزو",
   "words": 1
  },
  {
   "ayah": "3:159",
   "root": "فظظ",
   "words": 1
  },
  {
   "ayah": "4:2",
   "root": "حوب",
   "words": 1
  },
  {
   "ayah": "4:3",
   "root": "عول",
   "words": 1
  },
  {
   "ayah": "4:6",
   "root": "بدر",
   "words": 1
  },
  {
   "ayah": "4:21",
   "root": "فضو",
   "words": 1
  },
  {
   "ayah": "4:51",
   "root": "جبت",
   "words": 1
  },
  {
   "ayah": "4:56",
   "root": "نضج",
   "words": 1
  },
  {
   "ayah": "4:71",
   "root": "ثبي",
   "words": 1
  },
  {
   "ayah": "4:72",
   "root": "بطأ",
   "words": 1
  },
  {
   "ayah": "4:83",
   "root": "ذيع",
   "words": 1
  },
  {
   "ayah": "4:83",
   "root": "نبط",
   "words": 1
  },
  {
   "ayah": "4:100",
   "root": "رغم",
   "words": 1
  },
  {
   "ayah": "4:102",
   "root": "سلح",
   "words": 4
  },
  {
   "ayah": "4:119",
   "root": "بتك",
   "words": 1
  },
  {
   "ayah": "4:143",
   "root": "ذبذب",
   "words": 1
  },
  {
   "ayah": "5:3",
   "root": "خنق",
   "words": 1
  },
  {
   "ayah": "5:3",
   "root": "ذكو",
   "words": 1
  },
  {
   "ayah": "5:3",
   "root": "نطح",
   "words": 1
  },
  {
   "ayah": "5:3",
   "root": "وقذ",
   "words": 1
  },
  {
   "ayah": "5:26",
   "root": "تيه",
   "words": 1
  },
  {
   "ayah": "5:31",
   "root": "بحث",
   "words": 1
  },
  {
   "ayah": "5:33",
   "root": "نفي",
   "words": 1
  },
  {
   "ayah": "5:48",
   "root": "نهج",
   "words": 1
  },
  {
   "ayah": "5:82",
   "root": "قسس",
   "words": 1
  },
  {
   "ayah": "5:94",
   "root": "رمح",
   "words": 1
  },
  {
   "ayah": "5:103",
   "root": "سيب",
   "words": 1
  },
  {
   "ayah": "6:70",
   "root": "بسل",
   "words": 2
  },
  {
   "ayah": "6:71",
   "root": "حير",
   "words": 1
  },
  {
   "ayah": "6:95",
   "root": "نوي",
   "words": 1
  },
  {
   "ayah": "6:99",
   "root": "قنو",
   "words": 1
  },
  {
   "ayah": "6:99",
   "root": "ينع",
   "words": 1
  },
  {
   "ayah": "6:143",
   "root": "ضأن",
   "words": 1
  },
  {
   "ayah": "6:143",
   "root": "معز",
   "words": 1
  },
  {
   "ayah": "6:146",
   "root": "شحم",
   "words": 1
  },
  {
   "ayah": "7:18",
   "root": "ذأم",
   "words": 1
  },
  {
   "ayah": "7:26",
   "root": "ريش",
   "words": 1
  },
  {
   "ayah": "7:54",
   "root": "حثث",
   "words": 1
  },
  {
   "ayah": "7:58",
   "root": "نكد",
   "words": 1
  },
  {
   "ayah": "7:74",
   "root": "سهل",
   "words": 1
  },
  {
   "ayah": "7:133",
   "root": "ضفدع",
   "words"

Skeptic checks

  • [neutral] Per morphology dataset quran-morphology-8f38b39 (QAC 0.4 derivative); rarity is annotation-dependent.
  • [weakens] Single-context roots are expected in any corpus of this size; see the hapax note.
Your decision is recorded in data/research/reviews.json
Surahs with the highest and lowest root diversity per 100 words unreviewed deterministic candidate
{
 "highest": [
  {
   "distinctRoots": 36,
   "rootsPer100Words": 66.67,
   "surah": 91,
   "words": 54
  },
  {
   "distinctRoots": 22,
   "rootsPer100Words": 64.71,
   "surah": 95,
   "words": 34
  },
  {
   "distinctRoots": 24,
   "rootsPer100Words": 60,
   "surah": 100,
   "words": 40
  },
  {
   "distinctRoots": 100,
   "rootsPer100Words": 57.8,
   "surah": 78,
   "words": 173
  },
  {
   "distinctRoots": 41,
   "rootsPer100Words": 57.75,
   "surah": 92,
   "words": 71
  }
 ],
 "lowest": [
  {
   "distinctRoots": 477,
   "rootsPer100Words": 14.37,
   "surah": 7,
   "words": 3320
  },
  {
   "distinctRoots": 419,
   "rootsPer100Words": 13.74,
   "surah": 6,
   "words": 3050
  },
  {
   "distinctRoots": 442,
   "rootsPer100Words": 12.7,
   "surah": 3,
   "words": 3481
  },
  {
   "distinctRoots": 462,
   "rootsPer100Words": 12.33,
   "surah": 4,
   "words": 3747
  },
  {
   "distinctRoots": 587,
   "rootsPer100Words": 9.6,
   "surah": 2,
   "words": 6116
  }
 ],
 "minimumWords": 30,
 "surahsConsidered": 100
}

Skeptic checks

  • [weakens] Type diversity falls with text length by construction (roots repeat as texts grow); per-100-word normalization reduces but does not remove the confound. Comparisons across very different lengths remain unreliable.
  • [neutral] Per morphology dataset quran-morphology-8f38b39 (QAC 0.4 derivative); rarity is annotation-dependent.
Your decision is recorded in data/research/reviews.json