Release of the List of "Holy Quran" Words with Stem and Morphological Root
Dataset Introduction
Access to the core (stem) and morphological root of words is one of the valuable criteria that significantly improves the accuracy of Arabic text analysis. For this purpose, the morphological data related to Quranic words, which includes these new features, has been prepared using the Noor Morphological Analyzer and finally reviewed and completed by Arabic linguists. The data is provided in XML format for smoother and easier processing.
Key Features of the Dataset
| Feature | Description |
|---|---|
| Stemming (Pirasteh) | Prefixes and suffixes have been separated from the core of the word |
| Morphological Root (Risheh-e-Sarfi) | Provided for words that have a root in the Arabic language |
| Position Identification | All words of the Holy Quran can be accessed through Surah number, Verse number, and Word number |
Example
For example, the word "لِنُورِهِ" (Linoorihi) is presented in the dataset as shown below:

Hope
The "Noor Text Mining Group" hopes that the use of this achievement will be useful and effective for esteemed researchers and analysts.
Usage Rights
It should be noted that the intellectual property rights of this dataset belong to the Noor Computer Research Center for Islamic Sciences (Noor) , and non-commercial use is permitted provided that the center is acknowledged.
Access
The aforementioned dataset is accessible via the link below.