Similar Titles; Intelligent Discovery of Similar Article Titles from a Vast Collection of Articles
The Challenge of Text Document Processing
Numerous articles, news items, and text documents are produced and published daily in digital environments, making it impossible to easily examine the content of such a vast volume of information. The large number of texts, their linguistic diversity, varying lengths, and different encodings are among the difficulties of working with text documents.
Experts from various scientific fields have worked to solve this problem. Specialists in artificial intelligence, information retrieval, data mining, text mining, and text similarity detection have proposed solutions using information retrieval knowledge.
What are Similar Titles?
"Similar Titles" is one of these similarity detection systems, developed and offered based on the extensive data from the NoorMag journals database.
Similar titles is a capability for intelligently discovering similar article titles. Using text mining and artificial intelligence techniques, it suggests the most similar articles (from a title perspective) to the user when viewing each article.
Importance and Applications
Finding articles related to a given article is a research concern that must be addressed to organize comprehensive and non-repetitive research in the shortest possible time. The primary method for identifying relationships between articles is examining common words in their titles. This tool uses article titles to identify relationships between them.
Applications in News and Scientific Databases
Using similarity detection systems to discover hidden relationships among text data has various applications:
- News Databases: For identifying relationships between different news items (e.g., Google News or the "Related" section of Hamshahri News[1])
- Scientific Databases: Such as the "See also" section of Wikipedia
Innovation in NoorMag
While the only practical feature in the similarity detection process is the article titles, the designers of NoorMag have strived to go beyond the literal level of article titles and get closer to their meaning and subject matter.
Therefore, various experiments were conducted in the text mining department of the Noor Computer Research Center for Islamic Sciences to achieve this transition in a better way.
Techniques Used
Layered organization of semantic word clustering is an example of the techniques used in these experiments. This technique leads to the discovery of many "word co-occurrence" relationships.
Word Co-occurrence
Word co-occurrence means that the presence of one word implies the presence of another. For example, when the word "oil" appears, it is very likely that the word "gas" is also used.
On the other hand, the co-occurrence of two words indicates commonalities between them. In many words, these commonalities are semantic in nature. Therefore, the semantic clustering process becomes possible through their co-occurrence relationship.
Unique Features
The common techniques used in this feature differ from other search engines and have functionalities that are the result of local researchers' work:
- Word clustering
- Separating keywords from other words
- Making them more effective for similarity calculation
Future Perspective
These features are still under development; God willing, the similar titles feature will have greater accuracy and quality in future versions.
Hope
The Text Mining Department of the Noor Computer Research Center for Islamic Sciences hopes that by offering this feature, the path of research for scholars in academia and seminaries will be smoothed.
References
[1] Hamshahri News Agency