Plagiarism Detection
Text Matching and ClassificationPlagiarism Detection
Definition of Plagiarism
Plagiarism is the act of presenting another person's ideas, sentences, or work as one's own. This is a form of deception and academic dishonesty. (Ballard, 2010, p. 1)
Text reuse refers to the intentional or unintentional use of existing text to create new text. If proper documentation is not provided during this reuse, plagiarism occurs.
Educational and industrial institutions often face issues of plagiarism and copyright infringement. Copyright grants exclusive publishing rights to protect ideas and information. Authors may permit free use of their copyrighted works, but unauthorized reproduction by others constitutes copyright infringement.
Forms of Plagiarism
- Direct copying (word-for-word)
- Partial copying of a work
- Paraphrase plagiarism
- Mosaic plagiarism
- Source plagiarism
- Incomplete-citation plagiarism
- Phrase plagiarism
- Idea plagiarism
Automatic Plagiarism Detection Methods
Authorship attribution (or authorship identification) is the process of determining which author from a set of possible authors wrote a given text.
Stylometry is one method used for authorship detection. Stylometry defines features of an author's writing style and measures these features across two or more texts to determine similarity between them. The idea that style operates at a subconscious level makes it more measurable. In fact, writing style can be considered a fingerprint.
Automatic Plagiarism Detection Approaches
- Fingerprinting
- String matching
- Bag of words
- Citation analysis
- Stylometry

Sameem Noor (3) - A Machine-based Plagiarism Detection Tool
One way to identify works that contain plagiarism is to use the "Sameem Noor" database from the Noor Computer Research Center for Islamic Sciences.
This database, which uses machine learning methods to find similar texts, utilizes the Noor Specialized Journals Database (4) to compare similarity among user-submitted articles.
Future sources to be added to the database include:
- Noor Digital Library Database (5)
- Books converted into Noor software through the Cultural Services Department of the Center in collaboration with content producers
The strength of the Noor Computer Research Center for Islamic Sciences in this endeavor lies in its vast collection of machine-readable texts and vocabulary in the humanities and Islamic sciences. This wealth of resources provides excellent tools, raw materials for machine learning, and rich samples for matching and similarity detection.
