Evaluating lexical feature extraction for plagiarism detection in Arabic documents
Abstract
Plagiarism detection is the task of determining whether a document contains parts from other documents by employing different styles of plagiarism, such as copying certain parts and reordering or replacing words with synonyms, without citing the original text owner. This task is important in many applications, and there are two primary types of plagiarism detection methods: external and intrinsic. Plagiarism detection in Arabic documents is challenging because of Arabic’s rich morphological features, lexical variation, and syntactic complexity, which limit the effectiveness of some detection approaches. To address these challenges, this study introduces an external plagiarism detection framework built on an artificial neural network (ANN) model and a lexical feature extraction framework adapted to the linguistic features of Arabic. The proposed framework is evaluated using ExAraPlagDet-2015 benchmark, where a baseline model using support vector machine (SVM) is introduced for comparison. Experimental results demonstrate notable improvements in plagiarism detection performance of the proposed framework compared with SVM and other baseline methods. The proposed framework provides a precision value of 92% and an F-score value of 96%, verifying its effectiveness for Arabic plagiarism detection.
Keywords
Arabic documents; Artificial neural network; Longest common subsequence; Plagiarism detection; Support vector machine
Full Text:
PDFDOI: https://doi.org/10.11591/eei.v15i4.11767
Refbacks
- There are currently no refbacks.

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Bulletin of Electrical Engineering and Informatics (BEEI)
ISSN: 2089-3191
,
e-ISSN: 2302-9285
This journal is published by the
Institute of Advanced Engineering and Science (IAES)
.