GPlagNet: A Deep Semantic Neural Framework for Gujarati Document Plagiarism Detection

Main Article Content

Bhumi Shah
Gaurav Kumar Ameta

Abstract

The problem of plagiarism detection in low resource languages is still an important one because of the scarcity of linguistic resources and the lack of ability to detect semantic similarity in traditional lexical matching methods. Gujarati being one of the major Indian languages, does not have any strong plagiarism detection system which can detect paraphrased or similar documents in the same context. In this research the authors introduce GPlagNet, a deep semantic neural framework for Gujarati document plagiarism detection which integrates language-specific preprocessing with transformer-based representation learning. The proposed framework integrates language-specific normalization, tokenization, stop-word removal, stemming, and lemmatization before generating contextual document representations using the ALBERT transformer. The bi-encoder architecture produces fixed-length document representations, which are compared using cosine similarity to compute semantic similarity scores and rank candidate documents using a Top-K retrieval strategy. The proposed framework produced a score of 1.00 for an identical document and semantic similarity scores of 0.8054 and 0.8606 for lexically different but contextually related documents, indicating that the model captures semantic relatedness beyond lexical overlap. The results show that GPlagNet was able to identify many semantically related documents under the paraphrasing conditions represented in the evaluation dataset, indicating its potential as a promising approach for academic integrity and document similarity assessment in low-resource languages.

Downloads

Download data is not yet available.

Article Details

Section

Articles

How to Cite

Shah, B. and Gaurav Kumar Ameta (2026) “GPlagNet: A Deep Semantic Neural Framework for Gujarati Document Plagiarism Detection”, Journal of Engineering (Iraq), 32(10), pp. 52–74. doi:10.31026/j.eng.2026.10.03.

References

Ahnaf, A., Hasan, H.M.M., Sworna, N.S., and Hossain, N., 2023. An improved extrinsic monolingual plagiarism detection approach of the Bengali text. International Journal of Electrical and Computer Engineering, 13(4), pp. 4256–4267. https://doi.org/10.11591/IJECE.V13I4.PP4256-4267

AlSallal, M., Iqbal, R., Amin, S., James, O., and Palade, V., 2019. An integrated approach for intrinsic plagiarism detection. Future Generation Computer Systems, 96, pp. 700–712. https://doi.org/10.1016/j.future.2017.11.023.

Babic, K., and Mestrovic, A., 2024. Recursively autoregressive autoencoder for pyramidal text representation. IEEE Access, 12, pp. 71361–71370. https://doi.org/10.1109/ACCESS.2024.3402830

Bao, W., Dong, J., Xu, Y., Yang, Y., and Qi, X., 2024. Exploring attentive Siamese LSTM for low-resource text plagiarism detection. Data Intelligence, 6(2), pp. 488–503. https://doi.org/10.1162/dint_a_00242

Chang, C.Y., Jhang, S.J., Wu, S.J., and Roy, D.S., 2024. JCF: Joint coarse- and fine-grained similarity comparison for plagiarism detection based on NLP. Journal of Supercomputing, 80. https://doi.org/10.1007/s11227-023-05472-0.

Chauhan, U., Shah, S., Shiroya, D., Solanki, D., Patel, Z., Bhatia, J., Tanwar, S., Sharma, R., Marina, V., and Raboaca, M.S., 2023. Modeling topics in DFA-based lemmatized Gujarati text. Sensors, 23(5), pp. 1–17. https://doi.org/10.3390/s23052708

Devkar, S., 2014. ટેલિવિઝનની અસરકારકતા, પારિવારીક સંબંધોનું ઉઠમણું. Aksharnaad, 9 May 2014. https://www.aksharnaad.com/2014/05/09/effect-of-television/

Duan, X., Wang, M., and Mu, J., 2017. A plagiarism detection algorithm based on extended Winnowing. MATEC Web of Conferences, 128, pp. 1–5. https://doi.org/10.1051/matecconf/201712802019

El-Rashidy, M.A., Mohamed, R.G., El-Fishawy, N.A., and Shouman, M.A., 2024. An effective text plagiarism detection system based on feature selection and SVM techniques. Multimedia Tools and Applications, 83. https://doi.org/10.1007/s11042-023-15703-4.

El-Rashidy, M.A., Mohamed, R.G., El-Fishawy, N.A., and Shouman, M.A., 2022. Reliable plagiarism detection system based on deep learning approaches. Neural Computing and Applications, 34(21), pp. 18837–18858. https://doi.org/10.1007/s00521-022-07486-w.

Foltýnek, T., Dlabolová, D., and Ruoti, C., 2020. Testing of support tools for plagiarism detection. International Journal of Educational Technology in Higher Education, 17(1). https://doi.org/10.1186/s41239-020-00192-4

Garg, U., and Goyal, V., 2016. Maulik: A plagiarism detection tool for Hindi documents. Indian Journal of Science and Technology, 9(12), pp. 1–11. https://dx.doi.org/10.17485/ijst/2016/v9i12/86631

Haseeb, M., Manzoor, M.F., Farooq, S., and Abid, A., 2024. A versatile dataset for intrinsic plagiarism detection, text reuse analysis, and author clustering in Urdu. Data in Brief, 52, P. 109857. https://doi.org/10.1016/j.dib.2023.109857.

Hosam, E., Hadhoud, M., Atiya, A., and Fayek, M., 2022. Classification feature sets for source code plagiarism detection in Java. Journal of Engineering and Applied Science, 69(1), pp. 1–18. https://doi.org/10.1186/s44147-022-00155-8.

Iqbal, H.R., Maqsood, R., Raza, A.A., and Hassan, S.U., 2024. Urdu paraphrase detection: A novel DNN-based implementation using a semi-automatically generated corpus. Natural Language Engineering, 30(2), pp. 354–384. https://doi.org/10.1017/S1351324923000189

Jiffriya, M., Jahan, M.A., and Ragel, R.G., 2021. Plagiarism detection tools and techniques: A comprehensive survey. Journal of Science-FAS-SEUSL, 2(2), pp. 47–64. https://doi.org/10.48550/arXiv.1801.06323

Khaled, F., and Al-Tamimi, M.S.H., 2021. Plagiarism detection methods and tools: An overview. Iraqi Journal of Science, 62(8), pp. 2771–2783. https://doi.org/10.24996/ijs.2021.62.8.30

Kutbi, M., Al-Hoorie, A.H., and Al-Shammari, A.H., 2024. Detecting contract cheating through linguistic fingerprint. Humanities and Social Sciences Communications, 11(1), pp. 1–9. https://doi.org/10.1057/s41599-024-03160-9

Magdum, V., Dhekane, O.J., Hiwarkhedkar, S.S., Mittal, S.S., and Joshi, R., 2023. MahaNLP: A Marathi natural language processing library. Proceedings of the 13th International Joint Conference on Natural Language Processing (IJCNLP-AACL 2023), 5, pp. 34–40. https://doi.org/10.18653/v1/2023.ijcnlp-demo.5

Mansuri, P.F., 2024. ટેલિવિઝનની કૌટુંબિક જીવન ઉપર થતી અસર. Shikshan Sanshodhan: Journal of Arts, Humanities and Social Sciences, 7(6), pp. 5–8. https://doi.org/10.2018/SS/202406002

Manzoor, M.F., Farooq, S., Haseeb, M., and Abid, A., 2023. Exploring the landscape of intrinsic plagiarism detection: Benchmarks, techniques, evolution, and challenges. IEEE Access, 11, pp. 140519–140545. https://doi.org/10.1109/ACCESS.2023.3338855

Mentari, M., Rozi, I.F., and Rahayu, M.P., 2022. Cross-language text document plagiarism detection system using Winnowing method. Journal of Applied Intelligent System, 7(1), pp. 44–57. https://doi.org/10.33633/jais.v7i1.5950

Modh, J.C., Saini, J.R., and Kotecha, K., 2022. A novel readability complexity score for Gujarati idiomatic text. International Journal of Advanced Computer Science and Applications, 13(5), pp. 453–459. https://doi.org/10.14569/IJACSA.2022.0130553

Moravvej, S.V., Mousavirad, S.J., Moghadam, M.H., and Saadatmand, M., 2021. An LSTM-based plagiarism detection via attention mechanism and a population-based approach for pre-training parameters with imbalanced classes. Lecture Notes in Computer Science, 13110, pp. 690–701. https://doi.org/10.1007/978-3-030-92238-2_57

Moravvej, S.V., Mousavirad, S.J., Oliva, D., Schaefer, G., and Sobhaninia, Z., 2022. An improved DE algorithm to optimise the learning process of a BERT-based plagiarism detection model. Proceedings of the IEEE Congress on Evolutionary Computation (CEC 2022). https://doi.org/10.1109/CEC55065.2022.9870280

Mutsaddi, A., and Choudhary, A., 2025. Enhancing plagiarism detection in Marathi with a weighted ensemble of TF-IDF and BERT embeddings for low-resource language processing. Proceedings of the International Conference on Computational Linguistics (COLING), pp. 89–100. https://aclanthology.org/2025.loreslm-1.6/

Nahian, J.A., Srabon, M.M., Noori, S.R.H., and Masum, A.K.M., 2022. Review on multiple plagiarism: A performance comparison study. Proceedings of the 13th International Conference on Computing Communication and Networking Technologies (ICCCNT 2022). https://doi.org/10.1109/ICCCNT54827.2022.9984577

Paiva, J.C., Leal, J.P., and Figueira, Á., 2025. Clustering source code from automated assessment of programming assignments. International Journal of Data Science and Analytics, 20(2), pp. 1581–1592. https://doi.org/10.1007/s41060-024-00554-5.

Potthast, M., Barrón-Cedeño, A., Stein, B., and Rosso, P., 2011. Cross-language plagiarism detection. Language Resources and Evaluation, 45(1), pp. 45–62. https://doi.org/10.1007/s10579-009-9114-z

Saglam, T., Brodel, M., Schmid, L., and Hahner, S., 2024. Detecting automatic software plagiarism via token sequence normalization. Proceedings of the International Conference on Software Engineering, pp. 1384–1396. https://doi.org/10.1145/3597503.3639192

Saǧlam, T., Hahner, S., Wittler, J.W., and Kühn, T., 2022. Token-based plagiarism detection for metamodels. Proceedings of the ACM/IEEE International Conference on Model Driven Engineering Languages and Systems (MODELS 2022 Companion), pp. 138–141. https://doi.org/10.1145/3550356.3556508

Sajid, M., Sanaullah, M., Fuzail, M., Malik, T.S., and Shuhidan, S.M., 2025. Comparative analysis of text-based plagiarism detection techniques. PLoS ONE, 20(4), pp. 1–28. https://doi.org/10.1371/journal.pone.0319551.

Saqaabi, A.A., Stewart, C., Akrida, E., and Cristea, A.I., 2025. A deep learning approach for paragraph-level paraphrase generation for plagiarism detection. Neural Processing Letters, 57(3), pp. 1–42. https://doi.org/10.1007/s11063-025-11771-9.

Setu, D.M., 2025. A comprehensive strategy for identifying plagiarism in academic submissions. Journal of Umm Al-Qura University for Engineering and Architecture, 16(2), pp. 310–325. https://doi.org/10.1007/s43995-025-00108-1.

Shikshan Sanshodhan, 2019. Gujarati informative article, Paper ID: SS201906004. Shikshan Sanshodhan: Journal of Arts, Humanities and Social Sciences.

Veisi, H., Golchinpour, M., Salehi, M., and Gharavi, E., 2022. Multi-level text document similarity estimation and its application for plagiarism detection. Iran Journal of Computer Science, 5(2), pp. 143–155. https://doi.org/10.1007/s42044-022-00098-6

Wahle, J.P., Ruas, T., Foltýnek, T., Meuschke, N., and Gipp, B., 2022. Identifying machine-paraphrased plagiarism. Lecture Notes in Computer Science, 13192, pp. 393–413. https://doi.org/10.1007/978-3-030-96957-8_34

Wahle, J.P., Ruas, T., Kirstein, F., and Gipp, B., 2022. How large language models are transforming machine-paraphrased plagiarism. Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP 2022), pp. 952–963. https://doi.org/10.18653/v1/2022.emnlp-main.62

Yalcin, K., Cicekli, I., and Ercan, G., 2022. An external plagiarism detection system based on part-of-speech (POS) tag n-grams and word embedding. Expert Systems with Applications, 197. https://doi.org/10.1016/j.eswa.2022.116677

Zahid, M.M., 2023. An efficient machine learning approach for plagiarism detection in text documents. Journal of Computing and Biomedical Informatics, 4(2), pp. 241–248. https://www.jcbi.org/index.php/Main/article/view/153

Zimba, O., and Gasparyan, A.Y., 2021. Plagiarism detection and prevention: A primer for researchers. Reumatologia, 59(3), pp. 132–137. https://doi.org/10.5114/reum.2021.105974

Zouaoui, S., and Rezeg, K., 2022. Multi-agents indexing system (MAIS) for plagiarism detection. Journal of King Saud University - Computer and Information Sciences, 34(5), pp. 2131–2140. https://doi.org/10.1016/j.jksuci.2020.06.009.

Similar Articles

You may also start an advanced similarity search for this article.