Journal of Digital Content Management

Journal of Digital Content Management

Reviewing Generative AI Models for Answering Persian Historical Questions

Document Type : Original Article

Authors
1 Department of Computer Scienc, Faculty of Statistics, Mathematics and Computer Scienc, Allameh Tabatabai, Tehran, Iran
2 Department of History, Faculty of Perisan Literature and Foreign Languages, Allameh Tabatabai, Tehran, Iran
Abstract
Purpose: The present research proposes a retrieval-augmented generation approach to address this challenge for answering historical Persian questions.
Method: In this study, sixteen volumes of Persian translations of historical books were used as the dataset. After undergoing preprocessing the texts were prepared to be retrievable and usable for the question-answering process. The Qwen2.5 7b language model was employed as the core engine for generating answers, utilizing information retrieved from historical sources to produce more accurate and well-documented responses.
Findings: Quantitative evaluation demonstrated that the proposed model showed significant improvement compared to the base Qwen2.5 model. Specifically, the BLEU score increased from 0.48 to 0.90, and in terms of sentence similarity, the proposed model's performance was superior to that of the base model, managing to partially close the gap with the more powerful GPT-4 model. The results indicate that the proposed model not only outperforms its base version but also achieves competitive, near GPT-4 performance on certain metrics.
Conclusion: These improvements suggest that combining Qwen2.5 7b with a retrieval-based approach can be an effective strategy for enhancing the accuracy and reliability of language models in Persian applications.
Keywords
Subjects

Chen, X., Gao, P., Song, J., & Tan, X. (2024). HiQA: A hierarchical contextual augmentation RAG for multi-documents QA. ArXiv Preprint ArXiv:2402.01767.
Cherubini, M., Romano, F., Bolioli, A., De Mattei, L., & Sangermano, M. (2024). Improving the accessibility of EU laws: the Chat-EUR-Lex project.
Chouhan, A., & Gertz, M. (2024). LexDrafter: Terminology Drafting for Legislative Documents Using Retrieval Augmented Generation. In N. Calzolari, M.-Y. Kan, V. Hoste, A. Lenci, S. Sakti, & N. Xue (Eds.), Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) (pp. 10448–10458). ELRA and ICCL. https://aclanthology.org/2024.lrec-main.913/
De Freitas, B. A. T., & de Alencar Lotufo, R. (2024). Retail-GPT: leveraging Retrieval Augmented Generation (RAG) for building E-commerce Chat Assistants. ArXiv, abs/2408.08925. https://api.semanticscholar.org/CorpusID:271903742
Douze, M., Guzhva, A., Deng, C., Johnson, J., Szilvasy, G., Mazaré, P.-E., Lomeli, M., Hosseini, L., & Jégou, H. (2024). The Faiss library.
Ferreira, P., Zolfagharnasab, M. H., Gonçalves, T., Bonci, E., Mavioso, C., Cardoso, M. J., & Cardoso, J. S. (2025). Predicting Aesthetic Outcomes of Breast Cancer Surgery: A Robust and Explainable Image Retrieval Approach. Deep Breast Workshop on AI and Imaging for Diagnostic and Treatment Challenges in Breast Care, 268–278.
Giglou, H. B., Taffa, T. A., Abdullah, R., Usmanova, A., Usbeck, R., D’Souza, J., & Auer, S. (2024). Scholarly Question Answering using Large Language Models in the NFDI4DataScience Gateway. http://arxiv.org/abs/2406.07257
Habib, M. A., University, H., Pakistan, P., Amin, S., Jaipal, S., Khan, M. J., & Samad, A. (2024). TaxTajweez: A Large Language Model-based Chatbot for Income Tax Information In Pakistan Using Retrieval Augmented Generation (RAG) Muhammad Oqba.
Haghdadi, A., Zolfagharnasab, M. H., Damari, S., & Vakili, S. (2026). Intelligent Inventory Rotation and Revenue Optimization Using Integer Linear Programming: A Coffee Shop Case Study. In A. I. Pereira & others (Eds.), Optimization, Learning Algorithms and Applications (Vol. 2617, pp. 1–15). Springer Nature Switzerland. https://doi.org/10.1007/978-3-032-00137-5_3
Karbalaeipour, H., Damari, S., Zolfagharnasab, M. H., & Haghdadi, A. (2023). A Collection of 120 Psychology Patients with 17 Essential Symptoms to Diagnose Mania Bipolar Disorder, Depressive Bipolar Disorder, Major Depressive Disorder, and Normal Individuals. Harvard Dataverse. https://doi.org/10.7910/DVN/0FNET5
Lála, J., O’Donoghue, O., Shtedritski, A., Cox, S., Rodriques, S. G., & White, A. D. (2023). PaperQA: Retrieval-Augmented Generative Agent for Scientific Research. http://arxiv.org/abs/2312.07559
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. http://arxiv.org/abs/2005.11401
Li, Y., Li, Z., Zhang, K., Dan, R., Jiang, S., & Zhang, Y. (2023). ChatDoctor: A Medical Chat Model Fine-Tuned on a Large Language Model Meta-AI (LLaMA) Using Medical Domain Knowledge. Cureus. https://doi.org/10.7759/cureus.40895
Long, C., Subburam, D., Lowe, K., dos Santos, A., Zhang, J., Hwang, S., Saduka, N., Horev, Y., Su, T., Côté, D. W. J., & Wright, E. D. (2024). ChatENT: Augmented Large Language Model for Expert Knowledge Retrieval in Otolaryngology–Head and Neck Surgery. Otolaryngology - Head and Neck Surgery (United States), 171(4), 1042–1051. https://doi.org/10.1002/ohn.864
Lozano, A., Fleming, S. L., Chiang, C.-C., & Shah, N. (2023). Clinfo.ai: An Open-Source Retrieval-Augmented Large Language Model System for Answering Medical Questions using Scientific Literature. http://arxiv.org/abs/2310.16146
Mansurova, A., Tleubayeva, A., Nugumanova, A., Shomanov, A., & Seker, S. E. (2025). A Systematic Evaluation of Large Language Models and Retrieval-Augmented Generation for the Task of Kazakh Question Answering. Information, 16(11). https://doi.org/10.3390/info16110943
Montenegro, H., Zolfagharnasab, M. H., Teixeira, F., Pinto, G., Santos, J., Ferreira, P., Bonci, E.-A., Mavioso, C., Cardoso, M. J., & Cardoso, J. S. (2026). Automatic prediction and evaluation of aesthetic outcomes in plastic and oncological surgery: a systematic review. Archives of Computational Methods in Engineering. https://doi.org/10.1007/s11831-026-10508-8
Pinto, G., Zolfagharnasab, M. H., Teixeira, L. F., Cruz, H., Cardoso, M. J., & Cardoso, J. S. (2025). Towards Utilizing Robust Radiance Fields for 3D Reconstruction of Breast Aesthetics. Deep Breast Workshop on AI and Imaging for Diagnostic and Treatment Challenges in Breast Care, 279–288.
Saghayan, M. H., Zolfagharnasab, M. H., Khadem, A., Matinfar, F., & Rashidi, H. (2023). Diagnosing Bipolar Disorder from 3-D Structural Magnetic Resonance Images Using a Hybrid GAN-CNN Method. ArXiv Preprint ArXiv:2310.07359.
Shui, R., Cao, Y., Wang, X., & Chua, T.-S. (2023). A Comprehensive Evaluation of Large Language Models on Legal Judgment Prediction. In H. Bouamor, J. Pino, & K. Bali (Eds.), Findings of the Association for Computational Linguistics: EMNLP 2023 (pp. 7337–7348). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.findings-emnlp.490
Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., Scales, N., Tanwani, A., Cole-Lewis, H., Pfohl, S., Payne, P., Seneviratne, M., Gamble, P., Kelly, C., Scharli, N., Chowdhery, A., Mansfield, P., Arcas, B. A. y, Webster, D., … Natarajan, V. (2022). Large Language Models Encode Clinical Knowledge. Nature. http://arxiv.org/abs/2212.13138
VibrantLabs. (2024). Ragas: Supercharge Your LLM Application Evaluations.
Vidyarthi, A., Singh, K. M., & Moirangthem, D. S. (2026). SageRAG: Query rewriting for retrieval enhancement and retrieval-augmented generation for grounded responses in AI research assistance. Expert Systems with Applications, 308, 131160. https://doi.org/https://doi.org/10.1016/j.eswa.2026.131160
Wang, C., Long, Q., Xiao, M., Cai, X., Wu, C., Meng, Z., Wang, X., & Zhou, Y. (2024). BioRAG: A RAG-LLM Framework for Biological Question Reasoning. http://arxiv.org/abs/2408.01107
Wang, D., Liang, J., Ye, J., Li, J., Li, J., Zhang, Q., Hu, Q., Pan, C., Wang, D., Liu, Z., Shi, W., Shi, D., Li, F., Qu, B., & Zheng, Y. (2024). Enhancement of the Performance of Large Language Models in Diabetes Education through Retrieval-Augmented Generation: Comparative Study. Journal of Medical Internet Research, 26. https://doi.org/10.2196/58041
Wang, X., Chi, J., Tai, Z., Kwok, T. S. T., Li, M., Li, Z., He, H., Hua, Y., Lu, P., Wang, S., Wu, Y., Huang, J., Tian, J., Mo, F., Cui, Y., & Zhou, L. (2025). FinSage: A Multi-aspect RAG System for Financial Filings Question Answering. http://arxiv.org/abs/2504.14493
Yang, Q. A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Dong, G., Wei, H., Lin, H., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., Lin, J., … Wang, Z. (2024). Qwen2.5 Technical Report. ArXiv, abs/2412.15115. https://api.semanticscholar.org/CorpusID:274859421
Zhang, X., Zhang, Y., Long, D., Xie, W., Dai, Z., Tang, J., Lin, H., Yang, B., Xie, P., Huang, F., Zhang, M., Li, W., & Zhang, M. (2024). mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval. In F. Dernoncourt, D. Preo\c{t}iuc-Pietro, & A. Shimorina (Eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track (pp. 1393–1412). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.emnlp-industry.103
Zolfagharnasab, M. H., Aghanajafi, C., Kavian, S., Heydarian, N., & Ahmadi, M. H. (2020). Novel analysis of second law and irreversibility for a solar power plant using heliostat field and molten salt. Energy Science & Engineering, 8(11), 4136–4153.
Zolfagharnasab, M. H., Damari, S., Soltani, M., Ng, A., Karbalaeipour, H., Haghdadi, A., Saghayan, M. H., & Matinfar, F. (2025). A novel rule-based expert system for early diagnosis of bipolar and Major Depressive Disorder. Smart Health, 35, 100525. https://doi.org/https://doi.org/10.1016/j.smhl.2024.100525
Zolfagharnasab, M. H., Freitas, N., Gonçalves, T., Bonci, E., Mavioso, C., Cardoso, M. J., Oliveira, H. P., & Cardoso, J. S. (2024). Predicting Aesthetic Outcomes in Breast Cancer Surgery: A Multimodal Retrieval Approach. Deep Breast Workshop on AI and Imaging for Diagnostic and Treatment Challenges in Breast Care, 137–147.
Zolfagharnasab, M. H., Gonalves, T., Ferreira, P., Cardoso, M. J., & Cardoso, J. S. (2025). Towards Robust Breast Segmentation: Leveraging Depth Awareness and Convexity Optimization For Tackling Data Scarcity. Deep Breast Workshop on AI and Imaging for Diagnostic and Treatment Challenges in Breast Care, 41–51.
Zolfagharnasab, M. H., PourMohammadBagher, L., & Bahrani, M. (2024). Intelligent Travel Recommendations Using Neural Collaborative Filtering for Touristic Landmarks of Iran. Journal of Data Science and Modeling, 2(2), 119–147. https://doi.org/10.22054/jdsm.2025.82861.1059
Zolfagharnasab, M. H., Saghayan, M. H., Pedram, M. Z., Vafai, K., & Hoseinzadeh, S. (2023). A numerical study of the nanofluid mixtures inside a Buoyancy-driven cavity in the presence of a variable magnetic field. Energy Reports, 10, 973–988.

  • Receive Date 24 December 2025
  • Revise Date 25 June 2026
  • Accept Date 22 July 2026