Fusion of Graph-Based Ensembles Using Fuzzy Integrals for Persian Learning to Rank

Document Type : Research Paper

Authors

1 Department of Algorithms and Computation, School of Engineering Sciences, College of Engineering, University of Tehran, Tehran, Iran.

2 Professor, Department of Algorithms and Computation, School of Engineering Sciences, College of Engineering, University of Tehran, Tehran, Iran.

3 Assistant prof., ICT Research Institute, Tehran, Iran.

4 Assistant prof., Computer Engineering Department, Faculty of Engineering, College of Farabi, University of Tehran, Tehran, Iran.

10.22059/jitm.2026.412047.4443

Abstract

The rapid growth of Persian-language Web content has created an urgent need for effective Learning to Rank (LTR) systems. However, developing such systems for Persian presents significant challenges, including Persian’s morphological complexity, its unique Web link topology, and the scarcity of annotated resources relative to dominant languages such as English. Existing methods often rely on hand-crafted features, apply graph techniques in isolation, or employ simple ensemble fusion, lacking a unified framework to integrate structural semantics and handle model uncertainty. This paper proposes PersianRank, a novel LTR framework that bridges these gaps by modeling query-document relationships as a weighted bipartite graph to extract graph-based structural features. It applies deep neural networks to predict these graph features for unseen query-document pairs, enabling feature augmentation during the testing phase. A diverse ensemble of base rankers is then trained on the enriched feature set, and their predictions are fused via the Choquet fuzzy integral, which performs a non-linear, uncertainty-aware aggregation. Evaluated on the dotIR dataset, PersianRank achieves state-of-the-art performance, improving P@1 by 17.1% relative to the best baseline (MDPRank) and demonstrating notable top-rank accuracy through effective graph-based feature space expansion and uncertainty-aware model fusion. The framework is especially effective for top-rank retrieval, a critical aspect in user-centric search systems.

Keywords

Main Subjects


Asgari-Bidhendi, M., Fakhrian, F., & Minaei-Bidgoli, B. (2020). A Graph-based Approach for Persian Entity Linking. International Journal of Information and Communication Technology Research, 12(4), 60-69. http://ijict.itrc.ac.ir/article-1-472-en.html
Azarafza M., Feizi-Derakhshi MR., & Bagheri-Shendi M. (2020). TextRank-based Microblogs Keyword Extraction Method for Persian Language. In Proceedings of the 3rd International Congress on Science and Engineering (pp. 179-188).
Azarbonyad, H., Shakery, A., & Faili, H. (2014). Learning to Exploit Different Translation Resources for Cross Language Information Retrieval. International Journal of Information and Communication Technology Research, 6(1), 55-68. http://ijict.itrc.ac.ir/article-1-138-en.html
Azarbonyad, H., Shakery, A., & Faili, H. (2019). A learning to rank approach for cross-language information retrieval exploiting multiple translation resources. Natural Language Engineering, 25(3), 363-384. https://doi.org/10.1017/S1351324919000032
Barabási, A. L., & Pósfai, M. (2016). Network Science. Cambridge University Press.
Baradaran Hashemi, H., Yazdani, N., Shakery, A., & Pakdaman Naeini, M. (2010). Application of ensemble models in Web ranking. In Proceedings of the 2010 5th International Symposium on Telecommunications (pp. 726–731).
Bostan, S., Bidoki, A. M. Z., & Pajoohan, M. R. (2023). Improving Ranking Using Hybrid Custom Embedding Models on Persian Web. Journal of Web Engineering, 22(5), 797-820. https://doi.org/10.13052/JWE1540-9589.2253
Breiman, L. (2001). Random Forests. Machine Learning, 45(1), 5-32. https://doi.org/10.1023/A:1010933404324
Burges, C. J. C. (2010). From ranknet to lambdarank to lambdamart: An overview. Technical Report MSR-TR-2010-82. http://www.ccs.neu.edu/home/vip/teach/MLcourse/4_boosting/materials/msr-tr-2010-82.pdf
Darrudi, E., Hashemi, H. B., AleAhmad, A., Zare Bidoki, A. M., Habibian, A., Mahdikhani, F., & Rahgozar, M. (2009). dotIR collection for Persian Web retrieval. Technical Report UT-DBRG-402.  
Derhami, V., Khodadadian, E., Ghasemzadeh, M., & Zareh Bidoki, A. M. (2013). Applying reinforcement learning for Web pages ranking algorithms. Applied Soft Computing, 13(4), 1686-1692. https://doi.org/10.1016/J.ASOC.2012.12.023
Derhami, V., Paksima, J., & Khajah, H. (2015). Web pages ranking algorithm based on reinforcement learning and user feedback. Journal of AI and Data Mining, 3(2), 157–168. https://doi.org/10.5829/idosi.jaidm.2015.03.02.05
Derhami, V., Paksima, J., & Khajeh, H. (2019). RRLUFF: Ranking function based on Reinforcement Learning using User Feedback and Web Document Features. Journal of AI and Data Mining, 7(3), 421–442. https://doi.org/10.22044/JADM.2019.3547.1814
Eberhard, D. M., Simons, G. F., & Fennig, C. D. (2025). Languages of the World. In Ethnologue (28th ed.). SIL International. https://www.ethnologue.com
Freund, Y., Iyer, R., Schapire, R. E., Singer, Y., & Dietterich, T. G. (2003). An efficient boosting algorithm for combining preferences. The Journal of Machine Learning Research, 4, 933–969. https://doi.org/10.5555/945365.964285
Friedman, J. H. (2001). Greedy function approximation: A gradient boosting machine. The Annals of Statistics, 29(5), 1189–1232. https://doi.org/10.1214/AOS/1013203451
Ghanbari, E., & Shakery, A. (2018). Query-dependent learning to rank for cross-lingual information retrieval. Knowledge and Information Systems 2018 59:3, 59(3), 711–743. https://doi.org/10.1007/S10115-018-1232-8
Ghanbari, E., & Shakery, A. (2021). A Learning to rank framework based on cross-lingual loss function for cross-lingual information retrieval. Applied Intelligence 2021, 52(3), 3156–3174. https://doi.org/10.1007/S10489-021-02592-Z
Ibrahim, O. A. S., & Landa-Silva, D. (2017). ES-rank: Evolution strategy Learning to Rank approach. In Proceedings of the ACM Symposium on Applied Computing (pp. 944–950).
Joachims, T. (2002). Optimizing search engines using clickthrough data. In Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 133–142).
Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., & Liu, T. Y. (2017). LightGBM: A Highly Efficient Gradient Boosting Decision Tree. In Proceedings of the Advances in Neural Information Processing Systems (Vol. 30, pp. 3149–3157).
Keyhanipour, A. H. (2019a). Application of Learning to Rank on Persian Web Content. In Proceedings of the 5th IEEE International Conference on Web Research (Vol. 5, pp. 1–14).
Keyhanipour, A. H. (2019b). Effective Learning to Rank for the Persian Web Content. Journal of Information Technology Management, 11(4), 92–109. https://doi.org/10.22059/JITM.2019.284726.2377
Keyhanipour, A. H. (2020). Learning to Rank for the Persian Web Using the Layered Genetic Programming. Journal of Information and Communication Technology, 10(37), 45–70. https://jour.aicti.ir/en/Article/8179/FullText
Keyhanipour, A. H. (2024). Graph-based comparative analysis of learning to rank datasets. International Journal of Data Science and Analytics, 17(2), 165–187. https://doi.org/10.1007/S41060-023-00406-8/METRICS
Keyhanipour, A. H., Moshiri, B., & Rahgozar, M. (2015). CF-Rank: Learning to rank by classifier fusion on click-through data. Expert Systems with Applications, 42(22), 8597–8608. https://doi.org/10.1016/j.eswa.2015.07.014
Keyhanipour, A. H., Moshiri, B., Piroozmand, M., Oroumchian, F., & Moeini, A. (2016). Learning to rank with click-through features in a reinforcement learning framework. International Journal of Web Information Systems, 12(4), 448–476. https://doi.org/10.1108/IJWIS-12-2015-0046
Latifi, S., & Nematbakhsh, M. (2014). Query-independent learning to rank RDF entity results of SPARQL queries. In Proceedings of the 4th International Conference on Computer and Knowledge Engineering (pp.  297–301).
Manning, C. D., Raghavan, P., & Schütze, H. (2008). Introduction to Information Retrieval. Cambridge University Press.
Mirzababaei, B., Faili, H., & Ehsan, N. (2013). Discourse-aware Statistical Machine Translation as a Context-Sensitive Spell Checker. In Proceedings of Recent Advances in Natural Language Processing (pp. 475–482).
Moosavi, N. S. G.-S. G. (2009). A Ranking Approach to Persian Pronoun Resolution. Advances in Computational Linguistics, 41, 169–180.
Newman, M. (2018). Networks (Second Edition). Oxford University Press.
Qin, T., & Liu, T. Y. (2013). Introducing LETOR 4.0 Datasets. arXiv. http://arxiv.org/abs/1306.2597
Rahimi, R., Shakery, A., Dadashkarimi, J., Ariannezhad, M., Dehghani, M., & Esfahani, H. N. (2016). Building a multi-domain comparable corpus using a learning to rank method†. Natural Language Engineering, 22(4), 627–653. https://doi.org/10.1017/S1351324916000164
Rahimi, Z., & Shamsfard, M. (2023). A Neuro Symbolic Approach for Contradiction Detection in Persian Text. Journal of Universal Computer Science, 29(3), 242–264. https://doi.org/10.3897/jucs.90646
Rahimi, Z., & ShamsFard, M. (2024). A Knowledge-Based Approach for Recognizing Textual Entailments with a Focus on Causality and Contradiction. Available at SSRN 4526759. https://doi.org/10.21203/RS.3.RS-3826973/V1
Rezvani, M., & Hashemi, S. M. (2012). Enhancing accuracy of topic sensitive PageRank using Jaccard index and cosine similarity. In Proceedings of the 2012 IEEE/WIC/ACM International Conference on Web Intelligence (WI 2012) (pp. 620–624).
Rogers, R. (2010). Internet Research: The Question of Method—A Keynote Address from the YouTube and the 2008 Election Cycle in the United States Conference. Journal of Information Technology & Politics, 7(2–3), 241–260. https://doi.org/10.1080/19331681003753438
Shyam, K. P., Jagannathan, S., & Rajavel, M. (2013). A New Technique for Ranking Web Pages and Adwords. International Journal of Computer Applications, 82(12), 32–36. https://doi.org/10.5120/14171-2378
Syriopoulos, P. K., Kalampalikis, N. G., Kotsiantis, S. B., & Vrahatis, M. N. (2025). kNN Classification: a review. Annals of Mathematics and Artificial Intelligence, 93(1), 43–75. https://doi.org/10.1007/s10472-023-09882-x
Thakur, N., Kazi, S., Luo, G., Lin, J., & Ahmad, A. (2025). MIRAGE-Bench: Automatic Multilingual Benchmark Arena for Retrieval-Augmented Generation Systems. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 274–298).
Veisi, H., & Shandi, H. F. (2020). A Persian Medical Question Answering System. International Journal on Artificial Intelligence Tools, 29(6), 205-219. https://doi.org/10.1142/S0218213020500190
Wei, Z., Xu, J., Lan, Y., Guo, J., & Cheng, X. (2017). Reinforcement learning to rank with Markov decision process. In Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 945–948).
Wikipedia. (2026, June 6). Languages used on the Internet. Wikipedia. https://en.wikipedia.org/wiki/Languages_used_on_the_Internet#cite_note-w3techs_historical_trends-1
Wu, Q., Burges, C. J. C., Svore, K. M., & Gao, J. (2010). Adapting boosting for information retrieval measures. Information Retrieval, 13(3), 254–270. https://doi.org/10.1007/S10791-009-9112-1/FIGURES/4
Zeraatkar, A., Mirvaziri, H., & Ahsaee, M. G. (2019). Improvement of Page Ranking Algorithm by Negative Score of Spam Pages. Webology, Volume 16(2), 43–56. https://doi.org/10.14704/WEB/V16I2/A187