Advancing Automatic Sarcasm Detection in the French Language

Document Type : Research Paper

Authors

1 Assistant Prof., Senior Lecturer, Ph.D., Department of Mathematics and Computer Science, Faculty of Science, University of Maroua, P.O Box:814 Maroua, Cameroon; Laboratoire de Recherche en Sciences Informatiques et Applications (LRSIA), UAC, Abomey-Calavi, Benin.

2 Ph.D. Candidate, Department of Mathematics and computer Science, Faculty of Science, University of Maroua, Maroua, Cameroon.

3 Associate Prof., Department of Computer Science and telecommunications, National Advanced School of Engineering of Maroua, University of Maroua, Maroua, Cameroon.

4 Opscidia, 9 Rue des colonnes, Paris, France.

10.22059/jitm.2026.414898.4482

Abstract

The proliferation of user-generated content necessitates robust sentiment analysis and opinion mining. Sarcasm detection presents a significant challenge in natural language processing, as sarcastic statements often invert literal meaning. Accurate sarcasm detection is crucial for enhancing sentiment analysis, opinion mining, recommender systems, and public opinion monitoring. While transformer-based and deep learning models have advanced sarcasm detection in English and other languages, French remains underdeveloped due to limited annotated datasets and the absence of models tailored to its linguistic nuances. Consequently, existing methods struggle to generalize and fail to leverage complementary semantic information to improve sarcasm recognition. This paper introduces a novel hybrid multi-task learning framework for French sarcasm detection. This architecture integrates transfer learning with sentiment analysis by combining a fine-tuned CamemBERT encoder, an auxiliary sentiment-analysis module, and a Bi-LSTM classifier. This approach jointly models contextual and affective information to better distinguish between sarcastic and non-sarcastic text. A significant contribution of this study is the development and release of a large-scale French sarcasm dataset comprising approximately 22,000 manually collected and curated annotated instances from diverse sources, providing a valuable benchmark for future research. Another key contribution is the proposed novel hybrid multi-task architecture that integrates transfer learning, sentiment analysis, and Bi-LSTM-based classification for improved sarcasm detection. Empirical evidence shows that sentiment-aware representations significantly improve sarcasm detection over standard transformer approaches. Experiments on the TransCasm dataset and the new corpus demonstrate the model's effectiveness, achieving F1-scores of 96% and 90%, respectively, surpassing prior results.

Keywords

Main Subjects


Abuein, Q., Ra'ed, M., Migdady, A., Jawarneh, M. S., & Al-Khateeb, A. (2024). ArSa-Tweets: A novel
Aggarwal, A., Wadhawan, A., Chaudhary, A., & Maurya, K. (2020). “Did you really mean what you said?”: Sarcasm detection in Hindi-English code-mixed data using bilingual word embeddings. In Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020) (pp. 7-15).
Al-Ghadhban, D., Alnkhilan, E., Tatwany, L., & Alrazgan, M. (2017). Arabic sarcasm detection in Twitter. In 2017 International Conference on Engineering & MIS (ICEMIS) (pp. 1-7). IEEE.
Alharbi, A. I., & Lee, M. (2021). Multi-task learning using a combination of contextualised and static word embeddings for Arabic sarcasm detection and sentiment analysis. In Proceedings of the sixth Arabic natural language processing workshop (pp. 318-322).
          Arabic sarcasm detection system based on deep learning model. Heliyon10(17).
Ayyasamy S. (2021). Construction of a hybrid model for English news headline sarcasm detection by word
embedding technique. Journal of Electrical Engineering and Automation, 3:184–198, 11.
Băroiu, A. C., & Trăușan-Matu, Ș. (2022). Automatic sarcasm detection: Systematic literature review. Information13(8), 399.
Bharti, S. K., Sathya Babu, K., & Jena, S. K. (2017). Harnessing online news for sarcasm detection in Hindi tweets. In International conference on pattern recognition and machine intelligence (pp. 679-686). Cham: Springer International Publishing.
Derbala Yacoub, A., Elsayed Aboutabl, A., & O Slim, S. (2024). Multilingual sarcasm detection for enhancing sentiment analysis using deep learning algorithms. Journal of Communications Software and Systems20(4), 278-289.
Dubey, A., Joshi, A., & Bhattacharyya, P. (2019). Computational sarcasm for different languages: A survey.
Farabi, S., Ranasinghe, T., Kanojia, D., Kong, Y., & Zampieri, M. (2024). A survey of multimodal sarcasm detection. arXiv preprint arXiv:2410.18882.
Farhan, S., Shoukat, R., & Aslam, A. (2024). Automatic Sarcasm Detection on Cross-Platform Social Media Datasets: A GLoVe and Bi-LSTM Based Approach. Journal of Universal Computer Science (JUCS)30(5).
Frenda, S. (2023). Sarcasm and Implicitness in Abusive Language Detection: A Multilingual Perspective. Procesamiento del Lenguaje Natural70, 239-242.
Ganganwar, V., Manvainder, Singh, M., Patil, P., & Joshi, S. (2024). Sarcasm and humor detection in code-mixed Hindi data: A survey. In International Conference on Computing and Machine Learning (pp. 453-469). Singapore: Springer Nature Singapore.
Gong, X., Zhao, Q., Zhang, J., Mao, R., & Xu, R. (2020). The design and construction of a Chinese sarcasm dataset. In Proceedings of the twelfth language resources and evaluation conference (pp. 5034-5039).
Hassan, M. A., García-Méndez, S., & de Arriba-Pérez, F. (2026). Contribution to Sarcasm Detection in Arabic Using Natural Language Processing Techniques. Applied Sciences16(6), 2724. https://doi.org/10.3390/app16062724
Heraldi, F. D., & Ruskanda, Z. (2024). Effective intended sarcasm detection using fine-tuned llama 2 large language models. In 2024 11th International Conference on Advanced Informatics: Concept, Theory and Application (ICAICTA) (pp. 1-6). IEEE.
Ibrahim, S. A. M., Deshpande, M. M., & Pawar, V. N. (2025). Automated Sarcasm Detection in English Tweets Using CCNN and ELLSTM with Text and Emoji Embeddings. In 2025 International Conference on Intelligent and Innovative Technologies in Computing, Electrical and Electronics (IITCEE) (pp. 1-8). IEEE.
Jia, C., & Zan, H. (2022). Context-Based Sarcasm Detection Model in Chinese Social Media Using BERT and Bi-GRU Models. In CONF-SPML (pp. 42-50).
Jose, J. M., & Benedict, S. (2023). Deepasd framework: A deep learning-assisted automatic sarcasm detection in facial emotions. In 2023 8th International Conference on Communication and Electronics Systems (ICCES) (pp. 998-1004). IEEE.
Kumar, A., Sangwan, S. R., Singh, A. K., & Wadhwa, G. (2023). Hybrid deep learning model for sarcasm detection in Indian indigenous language using word-emoji embeddings. ACM Transactions on Asian and Low-Resource Language Information Processing22(5), 1-20.
Lin, S. K., & Hsieh, S. K. (2016). Sarcasm detection in Chinese using a crowdsourced corpus. In Proceedings of the 28th Conference on Computational Linguistics and Speech Processing (ROCLING 2016) (pp. 299-310).
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., ... & Stoyanov, V. (2019). Roberta: A robustly optimized BERT pre-training approach. arXiv preprint arXiv:1907.11692.
Martin, L., Muller, B., Suarez, P. O., Dupont, Y., Romary, L., de La Clergerie, É. V., ... & Sagot, B. (2020). CamemBERT: a tasty French language model. In Proceedings of the 58th annual meeting of the Association for Computational Linguistics (pp. 7203-7219).
Mazumder, D., Kumar, A., & Patro, J. (2024). Revealing the impact of synthetic native samples and multi-tasking strategies in Hindi-English code-mixed humour and sarcasm detection. arXiv preprint arXiv:2412.12761.
Mihi, S., Ait Ben Ali, B., & Laachfoubi, N. (2022). A Comparative Review of Tweets Automatic Sarcasm Detection in Arabic and English. In International conference on advanced intelligent systems for sustainable development (pp. 841-849). Cham: Springer Nature Switzerland.
Mihi, S., Ait Benali, B., & Laachfoubi, N. (2023). Automatic sarcasm detection in Arabic tweets: resources and approaches. Journal of Intelligent & Fuzzy Systems45(6), 9483-9497.
Naik, P. K., Chenjeri, S. S., Sruthy, S., & Mamatha, H. (2022). Sarcasm Detection in English Text using Tweets and Headlines. In 2022 International Conference on Distributed Computing, VLSI, Electrical Circuits and Robotics (DISCOVER) (pp. 192-196). IEEE.
Obeidat, R., Bashayreh, A., & Younis, L. B. (2022). The impact of combining Arabic sarcasm detection datasets on the performance of a BERT-based model. In 2022 13th International Conference on Information and Communication Systems (ICICS) (pp. 22-29). IEEE.
Pandey, R., Kumar, A., Singh, J. P., & Tripathi, S. (2025). A hybrid convolutional neural network for sarcasm detection from multilingual social media posts. Multimedia Tools and Applications84(16), 15867-15895.
Peled, L., & Reichart, R. (2017). Sarcasm SIGN: Interpreting sarcasm with sentiment-based monolingual machine translation. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 1690-1700).
Pokhriyal, H., & Jain, G. (2025). Supposititious sarcasm detection and sentiment analysis coping with Hindi language in social networks harnessing Zipf-Mandelbrot probabilistic optimisation and perplexity entropy learning. ACM Transactions on Asian and Low-Resource Language Information Processing24(2), 1-28.
Prajapati, A., Singh, A., Arora, V., & Bansal, N. (2024). Sarcasm detection in English-Hindi code-mixed text data. In 2024 2nd International Conference on Advances in Computation, Communication and Information Technology (ICAICCIT) (Vol. 1, pp. 94-99). IEEE.
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language models are unsupervised multitask learners. OpenAI blog1(8), 9.
Rašl, M., Žalik, M., & Keršič, V. (2021). Transformer-based Sarcasm Detection in English and Slovene Language. In Proceedings of the 2021 7th Student Computer Science Research Conference (p. 43).
Sharma, E., N. K. Gondhi, Chaahat, M. Murugappan, G. Naik, and L. Murugan. (2026). “Toward Effective Sarcasm Detection in Social Media: A Review of Artificial Intelligence-Based Automated Approaches.” Expert Systems43, no. 3: e70203. https://doi.org/10.1111/exsy.70203.
Simon, D., Castilho, S., Lohar, P., & Afli, H. (2022). TransCasm: A Bilingual Corpus of Sarcastic Tweets. In Proceedings of the LREC 2022 workshop on Natural Language Processing for Political Sciences (pp. 98-103).
Sukhavasi, V., & Dondeti, V. (2024). Effective automated transformer model-based sarcasm detection using multilingual data. Multimedia Tools and Applications83(16), 47531-47562.
Swami, S., Khandelwal, A., Singh, V., Akhtar, S. S., & Shrivastava, M. (2018). A corpus of English-Hindi code-mixed tweets for sarcasm detection. arXiv preprint arXiv:1805.11869.
Thorat, M., & Shaikh, N. F. (2024). Novel Deep Neural Network Approach for the Sarcasm Detection in Hindi Language. International Journal of Intelligent Systems and Applications in Engineering12(10s), 487–494. 
Tian, Y., & Yang, S. (2024). Chinese Sarcasm Detection Using Title-Attention and Bidirectional LSTM Based on Bert. In 2024 13th International Conference of Information and Communication Technology (ICTech) (pp. 34-38). IEEE.
Wu, B., Tian, H., Liu, X., Hu, W., Yang, C., & Li, S. (2024). Sarcasm detection in Chinese and English text with fine-tuned large language models. In 2024 IEEE 10th Conference on Big Data Security on Cloud (BigDataSecurity) (pp. 47-51). Ieee.
Yue, T., Shi, X., Mao, R., Hu, Z., & Cambria, E. (2024). SarcNet: a multilingual multimodal sarcasm detection dataset. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) (pp. 14325-14335).
Zhang, L., Zhao, X., Song, X., Fang, Y., Li, D., & Wang, H. (2022). A novel Chinese sarcasm detection model based on retrospective reader. In International Conference on Multimedia Modeling (pp. 267-278). Cham: Springer International Publishing.