The International Arab Journal of Information Technology (IAJIT)

..............................
..............................
..............................


Automated Detection of Albanian Gegë and Toskë Dialects Through Classical and Deep-Learning Classifiers

This study presents a framework for automatically distinguishing between the two major Albanian dialects, Gegë and Toskë , using a hybrid of classic classifiers and deep -learning classifiers. Through the course of the study, we have collected a balanced corpus of dialectal samples from social media posts and online texts. After preprocessing of the dataset, initially three classical algorithms were evaluated: Extreme Gradient Boosting (XGBoost) , due to its robust overfitting control and non -linear performance; Support Vector Machine (SVM), due to margin -based classification in high -dimensional text spaces; and a Recurrent Neu ral Network with Long Short -Term Memory (RNN -LSTM), to capture sequential dependencies in dialectal language. Moreover, after transforming raw text into numerical features, we have compared embeddings from four pretrained multilingual models: Bidirectional Encoder Representations from Transformers ( BERT )-base -multilingual -cased, Nomic -embed - text -v2-moe, Jina -embeddings -v3, and Cross -lingual Language Model -Robustly Optimized BERT Pretraining A pproach (XLM - RoBERTa -large ). Each embedding -classifier pairing was assessed on accuracy and F1 -score. Final results, showed that XGBoost from classical and XLM -RoBERTa -large from deep learning, achieved the highest accuracy. Nevertheless, other algorithms showed promising results such as RNN -LSTM an d BERT, w hich excelled at modeling sequential patterns. Moreover, SVM offered a compelling balance of performance and efficiency, demonstrating the practical value of using classical algorith ms instead of usage of advanced embeddings with targeted classifiers for d ialect identification .

[1] Ajvazi A. and Hardmeier C., “A Dataset of Offensive Language in Kosovo Social Media,” in Proceedings of the 13 th Language Resources and Evaluation Conference , Marseille, pp. 1860 -1869, 2022. https://aclanthology.org/2022.lrec -1.198/

[2] Alsuwaylimi A., “Arabic Dialect Identification in Social Media: A Hybrid Model with Transformer Models and BiLSTM,” Heliyon , vol. 10, no. 17, pp. 1 -18, 2024. https://doi.org/10.1016/j.heliyon.2024.e36280

[3] Althobaiti M., “Country -Level Arabic Dialect Identification Using Small Datasets with Integrated Machine Learning Techniques and Deep Learning Models,” in Proceedings of the 6 th Arabic Natural Language Processing Workshop , Kyiv, pp. 265 -270, 2021. https://aclanthology.org/2021.wanlp -1.30/

[4] Cerpja A. and Cepani A., “Albanian Dialect Classifications,” Dialectologia: Revista Electrònica , vol. 11, pp. 51 -87, 2023. https://www.raco.cat/index.php/Dialectologia/arti cle/view/427733

[5] Ciobanu A., Malmasi S., and Dinu L., “German Dialect Identification Using Classifier Ensembles,” in Proceedings of the 5 th Workshop on NLP for Similar Languages, Varieties and Dialects , New Mexico, pp. 288 -294, 2018. https://aclanthology.org/W18 -3933/

[6] Fetahi E., Hamiti M., Susuri A., Ajdari J., and Zenuni X., “AI -based Hate Speech Detection in Albanian Social Media: New Dataset and Mobile Web Application Integration,” International Journal of Interactive Mobile Technologies , vol. 18, no. 24, pp. 190 -208, 2024. https://online - journals.org/index.php/i -jim/article/view/50851

[7] Hazim L. and Ata O., “Bridging the Gap: Ensemble Learning -Based NLP Framework for AI -Generated Text Identification in Academia,” The International Arab Journal of Information Technology , vol. 22, no. 6, pp. 1055 -1069, 2025. https://doi.org/10.34028/iajit/22/6/2

[8] Hu H., Li W., Zhou H., Tian Z., and et al., “Ensemble Methods to Distinguish Mainland and Taiwan Chinese,” in Proceedings of the 6 th Workshop on NLP for Similar Languages, Varieties and Dialects , Michigan, pp. 165 -171, 2019. https://aclanthology.org/W19 -1417/

[9] Jashari A., “On the Structure of Adjectives’ Superlative Degree Formation by Means of Analytical Tools: A Contrastive Analysis of Albanian and some other Indo -European Languages,” Knowledge -International Journal , vol. 69, no. 5, pp. 847 -852, 2025. https://ojs.ikm.mk/index.php/kij/article/view/7308

[10] Kadriu A., Abazi L., and Abazi H., “Albanian Text Classification: Bag of Words Model and Word Analogies,” Business Systems Research , vol. 10, no. 1, pp. 74 -87, 2019. DOI: 10.2478/bsrj -2019 - 0006

[11] Kastrati M. and Biba M., “Natural Language Processing for Albanian: A State -of-the -Art Survey,” International Journal of Electrical and Computer Engineering , vol. 12, no. 6, pp. 6432 - 6439, 2022. http://doi.org/10.11591/ijece.v12i6.pp6432 -6439

[12] Nuci K., Landes P., and Di Eugenio B., “Ro BERT a Low Resource Fine Tuning for Sentiment Analysis in Albanian,” in Proceedings of the Joint International Conference on Computational Linguistics, Language Resources and Evaluation , Torino, pp. 14146 -14151, 2024. https://aclanthology.org/2024.lrec -main.1233/

[13] Nurce E., Keci J., and Derczynski L., “Detecting Abusive Albanian,” arXiv Preprint , vol. arXiv:2107.13592v3, pp. 1 -10 , 2021. https://arxiv.org/abs/2107.13592

[14] Plaku E., Jahaj K., Cela A., and Civici N., “A Machine Learning Framework for Automated News Article Title Classification in Albanian,” in Proceedings of the International Conference on Innovations in Intelligent Systems and Applications , Craiova, pp. 1 -6, 2024. https://ieeexplore.ieee.org/document/10683815

[15] Priban P. and Taylor S., “ZCU -NLP at MADAR 2019: Recognizing Arabic Dialects,” in Proceedings of the 4 th Arabic Natural Language Processing Workshop , Florence, pp. 208 -213, 2019. https://aclanthology.org/W19 -4623/

[16] Toska M., Nivre J., and Zeman D., “Universal Dependencies for Albanian,” in Proceedings of the 4th Workshop on Universal Dependencies , Barcelona, pp. 178 -188, 2020. https://aclanthology.org/2020.udw -1.20/

[17] Xu F., Wang M., and Li M., “Sentence -Level Dialects Identification in the Greater China Region,” International Journal on Natural Language Computing , vol. 5, no. 6, pp. 9 -20, 2016. DOI: 10.5121/ijnlc.2016.5602

[18] Zenuni X., Ajdari J., Ismaili F., and Raufi B., “Automatic Hate Speech Detection in Online Contents Using Latent Semantic Analysis,” Press Academia Procedia , vol. 5, no. 1, pp. 368 -371, 2017. https://dergipark.org.tr/en/pub/pap/article/371817