Abstract
This study compares BERTurk, mBERT, XLM-RoBERTa and ELECTRA-based language models for topic classification of Turkish news texts. A new dataset of 48,000 news articles in 12 categories was compiled and made publicly available. The Turkish pre-trained BERTurk model achieved the highest performance with a macro F1 score of 93.6%. The effect of agglutinative morphology on subword segmentation was examined; increasing vocabulary size improved performance, especially for short texts.
Türkçe Haber Metinlerinin Sınıflandırılmasında Dönüştürücü Tabanlı Dil Modellerinin Başarımı
Bu çalışmada Türkçe haber metinlerinin konu sınıflandırması için BERTurk, mBERT, XLM-RoBERTa ve ELECTRA tabanlı dil modelleri karşılaştırılmıştır. 12 kategoride 48.000 haber metninden oluşan yeni bir veri kümesi derlenmiş ve kamuya açık hale getirilmiştir. Türkçeye özgü ön eğitimli BERTurk modeli %93,6 makro F1 skoru ile en yüksek başarıyı göstermiştir. Eklemeli dil yapısının alt kelime bölütleme üzerindeki etkisi incelenmiş; sözlük boyutunun artırılmasının özellikle kısa metinlerde başarıyı iyileştirdiği görülmüştür.
Full Text
Declarations
- Ethics Approval
- Bu çalışma etik kurul onayı gerektirmemektedir.
- Conflict of Interest
- Yazarlar herhangi bir çıkar çatışması olmadığını beyan eder.
- Funding
- Bu çalışma Selçuk Üniversitesi BAP Koordinatörlüğü tarafından desteklenmiştir (Proje No: 24003).
References 5
- Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT 2019 (pp. 4171–4186). https://doi.org/10.18653/v1/N19-1423
- Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30.
- Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural Computation, 9(8), 1735–1780. https://doi.org/10.1162/neco.1997.9.8.1735
- Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep learning. MIT Press.
- Kingma, D. P., & Ba, J. (2015). Adam: A method for stochastic optimization. In International Conference on Learning Representations.
How to Cite
License
© 2024 Selin Öztürk, Cem Arslan. This article is distributed under the terms of the CC BY 4.0 license, which permits unrestricted use, distribution and reproduction in any medium, provided the original work is properly cited. License text