The Assessment of Feature Selection Methods on Agglutinative Language for Spam Email Detection: A Special Case for Turkish


ERGİN S., IŞIK Ş.

IEEE International Symposium on Innovations in Intelligent Systems and Applications (INISTA), Alberobello, İtalya, 23 - 25 Haziran 2014, ss.122-125 identifier identifier

  • Yayın Türü: Bildiri / Tam Metin Bildiri
  • Cilt numarası:
  • Doi Numarası: 10.1109/inista.2014.6873607
  • Basıldığı Şehir: Alberobello
  • Basıldığı Ülke: İtalya
  • Sayfa Sayıları: ss.122-125
  • Eskişehir Osmangazi Üniversitesi Adresli: Evet

Özet

In this study, the assessment of three different feature selection methods including Information Gain (IG), Gini Index (GI), and CHI square (CHI2) is made by utilizing two popular pattern classifiers, namely Artificial Neural Network (ANN) and Decision Tree (DT), on the classification of Turkish e-mails. The feature vectors are constructed by the bag-of-words feature extraction method. This paper is focused on the Turkish language since it is one of the widely used agglutinative languages all around the world. The results obviously reveal that CHI2 and GI feature selection methods are more efficacious than IG method for Turkish language.