Repository | Book | Chapter

A comparative study of language modeling to instance-based methods, and feature combinations for authorship attribution

Olga Fourkioti, Symeon Symeonidis, Avi Arampatzis

pp. 274-286

We present a comparative study of language modeling to traditional instance-based methods for authorship attribution, using several different basic units as features, such as characters, words, and other simple lexical measurements, as well as we propose the use of part-of-speech (POS) tags as features for language modeling. In contrast to many other studies which focus on small sets of documents written by major writers regarding several topics, we consider a relatively large corpus with documents edited by non-professional writers regarding the same topic. We find that language models based on either characters or POS tags are the most effective, while the latter provide additional efficiency benefits and robustness against data sparsity. Moreover, we experiment with linearly combining several language models, as well as employing unions of several different feature types in instance-based methods. We find that both such combinations constitute viable strategies which generally improve effectiveness. By linearly combining three language models, based respectively on character, word, and POS trigrams, we achieve the best generalization accuracy of 96%.

Publication details

DOI: 10.1007/978-3-319-67008-9_22

Full citation:

Fourkioti, O. , Symeonidis, S. , Arampatzis, A. (2017)., A comparative study of language modeling to instance-based methods, and feature combinations for authorship attribution, in J. Kamps, G. Tsakonas, Y. Manolopoulos, L. Iliadis & I. Karydis (eds.), Research and advanced technology for digital libraries, Dordrecht, Springer, pp. 274-286.

This document is unfortunately not available for download at the moment.