Skip to main navigation Skip to search Skip to main content

Work hard, play hard: Email classification on the avocado and enron corpora

  • Columbia University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

12 Scopus citations

Abstract

In this paper, we present an empirical study of email classification into two main categories “Business” and “Personal”. We train on the Enron email corpus, and test on the Enron and Avocado email corpora. We show that information from the email exchange networks improves the performance of classification. We represent the email exchange networks as social networks with graph structures. For this classification task, we extract social networks features from the graphs in addition to lexical features from email content and we compare the performance of SVM and Extra-Trees classifiers using these features. Combining graph features with lexical features improves the performance on both classifiers. We also provide manually annotated sets of the Avocado and Enron email corpora as a supplementary contribution.

Original languageEnglish
Title of host publicationProceedings of TextGraphs@ACL 2017
Subtitle of host publicationThe 11th Workshop on Graph-Based Methods for Natural Language Processing
EditorsMartin Riedl, Swapna Somasundaran, Goran Glavas, Eduard Hovy
PublisherAssociation for Computational Linguistics
Pages57-65
Number of pages9
ISBN (Electronic)9781945626609
StatePublished - 2020
Event11th Workshop on Graph-Based Methods for Natural Language Processing, TextGraphs 2017, in conjunction with the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017 - Vancouver, Canada
Duration: Aug 3 2017 → …

Publication series

NameProceedings of TextGraphs@ACL 2017: The 11th Workshop on Graph-Based Methods for Natural Language Processing

Conference

Conference11th Workshop on Graph-Based Methods for Natural Language Processing, TextGraphs 2017, in conjunction with the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017
Country/TerritoryCanada
CityVancouver
Period08/3/17 → …

Fingerprint

Dive into the research topics of 'Work hard, play hard: Email classification on the avocado and enron corpora'. Together they form a unique fingerprint.

Cite this