Skip to main navigation Skip to search Skip to main content

Keyword extraction for document clustering using submodular optimization

  • Stony Brook University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

1 Scopus citations

Abstract

With the rapid growth of information services, enormous amount of text corpus cannot simply be read and understand. Therefore, text clustering and visualization present a direct way to observe the documents as well as understand the topic by corresponding keywords. However, even a short paragraph contains a variety of words, which makes the keyword or topic extraction difficult to achieve. Therefore, we propose an algorithm to extract keywords efficiently and effectively, which makes use of the latent semantic indexing and submodular optimization. The visual layout allows users to simultaneously visualize (1) the overview of the whole dataset, (2) the detailed information in the specific scope of the collection of documents, and (3) the relationships of documents with their keywords.

Original languageEnglish
Title of host publication2017 New York Scientific Data Summit, NYSDS 2017 - Proceedings
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9781538631614
DOIs
StatePublished - Oct 25 2017
Event2017 New York Scientific Data Summit, NYSDS 2017 - New York, United States
Duration: Aug 6 2017Aug 9 2017

Publication series

Name2017 New York Scientific Data Summit, NYSDS 2017 - Proceedings

Conference

Conference2017 New York Scientific Data Summit, NYSDS 2017
Country/TerritoryUnited States
CityNew York
Period08/6/1708/9/17

Keywords

  • Document Clustering
  • Submodular
  • Visualization

Fingerprint

Dive into the research topics of 'Keyword extraction for document clustering using submodular optimization'. Together they form a unique fingerprint.

Cite this