Skip to main navigation Skip to search Skip to main content

The data warehouse of newsgroups

  • AT&T

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

3 Scopus citations

Abstract

Electronic newsgroups are one of the primary means for the dissemination, exchange and sharing of information. We argue that the current newsgroup model is unsatisfactory, especially when posted articles are relevant to multiple newsgroups. We demonstrate that considerable additional flexibility can be achieved by managing newsgroups in a data warehouse, where each article is a tuple of attribute-value pairs, and each newsgroup is a view on the set of all posted articles. Supporting this paradigm for a large set of newsgroups makes it imperative to efficiently support a very large number of views: this is the key difference between newsgroup data warehouses and conventional data warehouses. We identify two complementary problems concerning the design of such a newsgroup data warehouse. An important design decision that the system needs to make is which newsgroup views to eagerly maintain (i.e.,materialize). We demonstrate the intractability of the general newsgroup-selection problem, consider various natural special cases of the problem, and present efficient exact/approximation algorithms and complexity hardness results for them. A second important task concerns the efficient incremental maintenance of the eagerly maintained newsgroups. The newsgroup-maintenance problem for our model of newsgroup definitions is a more general version of the classical point-location problem, and we design an I/O and CPU efficient algorithm for this problem.

Original languageEnglish
Title of host publicationDatabase Theory - ICDT 1999 - 7th International Conference, Proceedings
EditorsPeter Buneman, Catriel Beeri
PublisherSpringer Verlag
Pages471-488
Number of pages18
ISBN (Print)3540654526, 9783540654520
StatePublished - 1998
Event7th International Conference on Database Theory, ICDT 1999 - Jerusalem, Israel
Duration: Jan 10 1999Jan 12 1999

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume1540
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference7th International Conference on Database Theory, ICDT 1999
Country/TerritoryIsrael
CityJerusalem
Period01/10/9901/12/99

Fingerprint

Dive into the research topics of 'The data warehouse of newsgroups'. Together they form a unique fingerprint.

Cite this