Skip to main navigation Skip to search Skip to main content

Classification of high-dimensional data with ensemble of logistic regression models

  • University of California at San Francisco
  • California State University Long Beach
  • United States Food and Drug Administration

Research output: Contribution to journalArticlepeer-review

20 Scopus citations

Abstract

A classification method is developed based on ensembles of logistic regression models, with each model fitted from a different set of predictors determined by a random partition of the feature space. The proposed method enables class prediction by an ensemble of logistic regression models for a high-dimensional data set, which is impossible by a single logistic regression model due to the restriction that the sample size needs to be larger than the number of predictors. The proposed classification method is applied to gene expression data on pediatric acute myeloid leukemia (AML) patients to predict each patient's risk for treatment failure or relapse at the time of diagnosis. Hence, specific prognostic biomarkers can be used to predict outcomes in pediatric AML and formulate individual risk-adjusted treatment. Our study shows that the proposed method is comparable to other widely used models in generalized accuracy and is significantly improved in balance between sensitivity and specificity. The proposed ensemble algorithm enables the standard classification model to be used for classification of high-dimensional data.

Original languageEnglish
Pages (from-to)160-171
Number of pages12
JournalJournal of Biopharmaceutical Statistics
Volume20
Issue number1
DOIs
StatePublished - Jan 2010

Keywords

  • Aggregation
  • Class prediction
  • Cross-validation
  • Decision threshold
  • Majority voting
  • Random partition

Fingerprint

Dive into the research topics of 'Classification of high-dimensional data with ensemble of logistic regression models'. Together they form a unique fingerprint.

Cite this