TY - GEN
T1 - ClusterSculptor
T2 - VAST IEEE Symposium on Visual Analytics Science and Technology 2007
AU - Nam, Eun Ju
AU - Han, Yiping
AU - Mueller, Klaus
AU - Zelenyuk, Alla
AU - Imre, Dan
PY - 2007
Y1 - 2007
N2 - Cluster analysis (CA) is a powerful strategy for the exploration of high-dimensional data in the absence of a-priori hypotheses or data classification models, and the results of CA can then be used to form such models. But even though formal models and classification rules may not exist in these data exploration scenarios, domain scientists and experts generally have a vast amount of non-compiled knowledge and intuition that they can bring to bear in this effort. In CA, there are various popular mechanisms to generate the clusters, however, the results from their nonsupervised deployment rarely fully agree with this expert knowledge and intuition. To this end, our paper describes a comprehensive and intuitive framework to aid scientists in the derivation of classification hierarchies in CA, using k-means as the overall clustering engine, but allowing them to tune its parameters interactively based on a non-distorted compact visual presentation of the inherent characteristics of the data in highdimensional space. These include cluster geometry, composition, spatial relations to neighbors, and others. In essence, we provide all the tools necessary for a high-dimensional activity we call cluster sculpting, and the evolving hierarchy can then be viewed in a space-efficient radial dendrogram. We demonstrate our system in the context of the mining and classification of a large collection of millions of data items of aerosol mass spectra, but our framework readily applies to any high-dimensional CA scenario.
AB - Cluster analysis (CA) is a powerful strategy for the exploration of high-dimensional data in the absence of a-priori hypotheses or data classification models, and the results of CA can then be used to form such models. But even though formal models and classification rules may not exist in these data exploration scenarios, domain scientists and experts generally have a vast amount of non-compiled knowledge and intuition that they can bring to bear in this effort. In CA, there are various popular mechanisms to generate the clusters, however, the results from their nonsupervised deployment rarely fully agree with this expert knowledge and intuition. To this end, our paper describes a comprehensive and intuitive framework to aid scientists in the derivation of classification hierarchies in CA, using k-means as the overall clustering engine, but allowing them to tune its parameters interactively based on a non-distorted compact visual presentation of the inherent characteristics of the data in highdimensional space. These include cluster geometry, composition, spatial relations to neighbors, and others. In essence, we provide all the tools necessary for a high-dimensional activity we call cluster sculpting, and the evolving hierarchy can then be viewed in a space-efficient radial dendrogram. We demonstrate our system in the context of the mining and classification of a large collection of millions of data items of aerosol mass spectra, but our framework readily applies to any high-dimensional CA scenario.
KW - High-dimensional data
KW - Space and environmental sciences
KW - Visual analytics
KW - Visual data mining
KW - Visualization in earth
UR - https://www.scopus.com/pages/publications/41549149938
U2 - 10.1109/VAST.2007.4388999
DO - 10.1109/VAST.2007.4388999
M3 - Conference contribution
AN - SCOPUS:41549149938
SN - 9781424416592
T3 - VAST IEEE Symposium on Visual Analytics Science and Technology 2007, Proceedings
SP - 75
EP - 82
BT - VAST IEEE Symposium on Visual Analytics Science and Technology 2007, Proceedings
Y2 - 30 October 2007 through 1 November 2007
ER -