Subpopulation Discovery in Epidemiological Data with Subspace Clustering

A prerequisite of personalized medicine is the identification of groups of people who share specific risk factors towards an outcome. We investigate the potential of subspace clustering for finding such groups in epidemiological data. We propose a workflow that encompasses clusterability assessment before cluster discovery and quality assessment after learning the clusters. Epidemiological usually do not have a ground truth for the verification of clusters found in subspaces. Hence, we introduce quality assessment through juxtaposition of the learned models to “models-of-randomness”, i.e. models that do not reflect a true cluster structure. On the basis of this workflow, we select subspace clustering methods, compare and discuss their performance. We use a dataset with hepatic steatosis as outcome, but our findings apply on arbitrary epidemiological cohort data that have tenths of variables and exhibit class skew.

eISSN:: 2300-3405
Language:: English

Publication timeframe:: 4 times per year
Journal Subjects:: Computer Sciences, Artificial Intelligence, Software Development

Journal RSS Feed

Subpopulation Discovery in Epidemiological Data with Subspace Clustering

Published Online: Dec 20, 2014

Page range: 271 - 300

Received: Aug 01, 2014

DOI: https://doi.org/10.2478/fcds-2014-0015

© by Uli Niemann

This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 3.0 License.