| dc.description.abstract |
This dissertation investigates the usage of Exploratory Data Analysis (EDA) in the biomedical domain.
EDA was developed as a complementary approach to Confirmatory Data Analysis (CDA).
While CDA mainly uses statistical analysis to confirm scientific hypotheses, EDA uses techniques like visualization to enable the exploration of data to guide
hypothesis generation.
This is illustrated using three approaches developed and published for this dissertation.
The first approach, squarely on the medical side of things, ICUVA, uses EDA to explore the status of the homeostatic system of a patient in an Intensive Care Unit (ICU) scenario.
This status is described by multiple different types of ICU measurements.
Usually, the course of these measurements across time is analyzed individually and often simple visualizations like line charts are used to explore the data.
This is not feasible for more than a few measurements, without losing the completeness of the situation.
Thus, ICUA uses time-curves, which embed the multivariate data into a two-dimensional space to identify points in time when the system as changes as a whole.
To understand phenotypes such as disease phenotypes observed in clinical practice, the organism also needs to be analyzed on a molecular level.
Various types of biomolecules are relevant for characterizing a phenotype.
These molecules stem from different levels of gene-expression.
Among those are transcripts, the messenger molecules copied from genes, proteins, the molecules that are synthesized from this information, and metabolites, small-molecules, like hormones, often produced with the help of enzymes.
These molecules can be mapped to biological pathways that describe how an organism works on a molecular level.
In the second approach, VisMOP, such data of biomolecules mapped to biological-pathways is analyzed simultaneously by performing multi-omics clustering.
These clustered pathways are visualized as a clustered, hierarchical graph where glyphs represent pathways and their multi-omics measurements.
Ultimately, this helps to find similarly regulated pathways which allows deriving hypotheses on how phenotypes are
established.
The associations of proteins, as they are present in biological pathways can also be modeled as an association-network, commonly called Protein-Protein Interaction (PPI)-network.
PPI-networks can for example be used to explore effects of defects in altered proteins on other proteins in the network, which can be used to generate hypotheses on disease mechanisms or other phenotypes.
To facilitate such exploration, we developed ProtEGOnist, a domain agnostic approach that utilizes an egocentric network layout to explore small world networks focusing on nodes-of-interest.
We demonstrated this on PPI-network data.
The discussion shows how EDA applications can be used in a complex biomedical scenario, exemplified using the three presented approaches.
Additionally, challenges when implementing EDA applications in an academic settings are described. Finally, this dissertation addresses how we coped with these challenges in the presented approaches. |
en |