Research
Overview
My current research focuses on unsupervised learning via Discrete Matrix Factorization, with a strong motivation from Multiple Myeloma (MM) genomics data.
Multiple Myeloma is a malignancy of post‑germinal centre B cells (plasma cells) and is driven by chromosomal and genetic alterations such as translocations, copy‑number abnormalities (CNAs), and point mutations. In particular, CNAs—gains or losses of chromosomes or chromosome arms, deletions of 1p, and gains of 1q—play an important role in prognosis, treatment choice, and disease relapse.
To analyse these discrete genomic profiles, I develop Bayesian methods for factorizing discrete matrices that:
- uncover latent structure in high‑dimensional binary and ternary data,
- provide compact and interpretable representations,
- remain robust in the presence of noise, and
- are tailored to copy‑number alteration data recorded as −1 (deletion), 0 (normal), and +1 (amplification).
Broader Applications
Although these methods are motivated by challenges in biomedical research, particularly cancer genomics, they are broadly applicable to any domain involving high-dimensional discrete data. Potential applications include recommender systems, text mining, social network analysis, cybersecurity, epidemiology, survey research, and other areas where uncovering interpretable latent structure and quantifying uncertainty are important.
Specific Projects
Current projects include:
