Methods
Our current projects:
Integration of RNA-Seq and morphology data using a shared autoencoder model
Project lead: Zahra Paylakhi
Major collaborators: Matt Tegtmeyer (Purdue, Dept of Biological Sciences)
Objective: To develop a deep learning framework that integrates gene expression and cellular morphology profiles through a shared latent space, enabling accurate cross-modal predictions and identification of biologically relevant features.
Summary: This project uses RNA-seq gene expression profiles and high-dimensional Cell Painting morphology data to train a shared autoencoder model for cross-modal prediction. The shared latent space enables translation between different types of datasets, allowing for accurate prediction of morphology features from gene expression data. Feature importance analyses and pathway enrichment are used to validate model interpretability and uncover mechanistic links between genetic and morphological variation.
AI- Driven framework for detecting nonlinear genetic interactions in complex traits
Project lead: Saranya Arirangan
Major collaborators: Boran Gao (Purdue, Dept of Biological Sciences; Dept of Statistics), Geyu Zhou (Purdue, Dept of Statistics, Dept of Biological Sciences), Matthew Tegtmeyer (Purdue, Dept of Biological Sciences), Luiz Brito (Purdue, Dept of Animal Sciences), Mitchell Tuinstra (Purdue, Dept of Agronomy)
Objective: To develop a scalable and interpretable AI driven framework for identifying disease-associated loci, including SNP-SNP interactions, in complex traits, and to uncover biologically meaningful multi-locus patterns through analysis of feature importance.
Summary: GWAS have identified many trait-associated loci, but these explain only a small portion of total heritability, often referred to as missing heritability. Epistasis, or non-linear interactions between multiple genetic variants, may contribute to this missing heritability. The true impact of epistatic effects on human traits is still poorly understood, with limited confirmed examples in large-scale studies. We propose a deep learning-based framework for epistasis discovery that leverages Variational Autoencoders (VAE) for unsupervised dimensionality reduction of high-dimensional genotype data, followed by Random Forests (RF) for non-linear feature selection and interaction modeling. The VAE learns compressed, latent representations that capture underlying structure in the genotype space, while the RF model uses these features to identify SNPs both individually and in interaction that influence complex phenotypes. Together, this hybrid approach aims to overcome the limitations of marginal testing by accounting for multi-locus dependencies and enhancing interpretability in causal SNP prioritization.
Protein language models for predicting the functional impact of synonymous mutations
Project lead: Zahra Paylakhi
Major collaborators: Michel Nivard (The University of Bristol)
Objective: To explore the application of protein language models for evaluating the structural and functional consequences of synonymous mutations, with a focus on model-driven feature extraction and prediction.
Summary: This project leverages large-scale protein language models (e.g., ESM-2, AlphaFold2) to study the effects of synonymous mutations, which can influence mRNA stability, translation efficiency, and co-translational protein folding without altering amino acid sequences. By generating structural and sequence embeddings for wild-type and mutant proteins, we compute mutation impact scores and identify potential deleterious variants. The approach aims to advance understanding of silent mutation biology and inform variant interpretation in precision medicine.
Beyond risk factors: Using large language models for narrative analysis of undetermined intent cases in NVDRS
Project lead: Rafael Geurgas
Major collaborators: Alina Arseniev-Koehler (Purdue, Dept of Sociology)
Summary: Undetermined deaths are difficult to classify as suicide, accident, or other causes, which limits the accuracy of public health data. This project uses the National Violent Death Reporting System (NVDRS) to analyze both structured variables and narrative reports from death investigations. By applying Large Language Models (LLMs) to these narratives, we aim to uncover hidden patterns that shed light on the social and contextual factors surrounding these deaths. The findings will help communities better understand how such cases are classified, improve suicide surveillance, and reveal how social and institutional processes influence the way deaths are recorded.
A Unified Framework for Robust Fine-Mapping under LD Reference Mismatch
Project lead: Jiachen Liu
Major collaborators: Boran Gao (Purdue, Dept of Statistics and Dept of Biological Sciences), Geyu Zhou (Purdue, Dept of Statistics and Dept of Biological Sciences)
Objective: To develop a unified statistical framework for genetic fine-mapping that jointly addresses multiple sources of LD reference mismatch, improving the robustness and power of causal variant identification in post-GWAS analyses.
Summary: Statistical fine-mapping is a critical step in translating GWAS findings into biological insights, but its reliability is often compromised by multiple sources of mismatch between summary statistics and external LD reference panels. This project develops a unified framework that jointly models these mismatch sources, rather than addressing them in isolation as existing methods do. With applications to complex traits across diverse ancestries, we aim to improve power, convergence, and robustness of fine-mapping under realistic data conditions, ultimately supporting more accurate therapeutic target discovery.

