Skip to Main Content

Qiao Liu, PhD

Assistant Professor of Biostatistics
DownloadHi-Res Photo

About

Titles

Assistant Professor of Biostatistics

Biography

I am an Assistant Professor in the Department of Biostatistics at Yale University. I currently lead the Liu Lab (https://liuq-lab.com/) at Yale Biostatistics.

Our research bridges modern AI and statistical science to build computational tools that are not only powerful and scalable but also trustworthy and interpretable. By combining the flexibility of AI with the rigor of statistics, we aim to drive transformative advances in biomedical research, enabling discoveries that were previously out of reach.

Our current research interests aim at understanding how genetic variation shapes gene regulation, cellular states, and ultimately human disease. A major challenge is that these processes are highly context dependent, varying across cell types, tissues, individuals, space, and time, and many of the molecular states connecting genotype to phenotype are only partially observed. We develop AI-powered statistical and computational methods for Causal Regulatory Genomics to uncover these hidden regulatory mechanisms.

Our work spans several complementary directions. Context-aware genomic foundation models, including RegFM (Gao et al., bioRxiv 2026) and EpiGePT (Gao et al., Genome Biol. 2024), learn transferable regulatory knowledge from large-scale genomic data. Generative AI for single-cell and multiomics, including scDiffusion-X (Luo et al., Nat. Commun. 2026), scMTG (Cui et al., bioRxiv 2026), scDEC (Liu et al., Nat. Mach. Intell. 2021), and MultiFlow (Wang et al., bioRxiv 2026), reconstructs and predicts cellular states across modalities, biological contexts, time, and perturbations. Trustworthy causal AI, including CausalBGM (Liu et al., J. Am. Stat. Assoc. 2026), BGM (Liu et al., arXiv 2026), CausalEGM (Liu et al., PNAS 2024), and Roundtrip (Liu et al., PNAS 2021), aims to move beyond association/prediction toward causal effect estimation and mechanistic interpretation.

Together, these directions form an integrated research paradigm for connecting genetic variation → regulatory effects → gene programs → cellular responses → phenotype, with the broader goal of turning increasingly large and complex genomic datasets into interpretable knowledge about human biology and disease.

Last Updated on August 27, 2026.

Appointments

Education & Training

Postdoctoral Scholar
Stanford University (2025)
Visiting PhD student
Stanford University (2021)
PhD
Tsinghua University (2021)

Research

Overview

  • Multiomics Integration: We develop statistical and AI-driven methods to integrate diverse molecular data modalities, including genetics, transcriptomics, epigenomics, radiomics at different scale and resolution. Our goal is to build scalable and interpretable computational frameworks that can jointly model heterogeneous multiomics data, recover missing modalities, quantify uncertainty, and reveal regulatory programs across molecular layers.
  • Causal Inference: We develop causal inference methodologies for high-dimensional biomedical data. By combining modern AI techniques and rigorious statistical inference (e.g., Bayesian inference), we aim to 1) estimate causal effects, especially under high-dimensional covaraites or even hidden confounding variable, 2) identify disease-relevant drivers (gene, regualtory elements, regulatory programs) and support interpretable biological discovery.
  • Single Cell Genomics: We develop computational tools for single cell analysis that capture cell hetergeneity, cell state transition, cell-cell/environment communications/response etc. Current reserach interests lie on identifying/discovering biological mechanisms to analyze time‑course data, lineage tracing, CRISPR and small‑molecule perturbation screens. Our models provide insights into gene regulation mechanisms by modeling cell development, transition, response, aging, etc.
  • Pharmacogenomics: We study how genetic and molecular variation influences drug response, treatment sensitivity, and disease progression. By integrating genomics, transcriptomics, and clinical or experimental drug-response data, we aim to develop predictive and causal models that can help identify therapeutic targets and support precision medicine.
  • Genomic Foundation Models: We build context-aware foundation models for regulatory genomics by leveraging large-scale sequence, epigenomic, transcriptomic, and perturbation datasets. Our goal is to develop context-aware models that can learn generalizable representations of gene regulation, predict molecular phenotypes across cell types and conditions, and enable downstream tasks such as genetic variant effect quantification, regulatory element discovery, and phenotype prediction.

Medical Research Interests

Artificial Intelligence; Causality; Computational Biology; Data Science; Deep Learning; Epigenomics; Gene Expression Regulation; Generative Artificial Intelligence; Genomics; Machine Learning; Single-Cell Analysis

Public Health Interests

Aging; Bioinformatics; Genetics, Genomics, Epigenetics; Modeling; Bayesian Statistics

Research at a Glance

Publications Timeline

A big-picture view of Qiao Liu's research output by year.

Publications

Featured Publications

2026

Academic Achievements & Community Involvement

Honors

  • honor

    Pathway to Independence Awards (K99/R00)

Get In Touch

Contacts

Academic Office Number

Locations

  • 300 George Street

    Academic Office

    Ste 501

    New Haven, CT 06511