AI agents for hypothesis generation from single-cell and spatial transcriptomics

Large language model agents that generate biomedical hypotheses from single-cell, spatial omics and drug–target data.

Workflow: knowledge augmentation from public databases (UniProt, Human Protein Atlas, ProteomicsDB, Open Targets, ClinicalTrials.gov, GTEx) feeding reasoning-based target prioritization and nomination with several large language models.
Baker DJ, et al. AI-driven discovery of GPNMB CAR T cells as a multi-cancer therapy. Cell (2026).

Background

Drug development remains slow, costly and high-risk, which motivates strategies that accelerate translation. Our work follows a unified framework centered on pathways as the common unit between disease biology and therapeutics.

Questions we are working on

  • How do a plain LLM, an LLM with retrieval-augmented generation (RAG) and an LLM grounded in a knowledge graph compare as agents for biomedical hypothesis generation?
  • Given single-cell, spatial omics and publicly available drug–target datasets, how do these agents compare with existing methods such as Medea and Biomni?

Key literature

  1. Baker DJ, et al. AI-driven discovery of GPNMB CAR T cells as a multi-cancer therapy. Cell (2026). [Hayat Lab]
  2. Sui P, et al. Medea: An omics AI agent for therapeutic discovery. bioRxiv (2026).
  3. Zhou L, et al. Autonomous Agents for Scientific Discovery: Orchestrating Scientists, Language, Code, and Physics. arXiv:2510.09901 (2025).
  4. Zitnik M. AI-enabled drug discovery reaches clinical milestone. Nat Med (2025).
  5. Gottweis J, et al. Accelerating scientific discovery with Co-Scientist. Nature (2026).
  6. Petrić Howe N, Thompson B. AI ‘scientists’ promise to accelerate research — how do they work? Nature (2026).
  7. Swanson K, et al. The Virtual Lab of AI agents designs new SARS-CoV-2 nanobodies. Nature (2025).