SusztakLab Biobank

Research

From mapping to mechanism to medicine

The Susztak Laboratory studies why kidneys fail and how to stop it. We start from the human kidney itself: more than 5,000 kidney samples with clinical information, profiled at every molecular layer, and read together with the genetics of millions of people, single-cell and spatial maps, longitudinal patient cohorts and artificial intelligence. Our aim is to turn that knowledge into precise diagnostics and targeted therapies for the 850 million people who live with kidney disease.

1. Artificial intelligence across scale: NephroBase · 2. Mapping the genetic architecture of kidney disease · 3. Defining cell-type-specific disease mechanisms · 4. Proximal tubule injury, immune crosstalk and fibrosis · 5. From mechanism to translational targets · Platforms

01The programme that ties the others together

Artificial intelligence across scale: NephroBase

Foundation models that read cells, slides, genomes and patients together.

No single dataset explains kidney disease. We have profiled tens of millions of cells, thousands of tissue sections and whole-slide images, and the molecular and clinical records of thousands of patients, and the real insight lies in reading them together. That is the purpose of our artificial intelligence programme, and it rests on four models that share one representation of the kidney.

NephroBase is the hub: a kidney-specific foundation model that embeds single-cell, spatial, molecular, imaging and clinical data in one latent space. Its single-cell component, a virtual cell model trained on tens of millions of kidney cells, recognises cell and tissue states it has never been shown, transfers knowledge across species and technologies, infers gene regulatory networks and predicts how a cell responds to a perturbation. NephroLens is our pathology foundation model, trained on a very large collection of kidney biopsy whole-slide images; it quantifies glomerular, tubular, interstitial and vascular lesions automatically and produces a structured report that a pathologist can review. For the genome, we adapt sequence-to-function models to kidney cell types so that the effect of a regulatory variant can be predicted in the cell where it acts, and we test those predictions against our own eQTL, chromatin and protein QTL maps.

The fourth element is the patient. By linking the molecular layers to longitudinal cohorts such as TRIDENT and to health-system records, we are building what we call a digital kidney twin: a model that follows an individual kidney through time, connects a biopsy, a genome and a clinical course to the mechanisms we understand, and tells us which of them we can treat and when. It is early, ambitious work, and it is the direction in which the whole laboratory is moving.

  • NephroBase: a kidney foundation model spanning more than 70 cell types, five data modalities and three species
  • NephroLens: automated, pathologist-reviewable quantification of kidney biopsy slides
  • Kidney-adapted sequence models that predict regulatory variant effects per cell type
  • Towards a digital kidney twin that links molecules to clinical trajectories

Read Klötzer et al., Nature Genetics 2025 · Liu et al., Science 2025

02

Mapping the genetic architecture of kidney disease

Which variants change kidney function, through which gene, in which cell?

Kidney disease is strongly inherited, yet for decades we knew almost none of the responsible genes. Genome-wide association studies changed that: by comparing millions of common variants between people with better and worse kidney function, our 2.2-million-person study identified more than a thousand regions of the genome where a variant changes kidney function. The difficulty is that almost none of these variants alter a protein. They sit in regulatory DNA, and the association alone cannot say which gene they act on, or in which of the kidney's thirty-odd cell types.

We answer that question in human kidney tissue. In hundreds of kidneys with both genotype and molecular data we ask, variant by variant, whether the genotype changes the expression of a gene, the methylation of a CpG site, the accessibility of chromatin or the abundance of a protein. When a GWAS signal and such a quantitative trait locus point to the same variant, the locus gains a gene, a cell type and a mechanism. This integration produced the Kidney Disease Genetic Scorecard, a ranked list of genetically supported target genes for every kidney function locus, and it is the fastest route we know from a risk variant to a drug target.

  • First maps of expression, methylation, chromatin and protein QTLs in human kidney tissue
  • 1,026 kidney function loci in 2.2 million people and a Genetic Scorecard of target genes
  • Prioritised causal genes validated in experimental models, including MANBA, DAB2, DACH1 and APOL1

Read Liu et al., Science 2025 · Liu et al., Nature Genetics 2022 · Sheng et al., Nature Genetics 2021 · Qiu et al., Nature Medicine 2018 · Hirohama et al., Nature Medicine 2025

03

Defining cell-type-specific disease mechanisms

Building the atlas of the human kidney in health and disease, one cell at a time.

The kidney is built from more than thirty cell types, and disease does not treat them equally. With single-cell and single-nucleus sequencing we have profiled millions of cells from healthy and diseased human kidneys, and from mouse and rat models, to see which cell types are lost, which change state and which appear only in disease. Two findings organise much of our current work: an injured, pro-inflammatory state of the proximal tubule that fails to repair, and a fibrotic microenvironment in which immune cells and activated fibroblasts reinforce each other and drive progression.

Spatial transcriptomics adds the missing coordinate. By measuring gene expression in intact tissue sections we can see which cells are neighbours and how those neighbourhoods reorganise in disease. In diabetic kidney disease this revealed a previously unrecognised subgroup of patients defined by B cell-rich immune niches, and in the developing kidney it showed how the microenvironment steers progenitor cells to their fate.

  • Single-cell atlases of adult, diseased and developing human kidney, integrated with mouse and rat
  • The injured proximal tubule state and the fibrotic microenvironment as drivers of progression
  • A B cell-rich subgroup of diabetic kidney disease revealed by spatial transcriptomics

Read Abedini, Levinsohn et al., Nature Genetics 2024 · Dumoulin et al., Nature 2026 · Levinsohn, Grindel et al., Nature Genetics 2026 · Klötzer et al., Nature Genetics 2025 · Park et al., Science 2018

04

Proximal tubule injury, immune crosstalk and fibrosis

How a damaged tubule cell turns into a driver of inflammation and scarring.

Most of the kidney's work is done by the proximal tubule, and most of what goes wrong in chronic kidney disease begins there. Our studies showed that the injured tubule cell is not a passive victim: it loses its metabolic identity, activates nucleic acid sensing and inflammatory signalling, and releases cytokines and chemokines that recruit macrophages, lymphocytes and basophils and activate the surrounding stromal cells. The result is collagen deposition and fibrosis, the final common pathway of kidney failure.

We dissect this cascade with cell-type-specific genetic models, lineage tracing and single-cell profiling after injury, and we ask which steps are reversible. Metabolic reprogramming, mitochondrial and lysosomal dysfunction, cytosolic DNA and RNA sensing, and the dialogue between tubule, immune and stromal cells are the mechanisms we study most closely, because each of them is a potential point of therapeutic intervention.

  • Metabolic and mitochondrial failure of the proximal tubule as an early event in fibrosis
  • Cytosolic nucleotide sensors (cGAS-STING, RIG-I) linking cell damage to inflammation
  • Tubule-immune and tubule-stromal signalling that recruits basophils, macrophages and fibroblasts

Read Doke et al., Nature Immunology 2022 · Balzer et al., Nature Communications 2022 · Dhillon, Park et al., Cell Metabolism 2021 · Miao, Balzer et al., Nature Communications 2021

05

From mechanism to translational targets

Testing genetically supported genes in models and defining when to intervene.

A gene that human genetics supports is a far better bet in the clinic than one chosen from a mouse experiment alone; drugs with genetic support are roughly twice as likely to succeed. We therefore take the genes that come out of our genetic and single-cell work and establish their causal role in experimental models, with cell-type-specific deletion or activation, before we think about therapy. MANBA, DAB2, DACH1 and APOL1 are examples of kidney disease genes that moved this way from a statistical signal to a mechanism.

Our rat and mouse studies of established drugs, renin-angiotensin-aldosterone blockade, mineralocorticoid receptor antagonists and soluble guanylate cyclase activators, show at single-cell resolution how a treatment protects the kidney and in which cells it acts. The same approach tells us when a mechanism is active, and therefore when an intervention is likely to work, which is as important as the target itself.

  • Causal validation of MANBA, DAB2, DACH1 and APOL1 in animal models
  • Single-cell maps of drug action for RAAS inhibition, MRAs and sGC activators
  • Disease-stage-specific windows for intervention

Read Balzer et al., J Am Soc Nephrol 2026 · Abedini et al., J Clin Invest 2024 · Balzer et al., Cell Reports Medicine 2023 · Sheng et al., Nature Genetics 2021

Research platforms

The projects above rest on four shared platforms: a human kidney biobank, two patient cohorts and the computational engine that reads them together.

The molecular foundation

Human Kidney Biobank

Progress in nephrology has long been limited by the scarcity of human kidney tissue. Over two decades we have assembled one of the largest research collections of human kidney samples in the world, each with curated clinical information, preserved fresh-frozen, in FFPE blocks, in RNAlater and as primary cell cultures, so that the same kidney can be studied by genomics, epigenomics, proteomics, metabolomics and digital pathology. Every atlas on this site grew out of this collection.

  • 5,000+ human kidney samples
  • Genotype, methylation, chromatin, expression, protein
  • Bulk, single-nucleus and spatial profiling
Browse the data resources →

Longitudinal precision for diabetic kidney disease

TRIDENT

The Transformative Research in Diabetic Nephropathy consortium enrols people with diabetes at the time of a clinical kidney biopsy and follows them for years, pairing the biopsy tissue with serial urine and plasma samples and clinical outcomes. It is the cohort that lets us ask which molecular features predict who will progress and who will respond to treatment, and it is the basis of our biomarker and patient-stratification work.

  • Multi-centre prospective cohort
  • Kidney tissue, urine and plasma over time
  • Progression and treatment-response outcomes
TRIDENT website →

Health-system-scale validation

Penn Medicine Kidney Biobank

Through the Penn Medicine Biobank, kidney tissue and biofluids are linked to the electronic health records of a large and diverse health system. This is where molecular findings meet real-world outcomes: a risk score or a biomarker discovered in the laboratory can be tested in tens of thousands of patients, across ancestries and across the full course of disease.

  • EHR-linked tissue and biofluids
  • Outcome analyses at health-system scale
  • Validation across diverse populations
Penn Medicine Biobank →

The analytical engine

NephroBase

NephroBase is the computational and AI platform that holds the laboratory's data together: a kidney-specific foundation model trained on tens of millions of cells and the clinical, imaging and molecular data around them. It provides unified representations across more than 70 kidney cell types, predictive models of disease progression, automated quantification of kidney pathology and a ranking of therapeutic targets.

  • Kidney-specific foundation model
  • 70+ cell types in one representation
  • Progression, pathology and target prediction
Read about the AI programme →