The mission stated at the start of this analysis was to extract the most kidney biology from 128 heterogeneous samples and to say clearly what still cannot be extracted. This final chapter is the accounting: what the nine preceding chapters established, and what remains beyond all five platforms.

What the analysis established

The single-cell platforms resolved the kidney’s composition directly, at 90% assignment, without supervision: the loop of Henle dominant in the adult, the stroma dominant in the fetus, the podocyte and tubule programs in their expected places (Chapter 3). The whole-transcriptome platforms added the genome-scale spatial program, including renin, podocyte, loop-of-Henle, and distal-tubule genes and the metabolic zonation that makes the outer medulla the vulnerable zone (Chapter 9). The region platform recovered compartment identity with no coordinates at all (Chapter 5). And the cleanest experiment delivered the cleanest result: within four hours of ischemia, the vulnerable S3 segment loses identity and converts to an injured state, the textbook’s most vulnerable cell, at single-cell resolution, on two independent platforms (Chapter 7).

These are not claims the analysis was given. They are the biology the data carry, recovered by a framework that treats every platform as a lens and every method as an instrument to be interrogated.

What the integration adds

The final analytical step was to bring the platforms into a common space: scVI per species and tier across samples, and a module-based resolution ladder that places every ROI, spot, bin, and cell on shared biological axes.

The limiting fact in this table is as important as the positive results: the integration is bounded by shared genes. The mouse Xenium integration runs on 55 shared panel genes, because across 19 Xenium sections that is all the panels share. This is the cross-platform ceiling in its purest form, and it is reported, not hidden.

Mechanically, the integration is scVI, a variational autoencoder from the scvi-tools stack, fit separately for each species and platform tier with the sample as the batch variable, on the genes those samples share. Where a tier held a single sample, or the model failed to fit, I fell back to PCA on the same shared genes rather than reporting nothing. The 55 shared genes on the mouse Xenium tier are the bound in its purest form: nineteen sections and roughly 1.7 million cells, with a panel intersection smaller than a single panel’s size.

Cell-cell communication: a null result

I also asked whether the single-cell samples could support ligand-receptor communication inference. The answer was uniformly negative, and I report it as such.

Across all 22 single-cell samples, the permutation test for ligand-receptor enrichment returned zero significant interactions. The neighborhood analysis still identified eight compositionally distinct niches per sample, with the dominant cell type of each recorded in the full table. Intercellular signaling was not detectable at this panel depth. A communication analysis that reported significant axes here would be reporting noise; the null is the result, and that is what I report.

Mechanically, the inference used squidpy’s ligand-receptor routine against the cellphonedb consensus interaction database, with a permutation test to establish what chance alone would produce. I report the null as a result because it is one: at panel depths of a few hundred genes, the ligand-receptor pairs the panel actually measures are too few to detect a consistent signal, which is the whole reason the null is informative rather than disappointing.

What remains beyond these platforms

Even with all five platforms, 128 samples, and the full analytical stack, there are questions the data cannot answer, and a professional analysis names them:

  1. The panel gap. A 300-gene panel and a 19,000-gene transcriptome share a few hundred genes. Gene-level discovery is confined to the platforms that measured the genes; no analysis can manufacture information that was never measured.

  2. No matched multi-platform sections. Each platform profiled its own tissue. Cross-platform comparisons therefore conflate platform and biology; only the matched-reference sub-analyses (e.g., the IRI Xenium-to-Visium concordance of Chapter 4) control for this, and they are the exceptions, not the rule.

  3. Deconvolution is reference- and method-bound. Spot-resolution cell types are only as reliable as the reference that matches the biology and as the method used to unmix; the concordance range, measured per method in Chapter 4 (Tangram 0.06 to 0.88, RCTD −0.18 to 0.64), is the cost of reference and method mismatch. The choice between deconvolution strategies, and why no single one can be trusted blindly, is discussed in Chapter 4.

  4. GeoMx has no coordinates. The region lens reads compartments, not neighborhoods, not domains, not gradients within a region.

  5. Image resolution caps segmentation. Standard Visium images cannot resolve individual cells; only Visium HD’s finer images can, and even there the cell matrices are derived, not directly imaged.

  6. No orthogonal experimental validation. This is a re-analysis of public data; every conclusion here is a statistical readout, not a perturbed experiment. Nothing in this analysis is a therapeutic claim.

  7. External references are integrated. Annotation and deconvolution run against whole-transcriptome human and mouse kidney references retrieved from the CELLxGENE Census, with original-study labels harmonized to the nephron-module vocabulary. The planned cross-reference validation is now computed: a label-transfer adjusted Rand index of 0.26 on average across the single-cell samples (Chapter 3) and a Tangram deconvolution concordance with a median ρ of 0.71 across the Visium sections (Chapter 4). Two reference boundaries remain: the Census slices carry no S3-specific or injured-proximal-tubule labels, so reference-bound deconvolution cannot resolve those sub-states, and the reference composition is a balanced cross-study slice rather than a per-section ground truth.

Conclusion

The analysis supports a substantial body of kidney biology that can be recovered across platforms and evaluated through internal consistency. At the same time, the available data do not justify conclusions that extend beyond the measured molecular panels, reference datasets, or experimental designs. A central contribution of the analytical framework is therefore to make these evidentiary boundaries explicit. This is reflected in the reporting of concordance measures alongside deconvolution results, ARI values for domain comparisons, explicit panel limitations for targeted-panel gene lists, and the documentation of ribosomal artifacts rather than their interpretation as biological signals.

That boundary is an integral part of the analysis, not a limitation to be set aside. A rigorous examination of 128 heterogeneous samples requires both the systematic extraction of patterns supported by the data and the explicit identification of conclusions that cannot be established from the available evidence. Together, these findings define the requirements for subsequent studies, including matched multi-platform sections, expanded molecular panels, and appropriately matched reference datasets. The resulting kidney spatial-transcriptomics resource should therefore be regarded as a foundation for further investigation rather than a complete representation of kidney spatial biology. Its value lies in providing a systematically characterized and quantitatively assessed basis on which subsequent studies can build.