Scientific Infrastructure
Tools & Computational Resources
Bioconductor packages, PyPI libraries, containerized workflows, and benchmark interactomes authored and maintained by Dr. Qingzhou Zhang.
SpatialRenal ↗
High-resolution spatial transcriptomics atlas of the kidney reconciling Visium HD, Xenium, MERFISH, GeoMx, and Nanostring data for detailed spatial structure and cellular organization.
Epi-Flow ↗
Unified reproducible pipeline for chromatin accessibility and occupancy assays: ATAC-seq, CUT&RUN, and ChIP-seq. Container-ready with single-command execution.
scfeatureprofiler ↗
Multi-interface Python package for deep characterization and statistical profiling of gene expression patterns, marker distribution, and co-expression in single-cell data.
ComplexMap ↗
Comprehensive R toolset for functional analysis, topological exploration, and publication-ready visualization of protein complex and interactome data.
SMAD (Bioconductor) ↗
Bioconductor package for statistical scoring of affinity purification–mass spectrometry (AP-MS) and proximity-dependent biotinylation (BioID) to discover high-confidence PPIs.
ShinyCITExpresso ↗
Interactive Shiny application for multi-modal CITE-seq data exploration, offering joint visualization of surface antibody-derived tags (ADT) and mRNA transcriptomes.
GEX_Index_Builder ↗
Automated RNA-seq reference index builder for STAR and Salmon with container-aware memory allocation, automated GENCODE/Ensembl retrieval, and version locking.
RefInt ↗
Curated reference interactome benchmark suite with experimentally validated true positives and negative controls for benchmarking protein-protein interaction prediction algorithms.
ExpressoGEO ↗
R package for streamlined retrieval, parsing, and quality-controlled preprocessing of Gene Expression Omnibus (GEO) datasets.
The Kidney Spatial Transcriptomics Analysis
Systematic re-analysis of 128 publicly available spatial transcriptomics samples across five major platforms: GeoMx, Visium, Xenium, NanoString, and MERFISH. Mapping spatial cellular niches, microenvironmental boundaries, and cross-platform concordance.
CAR T Resistance Mechanisms in Alveolar Soft Part Sarcoma (GCAR1 First-in-Human Trial)
Deep spatiotemporal multi-omics re-analysis of the GCAR1 trial integrating longitudinal scRNA-seq, single-cell TCR V(D)J repertoires, and sub-micron Visium HD spatial transcriptomics. Mapped CAR-T exhaustion trajectories, perivascular resistance barriers, and epitope spreading.
Atherosclerotic Plaque Smooth Muscle Cell Plasticity (ApoE KO Mice)
Characterized cellular heterogeneity and lineage transitions among clonally expanding vascular smooth muscle cells (VSMCs) in atherosclerotic lesions of ApoE-deficient mice using SMART-seq2 full-length single-cell sequencing.
ATF3 Knockdown Transcriptomic & Secretome Profiling in VSMCs
Comprehensive bulk RNA-seq and functional enrichment analysis characterizing transcriptomic and secretomic rewiring after ATF3 knockdown in vascular smooth muscle cells, deciphering downstream metabolic and contractile phenotypes.
Unified Epigenomics Workflow: ATAC-seq, CUT&RUN, ChIP-seq
Single-command automated pipeline supporting multiple chromatin profiling assays with quality metrics, peak calling, consensus matrix generation, and motif enrichment.
Automated RNA-seq Reference Index Generator
Container-aware indexing tool for STAR and Salmon with automatic Ensembl/GENCODE releases, memory detection, and version-controlled indices.
Pre-Flight Omics Study Design Checklist
Review these critical gates before submitting library pools for sequencing:
Verify that condition of interest (treatment vs control) is never perfectly correlated with batch, technician, extraction date, or sequencing lane.
Confirm that N represents independent biological organisms/donors, not technical replicate splits or multiple cells from a single animal treated as independent observations.
For single-cell perturbation experiments, confirm that control pools sample the baseline cell-state distribution sufficiently to detect shifts in rare subpopulations (< 2%).
Determine whether total RNA content is expected to change globally (e.g. Myc overexpression, acute cell death) and plan synthetic spike-ins or orthogonal cell counting accordingly.
Curated Reading: Foundations of Predictive Biology
Essential literature for transitioning from descriptive omics to dynamical systems and predictive modeling:
Waddington landscapes, attractor basins, and cell-state manifolds
Foundational theory framing cellular identity as low-dimensional dynamical attractors rather than static taxonomy.
Combinatorial CRISPR screens and cellular response vector fields
Empirical methodologies measuring multi-dimensional transcriptional shifts in response to targeted genetic knockouts.
Disentangling direct regulatory mechanisms from downstream trans-cascades
Statistical frameworks for inferring directed gene regulatory networks and counterfactual predictions.
Stress-testing single-cell foundation models against simple linear baselines
Critical empirical benchmarks evaluating whether large pre-trained models capture true biology or batch confounders.