IBD Mycobiome Dashboard

Dashboards
R
Shiny
Quarto
statistics
bioinformatics
data visualization
Shiny dashboard and statistics site linking gut fungi, diet, and inflammation in IBD, built with a hospital partner.
Author

Ian Gault

Published

September 24, 2026

Live Dashboard Statistics Site View on GitHub

This project explores how the gut mycobiome (the fungal community in the gut), diet, and inflammatory biomarkers relate to inflammatory bowel disease (IBD). It combines a data pipeline, an interactive Shiny dashboard with participant-level and cohort-level views, and a Quarto site with one post per statistical analysis.

The original work was a 2026 UBC Master of Data Science capstone built with a hospital research partner by Tiffany Chu, Victoria Farkas, Ian Gault, and Derrick Jaskiel. The original used real patient data under a data-use agreement and is kept private. This public version runs the same code on synthetic data, so the numbers and figures are not real findings.

What it does

  • Data pipeline - R scripts that clean and merge four sources (stool mycobiome relative abundance, dietary intake, inflammatory biomarkers, and participant characteristics) into analysis-ready files, orchestrated with GNU Make
  • Shiny dashboard - participant-level and cohort-level views of fungal composition, diet, and biomarkers
  • Statistics site - alpha diversity (Kruskal-Wallis), beta diversity (PERMANOVA at four taxonomic levels), symptom and nutrient associations, PCA and clustering, and an evidence-ranking matrix

Technical stack

Area Tools
Data pipeline R (tidyverse), GNU Make
Statistics PERMANOVA and Bray-Curtis distances (vegan), Kruskal-Wallis, PCA, clustering
App framework Shiny
Reporting Quarto
Reproducibility renv, testthat unit tests
CI/CD GitHub Actions (rebuilds and publishes the statistics site to GitHub Pages)
Deployment Posit Connect Cloud
Synthetic data Python (pandas, numpy)

What I changed for the public version

  • Wrote a synthetic data generator that matches the original raw file schemas, so the full pipeline runs end to end
  • Removed the dashboard login (shinymanager), since there is no real data to protect
  • Added a GitHub Actions workflow that rebuilds and publishes the statistics site
  • Removed patient data, results, and the capstone report