Participant characteristics: missingness audit

eda
characteristics
Descriptive audit of missing values in the cleaned participant characteristics file, including variable-level rates, participant-level burden, and core clinical coverage.
Author

IBD Capstone team

Published

June 16, 2026

1 Overview

This post documents missingness in data/processed/cleaned_characteristics.csv after the cleaning pipeline in src/characteristics/characteristics.R. The goal is descriptive: quantify where data are absent, which variable groups are most affected, and how complete the cohort is on fields used in downstream analyses. No hypothesis test is performed.

Figures generated by src/characteristics/01_missingness_eda.R via make characteristics.

2 Data

  • Unit of analysis: Participant (one row per survey record in the characteristics file).
  • Input: data/processed/cleaned_characteristics.csv (make characteristics on data/raw/SYN_Participant Characteristics.xlsx).
  • Scope: All 146 columns after cleaning, including demographics, IBD clinical scores, symptom-frequency items, food-avoidance blocks, and ancillary survey fields.

Cohort size: 12 participants, 146 variables.

Columns with no missing values: 128 (87.7% of variables).

Median participant missingness: 9.9% of variables missing per person (max 11.6%).

Complete on core clinical/demographic fields (10 variables): 12 / 12 participants.

3 Methods

Item Choice
Missing definition NA after cleaning (blank cells, failed coercion, or scrubbed impossible values)
Column summaries Count and percent missing per variable
Participant summaries Percent of variables missing per participant
Variable grouping Rule-based labels (demographics, IBD clinical, symptom frequency, food avoidance, etc.)
Statistical test None — exploratory audit only

4 Figures

4.1 Missingness bands across variables

How many columns fall into each missingness range.

4.2 Variable-level missing rates

All variables with at least one missing value, coloured by survey domain.

4.3 Participant-level burden

Distribution of how many variables are missing per participant.

4.4 Core clinical and demographic fields

Missingness pattern for fields most often used in symptom and cohort summaries.

5 Highest-missing variables

Variable Missing n Missing %
excluded_fruits_active 12 100
excluded_fruits_rem 12 100
excluded_gluten_active 12 100
excluded_gluten_rem 12 100
excluded_nuts_seeds_active 12 100
excluded_nuts_seeds_rem 12 100
excluded_spicy_foods_rem 12 100
excluded_vegetables_rem 12 100
excluded_whole_grains_active 12 100
excluded_whole_grains_rem 12 100
if_‘other’,_please_specify 12 100
excluded_spicy_foods_active 9 75
excluded_fat_foods_rem 6 50
excluded_lactose_active 6 50
excluded_lactose_rem 6 50

6 Interpretation

Missingness varies across the survey. After cleaning, the demographic fields (age, gender, ethnicity, comorbidities) have no missing values, and BMI is complete. Conditional follow-up fields, such as the food-avoidance “excluded …” columns, are a median 100% missing. These only apply when a participant reports avoidance, so high missingness is expected.

All 30 symptom-frequency items are complete. 12 of 12 participants have complete data across the 10 core clinical fields listed above.

Implications for analysis:

  • Treat food-avoidance and free-text exclusion columns as sparse optional modules, not core covariates.
  • For symptom scores, report available N per variable and consider sensitivity analyses restricted to complete cases.
  • The cleaning pipeline imputes or recodes some fields (e.g. median age for missing age, "missing" ethnicity, "none" comorbidities); this audit reflects post-cleaning missingness on fields left as NA.

7 Reproducibility

  • Cleaning: src/characteristics/characteristics.R
  • Missingness figures: src/characteristics/01_missingness_eda.R
  • Rendered: make characteristics then quarto render stats from repository root