| Variable | Missing n | Missing % |
|---|---|---|
| excluded_fruits_active | 12 | 100 |
| excluded_fruits_rem | 12 | 100 |
| excluded_gluten_active | 12 | 100 |
| excluded_gluten_rem | 12 | 100 |
| excluded_nuts_seeds_active | 12 | 100 |
| excluded_nuts_seeds_rem | 12 | 100 |
| excluded_spicy_foods_rem | 12 | 100 |
| excluded_vegetables_rem | 12 | 100 |
| excluded_whole_grains_active | 12 | 100 |
| excluded_whole_grains_rem | 12 | 100 |
| if_‘other’,_please_specify | 12 | 100 |
| excluded_spicy_foods_active | 9 | 75 |
| excluded_fat_foods_rem | 6 | 50 |
| excluded_lactose_active | 6 | 50 |
| excluded_lactose_rem | 6 | 50 |
Participant characteristics: missingness audit
1 Overview
This post documents missingness in data/processed/cleaned_characteristics.csv after the cleaning pipeline in src/characteristics/characteristics.R. The goal is descriptive: quantify where data are absent, which variable groups are most affected, and how complete the cohort is on fields used in downstream analyses. No hypothesis test is performed.
Figures generated by
src/characteristics/01_missingness_eda.Rviamake characteristics.
2 Data
- Unit of analysis: Participant (one row per survey record in the characteristics file).
- Input:
data/processed/cleaned_characteristics.csv(make characteristicsondata/raw/SYN_Participant Characteristics.xlsx). - Scope: All 146 columns after cleaning, including demographics, IBD clinical scores, symptom-frequency items, food-avoidance blocks, and ancillary survey fields.
Cohort size: 12 participants, 146 variables.
Columns with no missing values: 128 (87.7% of variables).
Median participant missingness: 9.9% of variables missing per person (max 11.6%).
Complete on core clinical/demographic fields (10 variables): 12 / 12 participants.
3 Methods
| Item | Choice |
|---|---|
| Missing definition | NA after cleaning (blank cells, failed coercion, or scrubbed impossible values) |
| Column summaries | Count and percent missing per variable |
| Participant summaries | Percent of variables missing per participant |
| Variable grouping | Rule-based labels (demographics, IBD clinical, symptom frequency, food avoidance, etc.) |
| Statistical test | None — exploratory audit only |
4 Figures
4.1 Missingness bands across variables
How many columns fall into each missingness range.
4.2 Variable-level missing rates
All variables with at least one missing value, coloured by survey domain.
4.3 Participant-level burden
Distribution of how many variables are missing per participant.
4.4 Core clinical and demographic fields
Missingness pattern for fields most often used in symptom and cohort summaries.
5 Highest-missing variables
6 Interpretation
Missingness varies across the survey. After cleaning, the demographic fields (age, gender, ethnicity, comorbidities) have no missing values, and BMI is complete. Conditional follow-up fields, such as the food-avoidance “excluded …” columns, are a median 100% missing. These only apply when a participant reports avoidance, so high missingness is expected.
All 30 symptom-frequency items are complete. 12 of 12 participants have complete data across the 10 core clinical fields listed above.
Implications for analysis:
- Treat food-avoidance and free-text exclusion columns as sparse optional modules, not core covariates.
- For symptom scores, report available N per variable and consider sensitivity analyses restricted to complete cases.
- The cleaning pipeline imputes or recodes some fields (e.g. median age for missing age,
"missing"ethnicity,"none"comorbidities); this audit reflects post-cleaning missingness on fields left asNA.
7 Reproducibility
- Cleaning:
src/characteristics/characteristics.R - Missingness figures:
src/characteristics/01_missingness_eda.R - Rendered:
make characteristicsthenquarto render statsfrom repository root