Project dossier
OngoingMicrobiome IBD Classifier
Can gut microbiome profiles tell Crohn's disease from ulcerative colitis across independent cohorts? A reported 0.91 ROC-AUC turned out to be evaluation leakage; the honest number is 0.64.
0.91 → 0.64
ROC-AUC before and after removing cohort leakage from the evaluation
- My role
- Sole author [TODO: confirm; the README clone URL points at a collaborator's fork]
- Team
- [TODO: solo / course project / collaborators]
- When
- [TODO: e.g. Jan to Apr 2026]
- ML
- Data
Distinguishing Crohn's disease from ulcerative colitis using stool microbiome profiles, evaluated across independent research cohorts. The first version of this project reported a 0.91 ROC-AUC. That number was wrong, and finding out why is the most useful thing the project produced.
[TODO: one line of context: course, term, solo or with whom.]
At a glance
- 0.913 to 0.637 binary ROC-AUC after rebuilding the evaluation so no study cohort appears on both sides of a split
- 2,105 IBD samples, 2,893 species-level features, 12 cohorts. Two cohorts hold roughly 80% of the IBD data
- Three separate leaks found and closed: feature selection, early stopping, and threshold tuning had each been allowed to see test data
- The corrected model beats the majority-class baseline by 1.6 accuracy points. The 3-class model does not clear its baseline at all
Problem
Crohn's disease and ulcerative colitis are the two main forms of inflammatory bowel disease. They share symptoms, but UC is confined to the large intestine while Crohn's can affect any part of the digestive tract, and the treatments differ. Colonoscopy, biopsy, and endoscopy are invasive and expensive, and still fail to separate the two in 10 to 20% of cases, leaving patients labelled "indeterminate colitis" until more symptoms emerge.
Stool sampling is non-invasive, cheap, and can pick up IBD at an inactive stage. The question was whether species-level gut microbiome abundances carry enough signal to tell the two diseases apart, and, more importantly, whether that signal survives a move from one research cohort to another.
Approach
The data, and the fact that bounds everything
Two public tables: patient metadata (project, run, phenotype, country, sequencing type) and long-format species abundances. Filtering to IBD-relevant projects gave 7,478 samples across 46 projects. An inner merge with the abundance table dropped that to 2,838 samples across 12 projects, because 4,640 samples had no abundance profile. The merge also removed nearly every amplicon sample, leaving 100% shotgun metagenomics, which is why species-level features are used throughout.
The drop from 46 projects to 12 is the single most consequential fact in the project. Of the 2,105 IBD samples, two projects hold about 80%, and UC appears in only 7 of the 12. However many rows there are, grouped cross-validation is testing against a small number of independent sources.
Finding the leak
The original pipeline used a random train/test split. Samples from the same study share sequencing platform, protocol, and population, so a model can score well by recognising the batch rather than the disease. On top of that, LASSO feature selection and XGBoost tuning had both been run with the test set visible. The reported 0.913 binary and 0.899 three-class ROC-AUC were largely the model memorising which cohort a sample came from.
I rebuilt the evaluation so that the study cohort is the unit that cannot leak:
| Leak | Control |
|---|---|
| Cohort across train and test | StratifiedGroupKFold grouped on project_id; a project is never split |
| Feature selection saw test rows | LASSO refit inside each fold on that fold's training rows only |
| Early stopping saw test rows | XGBoost stops on a grouped validation set carved from training |
| Threshold tuned on test | Decision threshold chosen on validation, applied to test unchanged |
| Degenerate folds | A geometry search rejects fold counts that strand a class, and reports when no clean option exists |
| Reporting | Pooled out-of-fold predictions, so every sample is scored once by a model that never saw it |
Every number below is out-of-fold. The three leaks are worth naming individually because each one looked like ordinary modelling code. None of them was a bug in the usual sense; they were the default way to write the pipeline.
Modelling choices
XGBoost as the primary model, Random Forest as the benchmark. With about 2,900 mostly-zero taxon columns, gradient boosting handles the sparsity natively and down-weights the many uninformative taxa, and its logistic objective gives probabilities for ROC-AUC and threshold analysis. L1-penalised logistic regression (C=0.1, GPU via cuML) does feature selection; refit on the full data for interpretation it keeps 746 taxa, but inside cross-validation it is refit per fold, so the selected set differs by fold on purpose. Neural networks and SVMs were not prioritised: with 12 independent cohorts the overfitting risk outweighs the sample count, and the goal was cohort-aware evaluation, not an exhaustive model bake-off.
What broke, twice. An early comparison of raw, log1p, and centred-log-ratio transforms showed no difference, and it could not have: tree ensembles split on thresholds and are invariant to any monotonic per-feature transform, and the one step that is sensitive to it, LASSO, was being fed raw values regardless of the setting. Once both paths were wired correctly (and CLR added specifically because it mixes information across taxa), the comparison ran for real. The answer was still negative, all three within one per-fold standard deviation. Separately, Random Forest recalls were byte-identical with and without class weighting, which means the installed cuML build silently ignores class_weight; the notebook now checks whether the parameter is exposed and says so.
Results
Baseline means always predicting the majority class. It is the number any model has to beat to be worth anything.
Binary, Crohn's vs UC (2,105 samples, clean 2-fold cohort geometry, 6 test projects per fold):
| Model | Accuracy | Baseline | Crohn F1 | UC F1 | ROC-AUC |
|---|---|---|---|---|---|
| Majority baseline | 0.624 | 0.500 | |||
| Random Forest | 0.61 | 0.624 | 0.69 | 0.49 | n/a |
| XGBoost | 0.64 | 0.624 | 0.74 | 0.46 | 0.637 |
Per-fold ROC-AUC 0.632 ± 0.001. The AUC is the meaningful figure: the model ranks UC above Crohn's better than chance, even though its hard predictions sit within two points of guessing the majority class.
Three-class, Healthy vs Crohn's vs UC (2,838 samples): XGBoost macro ROC-AUC 0.646 (chance 0.5), but accuracy 0.44 against a 0.463 baseline. Random Forest clears the baseline by 0.7 points. No fold count produced a clean cohort split for this task, one fold trains on 36% of the data, and a single reshuffle of fold assignment moved UC recall from 0.06 to 0.60 with no change to the model. Only the pooled out-of-fold figures should be quoted; the per-fold spreads are noise.
What the numbers do not say.
- Grouped CV cut apparent performance roughly in half, 0.913 to 0.637 binary and 0.899 to 0.646 three-class. That gap is a direct measure of how much the original result was batch memorisation.
- Fold assignment moves the result more than modelling does. With this few cohorts, which projects land in which fold matters more than any hyperparameter. Any single number here is one draw from a wide distribution.
- Threshold tuning is a trade, not a gain. Lowering the cutoff lifts UC recall to 0.98 while Crohn's recall collapses from 0.79 to 0.08; ROC-AUC does not move, because sliding along a fixed curve does not improve it. Which operating point to prefer is a clinical judgement about which misdiagnosis costs more.
- Crohn's and UC overlap at the species level. Two inflammatory conditions in the same organ system, with a 10 to 20% clinical misdiagnosis rate under colonoscopy. Expecting stool composition alone to separate them cleanly was optimistic.
- Not a diagnostic tool. Batch effects are controlled for in evaluation, not corrected in the data. Diet, geography, antibiotic use, and disease activity are all absent from the features.
The withdrawn numbers stay in the README on purpose. The gap between the two sets of results is the most instructive thing the project has to show.
What I'd do next
- Add genus-level features. Roughly half the abundance table is genus rows that the current filter discards. Genus is less sparse and less sensitive to cross-study differences in classifier reference databases, so it may carry signal that transfers across cohorts where species-level detail turns into noise. Run genus alone and genus plus species under the identical protocol to separate "more information" from "better resolution."
- Repeated cross-validation across seeds. One fold reshuffle moved UC recall by 54 points. A distribution over seeds is more honest than any point estimate.
- Batch correction inside the folds. Estimate correction parameters on training cohorts only, then apply to held-out cohorts. Fitting a correction globally would reintroduce exactly the optimism this evaluation was rebuilt to remove.
- More independent cohorts. The real fix, and unavailable at present. More rows from the existing 12 projects would not help; the limit is the number of sources.
Stack
Python · pandas, NumPy · scikit-learn · XGBoost · cuML / RAPIDS on a Colab T4 for LASSO and Random Forest · matplotlib, seaborn · FastAPI prediction API serving the binary model · React demo on Vercel