What actually determines whether a clinical ML model works in the field — the classifier, or everything you do to the data before the classifier ever runs?
A new paper out of ADNI territory gives me an excuse to argue the second answer. Debopriya Ghosh's master's work proposes a model to detect the early stages of Alzheimer's disease from clinical details, neuropsychological test scores, and neuroimaging-related measures, drawn from the Alzheimer's Disease Neuroimaging Initiative (ADNI). The headline result is a stacking ensemble — Logistic Regression, Extra Trees, Bagging KNN, and LightGBM as base classifiers — compared against a plain artificial neural network. That's the part that gets cited. It's not the part that matters most.
Let me define the terms first, because the reader should never have to guess. Alzheimer's disease is a slow brain disorder that mainly attacks memory, thinking, and language; in the early stage the symptoms look like normal ageing, which is exactly why people get diagnosed late. Early detection here does not mean cure — there is no complete cure for AD — it means buying the clinician time to manage the condition. So the function of this whole system in the larger loop is not diagnosis-as-verdict. It's a triage signal: flag the borderline patient early enough that a human specialist looks harder.
That framing changes what "good" means. A triage model that misses early cases (low recall) is worse than useless — it gives false reassurance. This is why the paper reports precision, recall, F1-score, and AUC-ROC rather than raw accuracy, and it's the correct instinct. Accuracy on an imbalanced medical dataset is a vanity metric. If 90% of your labelled patients are one class, a model that always predicts that class scores 90% and catches zero of the cases you actually care about.
Which brings me to the three decisions in this pipeline that I think carry the real weight — all of them upstream of the classifier.
missing values: iterative imputation is a modelling choice, not a cleanup step
The ADNI dataset has missing values, and the paper handles them with iterative imputation. In practice, this is where a lot of clinical ML quietly goes wrong. Iterative imputation models each feature-with-gaps as a function of the other features and fills the blanks with predictions. That's powerful — and it's also a second model living inside your pipeline, one that can leak information and manufacture correlations that were never in the raw data.
The trade-off is concrete. Drop rows with missing values and you throw away patients — often the sickest ones, because they miss appointments and skip tests. Impute aggressively and you invent structure. In a production clinical system, the imputer has to be fit on training data only and frozen, then applied to new patients at inference. Fit it on the full dataset and your reported AUC is inflated by a leak you'll never see in deployment. This is the single most common way I've seen medical pipelines look great on paper and collapse in the field.
class imbalance: SMOTE fixes the loss, not the world
The paper handles class imbalance with Borderline SVM-SMOTE. SMOTE — Synthetic Minority Over-sampling Technique — generates synthetic minority-class examples by interpolating between real ones. The "borderline" variant focuses on the hard cases near the decision boundary, which is a smarter place to spend synthetic samples.
Here's the applied caution I'd attach to any pipeline that uses it: oversample only the training split, never the test split, and never before the train/test cut. Synthetic points that leak across the split turn your evaluation into a lie. And oversampling changes the prior — the model now sees a balanced world that does not exist in the clinic. If you don't recalibrate the output probabilities back to the real base rate, your triage threshold is meaningless. SMOTE is a training-time convenience, not a statement about disease prevalence.
feature selection is the biomarker discovery, and that's the actual payoff
The paper uses wrapper-based and embedded methods to keep only important features for training, and states a second aim beyond raw prediction: identify important biomarkers for early diagnosis. This is the part I care about most, and it's underrated in the abstract's ordering.
A clinician does not want a black box that emits "71% Alzheimer's." They want to know which neuropsychological scores and imaging measures drove the flag, because that maps to something they can act on and defend. Wrapper methods (which repeatedly train the model on candidate feature subsets) and embedded methods (like the feature importances that fall out of tree ensembles and LightGBM) both produce a ranked shortlist. In a shipped system, that shortlist is the interface between the model and the human — it's what makes the output auditable.
I've built this exact shape of thing outside medicine. In Sentinel-5P AQI we derive an actionable air-quality index from satellite bands where ground sensors don't exist; the value isn't a single number, it's knowing which spectral inputs justified it and being able to defend the reading to a municipal auditor. Same structure here: sensitive, sparse data, a model, and a downstream human who needs the reasons, not just the score. The feature-selection stage is where a predictive toy becomes an instrument.
the ensemble question, answered honestly
So why stack four classifiers and also train a neural network? Because on tabular clinical data of this size, gradient-boosted trees and their ensembles usually beat deep nets, and the paper is right to compare rather than assume. This is not a knock on neural networks — it's a hardware-and-data-budget observation. A stacking ensemble of Logistic Regression, Extra Trees, Bagging KNN, and LightGBM trains in minutes on a laptop, is easy to inspect, and hands you feature importances for free. A deep ANN needs more data to justify its capacity and is harder to explain to a clinician. For a triage tool that has to run in a hospital IT environment and survive an audit, the ensemble is not the compromise choice — it's the engineering choice.
Knowledge here is infrastructure: a bounded set of information — ADNI's clinical, neuropsychological, and imaging measures — combined into a framework that a clinician can use to look harder at the right patient earlier. The classifier is one layer of that stack. The imputation, the resampling, and the feature selection are the load-bearing ones.
What I'd push back on, gently, is the instinct to read this paper as "ensemble beats ANN." That's the least transferable finding in it. The transferable lessons are the ones every applied engineer relearns the hard way: fit your imputer and your resampler inside the cross-validation fold, recalibrate your probabilities to the real prevalence, and treat feature selection as the product, not the housekeeping. Get those three right and almost any reasonable classifier will do. Get them wrong and the fanciest ensemble in the world will lie to you with a straight face.
The worst outcome for a tool like this isn't a slightly-lower AUC. It's a leak-inflated number that convinces a clinic to trust it, and a false-negative rate that only reveals itself one late diagnosis at a time. In production, the boring preprocessing decisions are the safety system.
Sources
- Computer Science > Machine Learning — Manual / ad-hoc · 2026-07-03