Data validation & prediction modelling
From a difficult raw dataset to a publication-ready clinical study
A 1,800-patient multicentre retrospective dataset, two weeks to a submission deadline, and a preliminary analysis that turned out to be wrong.
A clinical research team asked us to analyse a multicentre retrospective dataset of approximately 1,800 patients and build a risk model for postoperative complications. The data were described as ready for analysis. Structured validation found duplicate records, incompatible coding between hospitals, physiologically impossible values, inconsistent date formats, continuous variables stored as text, patients who did not meet the eligibility criteria, and unconsidered confounders. After cleaning and rebuilding the analysis plan, the final results differed materially from the team's preliminary analysis — several apparently significant predictors were no longer independently associated with the outcome, while other clinically meaningful predictors emerged once the model was correctly specified.
The brief
A clinical research team approached us with a multicentre retrospective dataset of approximately 1,800 patients. Their objective was to identify predictors of postoperative complications and develop a clinically useful risk model.
Data collection across several hospitals was already complete, so our role was not to recollect anything. We were responsible for:
- Validating the collected data
- Cleaning and restructuring the database
- Checking the statistical methodology
- Conducting the final statistical analysis
- Producing publication-ready tables and figures
- Supporting interpretation and manuscript preparation
The client believed the database was ready for analysis. Validation identified several issues that could have materially changed the study conclusions.
What validation found
Because the data had been collected by different researchers across multiple centres, a number of inconsistencies were present:
- Duplicate patient records
- Different coding systems between hospitals
- Implausible laboratory and clinical values
- Missing outcome information
- Dates entered in different formats
- Continuous variables entered as text
- Inconsistent definitions of comorbidities
- Patients who did not actually meet the eligibility criteria
- Several variables with substantial missingness
- Potentially important confounders that had not been considered in the original analysis
The original statistical plan relied mainly on univariate comparisons followed by a standard logistic regression model. On evaluation, some model assumptions were not adequately addressed, and the number of candidate predictors was high relative to the number of outcome events. Used as it stood, that approach would have produced an unstable, overfitted model — one that looks convincing in the development sample and fails everywhere else.
What we did
We created a structured data-validation protocol before performing any statistical analysis. Every variable was reviewed against its expected range, coding system, clinical definition, and relationship with other variables.
Values such as impossible ages, laboratory measurements outside physiologically plausible ranges, procedures recorded before admission dates, and contradictory medical-history variables were automatically flagged and then manually reviewed by a clinical researcher — flagged records were examined, not deleted.
Duplicate records were identified using combinations of patient identifiers, dates, demographic characteristics and clinical variables, rather than identifier matching alone. For multicentre variables coded differently between hospitals, we built a standardised data dictionary so that all centres mapped onto the same definitions.
Missing data were then assessed variable by variable. Instead of automatically deleting every patient with missing information, we evaluated the extent and pattern of missingness and applied an appropriate missing-data strategy where it was justified.
The statistical solution
After cleaning, we rebuilt the analysis plan around the actual research question. The analysis included:
- Descriptive analysis of the cohort
- Appropriate univariate comparisons
- Assessment of continuous-variable distributions
- Examination of non-linear relationships
- Multicollinearity assessment
- Clinically informed confounder selection
- Multivariable regression modelling
- Sensitivity analyses
- Internal model validation
- Assessment of discrimination and calibration
Rather than selecting predictors purely by univariate p-value, clinically important variables were considered alongside the statistical evidence. Where appropriate, continuous predictors were kept continuous instead of being arbitrarily categorised. For the prediction component, model performance was assessed through discrimination, calibration and internal validation — not by listing which predictors reached significance.
A second researcher independently reproduced the major analyses and checked every final number against the cleaned database.
The team and the audit trail
The project required a multidisciplinary team:
- A clinical researcher to evaluate medical definitions and clinical plausibility
- A biostatistician to develop and perform the statistical analysis
- A data analyst to clean and restructure the database
- A senior researcher to independently review methodology and results
- A scientific writer to ensure the analysis was accurately represented in the manuscript
The analysis was performed using reproducible statistical scripts so that every major result could be traced back to the cleaned dataset. A complete audit trail was maintained: raw data → identified issue → correction or decision → final dataset → statistical output → manuscript table or figure.
The two-week constraint
The hardest part of this project was the timeline. The researchers approached us shortly before a submission deadline, leaving approximately two weeks to validate the dataset, complete the analysis and prepare the major results. A database of this size and complexity would normally benefit from a substantially longer validation period.
We managed it with a parallel workflow rather than by cutting quality control:
- Data validation and cleaning were performed simultaneously by different team members.
- Clinical questions were escalated immediately to the medical research team instead of queued.
- Statistical scripts were developed while the final cleaned dataset was still being prepared.
- Tables and figures were generated automatically from the statistical code wherever possible.
- An independent reviewer checked the major results before delivery.
This reduced turnaround substantially without removing the major quality-control steps.
The outcome
After validation, a meaningful proportion of the original records required at least one correction, recoding decision or eligibility review.
More importantly, the final analysis differed from the researchers' preliminary analysis. Some variables that initially appeared statistically significant were no longer independently associated with the outcome after appropriate adjustment, while other clinically meaningful predictors became apparent once the model was correctly specified.
The final package included:
- A cleaned and documented database
- Data dictionary
- Reproducible statistical code
- Complete statistical output
- Publication-ready tables and figures
- Sensitivity analyses
- Interpretation of the main findings
- Statistical methods and results sections ready for the manuscript
The manuscript was then prepared around the validated analysis rather than the preliminary results.
What made the case difficult
The main limitation was time. Two weeks for a large multicentre dataset is far from ideal; a longer timeline would have allowed additional rounds of source-data verification and potentially more extensive external validation of the prediction model.
There were also limitations inherent to retrospective data. No statistical method can eliminate unmeasured confounding or variables that were never collected. Those limitations were reported explicitly rather than compensated for statistically — which is the only honest option.
The key lesson
The most important part of this project was not performing a sophisticated statistical test. It was identifying that the dataset and the original analytical strategy contained issues before those issues became published conclusions.
We do not simply analyse the spreadsheet we receive. We first determine whether the data are clinically plausible, methodologically appropriate, statistically analysable, and capable of supporting the conclusion the researchers want to investigate. Only after that validation do we proceed to the final analysis.
The same standard applies to your project
Structured validation before analysis, methods matched to the question, assumptions tested, sensitivity analyses run, every primary result independently reproduced, and reproducible scripts handed over with the output.
More evidence
Other documented projects.
Research at scale: ~100 papers a quarter, 80% published
Supporting an academic institution across design, analysis, writing and journal selection — at a volume that only works with a structured workflow.
Read the case study Manuscript rescueRejected by more than ten journals — accepted by the next one
Repeated rejection is usually a diagnosis problem. This manuscript had four separate ones — and a journal strategy that was never going to work.
Read the case studyLet’s make your next result defensible.
Book a free 15-minute consultation. Tell us the research question, the data you have, and the deadline — we will tell you honestly what is achievable.
Or call +20 100 163 8864 · Sunday–Thursday, 09:00–18:00 (GMT+2, Cairo)