← All publications

Preprint

Evidence Behind the Automation of Clinical Trial Statistical Programming: A Scoping Review of Technology Adoption, Validation Frameworks, and AI/ML Integration (2020–2025)

Jaime Yan, Jason Zhang, Tingting Tian

medRxiv · 2025 · cited by a Cytel author at PHUSE US Connect 2026

DOI · 10.64898/2025.12.24.25342988

Abstract

This scoping review (reported per PRISMA-ScR) maps the 2020–2025 evidence base on automating clinical trial statistical programming — from macro-based tooling to LLM-driven approaches — characterizing study types, claimed efficiency gains, evaluation rigor, and open gaps. From 1,247 records, 262 studies were included; reported gains include 15–25% development-time reduction for pharmaverse TLF tools, 30–50% effort reduction for risk-based validation with CI/CD, and 75–85% SDTM conversion-time reduction for REDCap2SDTM, alongside 88–93% F1 for domain-specific LLMs on clinical NLP versus 60–85% code-generation accuracy for general models. Evidence quality is predominantly Low to Very Low: only 12 of 527 validation papers (2.3%) report quantitative outcomes, and no RCTs comparing validation approaches exist, defining the critical research priorities for the field.

Keywords

scoping review · TLF automation · validation · statistical programming · clinical trials · automation · LLM

Cite

Jaime Yan, Jason Zhang, Tingting Tian. "Evidence Behind the Automation of Clinical Trial Statistical Programming: A Scoping Review of Technology Adoption, Validation Frameworks, and AI/ML Integration (2020–2025)." medRxiv, 2025. doi:10.64898/2025.12.24.25342988.