Conference Paper
A Novel Pipeline for Generating Realistic Synthetic CDISC ADaM Datasets Using Large Language Models and Knowledge Graphs
Jaime Yan, Chao Su
PhUSE US Connect 2025 · 2025 · ML12
Abstract
Creating synthetic ADaM data that faithfully represents real trial characteristics is challenging. This paper combines knowledge graphs built from clinical trial documentation (protocols, SAPs, CRFs) with Faker-based generation, using LLMs to enrich and reorganize the JSON schemas that drive data generation. A three-way comparison of direct JSON, LLM-enhanced JSON, and template-based generation shows the template-based variant achieving the highest overall quality score (0.70 vs 0.45 for direct JSON-schema generation), with improved structural integrity and cross-dataset relationships.
Keywords
synthetic data · ADaM · CDISC · knowledge graphs · LLM · Faker
Cite
Jaime Yan, Chao Su. "A Novel Pipeline for Generating Realistic Synthetic CDISC ADaM Datasets Using Large Language Models and Knowledge Graphs." PhUSE US Connect 2025, 2025. doi:10.5281/zenodo.22182901.