← All publications

Conference Paper

A Novel Pipeline for Generating Realistic Synthetic CDISC ADaM Datasets Using Large Language Models and Knowledge Graphs

Jaime Yan, Chao Su

PhUSE US Connect 2025 · 2025 · ML12

DOI · 10.5281/zenodo.22182901 Download PDF

Abstract

Creating synthetic ADaM data that faithfully represents real trial characteristics is challenging. This paper combines knowledge graphs built from clinical trial documentation (protocols, SAPs, CRFs) with Faker-based generation, using LLMs to enrich and reorganize the JSON schemas that drive data generation. A three-way comparison of direct JSON, LLM-enhanced JSON, and template-based generation shows the template-based variant achieving the highest overall quality score (0.70 vs 0.45 for direct JSON-schema generation), with improved structural integrity and cross-dataset relationships.

Keywords

synthetic data · ADaM · CDISC · knowledge graphs · LLM · Faker

Cite

Jaime Yan, Chao Su. "A Novel Pipeline for Generating Realistic Synthetic CDISC ADaM Datasets Using Large Language Models and Knowledge Graphs." PhUSE US Connect 2025, 2025. doi:10.5281/zenodo.22182901.