Submission Week: Reviewers Start Here
SUBMISSION WEEK · eCTD MODULE 5
Reviewers Start Here
Metadata is read before the datasets.
Reviewers use define.xml as the proxy for program control.
If define.xml is not valid
Typical comments before filing
No valid define.xml means a folder of transport files whose labels nobody can resolve.
Variable: no origin | Method: “Derived per program” | Codelist: drifted at amendment
Lesson scope — define-XML elements · spec-driven generation · ADRG · pre-ship defect sweep
Opening scene using a concrete industry situation: a submission package at eCTD Module 5 is gathering reviewer comments about its define.xml. The senior programming lead frames why metadata, not datasets, is read first.
Speaker notes
This is submission week, and in the Electronic Common Technical Document, or eCTD, Module 5, reviewers start here, with Define-XML, the metadata standard built on Extensible Markup Language. They read the metadata before they open a single dataset, using Define-XML as a proxy for how controlled the whole program is. If Define-XML is not valid, the package becomes a folder of transport files whose labels nobody can resolve. Typical pre-filing comments include a variable with no origin, a method that just says derived per program, and a codelist that drifted from the data at an amendment. So in this lesson we will cover the Define-XML elements, spec-driven generation, the Analysis Data Reviewer's Guide, or ADRG, and a pre-ship defect sweep.
The Contract, Not the Documentation
The Contract, Not the Documentation
What define-XML is: one XML file that describes every dataset and variable in the package.
It defines purpose, label, type, length, origin, controlled terminology, and derivation.
Machine-Readable
Reviewer tools parse the XML. A wrong codelist reference is a broken binding — not a cosmetic defect.
Everything Is a Claim
An origin of CRF claims the annotated CRF backs it. A MethodDef claims the derivation ran as described.
Reviewers sample claims against data — metadata is evidence of study-program discipline.
Defines what define-XML actually is and the two properties that follow from reviewer tools consuming it: it is machine-readable, and everything in it is a claim about the program.
Speaker notes
Define-XML, the Extensible Markup Language standard, is not just documentation; it is the machine-readable contract for your submission package. In version 2.1, one XML file describes every dataset and variable: purpose, label, type, length, origin, controlled terminology, and derivation. Reviewer tools parse that XML, so a wrong codelist reference is a broken binding, not a cosmetic defect. Every element is a claim: an origin of Case Report Form, or CRF, claims the annotated CRF backs that value, and a MethodDef claims the derivation happened as described. Reviewers sample those claims against the data, so your metadata reads as evidence about the discipline of the study program.
Five Elements Carry the Weight
Five Elements Carry the Weight
Element → spec object → reviewer verification for Define-XML
| Element | Define-XML role | Review in the spec column |
|---|---|---|
| ItemGroupDef | One dataset: class, structure, purpose, sort keys | Dataset header |
| ItemDef | One variable: label, datatype, length | Origin col: CRF / derived / assigned / predecessor |
| CodeList | Controlled terminology (CT) + version | CT column: CDISC / sponsor / MedDRA |
| ValueListDef + WhereClauseDef | Value-level rule varies by parameter | Transformation column |
| MethodDef / CommentDef | Derivation or comment text | Transformation & notes columns |
Blank spec-column cell → that element ships wrong
Presents the element-to-spec-to-reviewer map for ItemGroupDef, ItemDef, CodeList, ValueListDef with WhereClauseDef, and MethodDef/CommentDef.
Speaker notes
Five elements carry the weight in Define-XML. The ItemGroupDef maps to a dataset, with class, structure, purpose, and sort keys from the dataset header, while the ItemDef describes a variable with label, datatype, length, and origin columns. That origin tells you if a value came from a Case Report Form, or CRF, was derived, assigned, or inherited from a predecessor dataset. The CodeList binds a variable to Controlled Terminology, or CT, such as a Clinical Data Interchange Standards Consortium, or CDISC, list, a sponsor-defined list, or an external dictionary like the Medical Dictionary for Regulatory Activities, or MedDRA, with its version pinned. ValueListDef and WhereClauseDef handle value-level rules when a rule varies by parameter, from the Transformation column, and MethodDef or CommentDef carry derivation text from Transformation and notes columns. If a spec-column cell is blank, that element ships wrong.
One Variable, Many Rules: Value-Level Metadata
One Variable, Many Rules
Value-Level Metadata
| Metadata Element | Role at Value Level |
|---|---|
| ValueListDef | Holds one definition for each branch of the variable. |
| WhereClauseDef | States the condition that selects which branch to use. |
| MethodDef | Carries derivation text a reviewer reads instead of the SAS program. |
| CommentDef | Carries comment text attached to the value-level definition. |
Why does one parameter differ from another? The answer is in the spec row — not in the SAS program.
Explains why one variable may need separate metadata per parameter and how ValueListDef, WhereClauseDef, MethodDef, and CommentDef implement that branch.
Speaker notes
This slide introduces value-level metadata. Value-level metadata covers the case where one variable carries different definitions for different parameters. A ValueListDef attaches a separate definition to each branch of that variable. A WhereClauseDef states the condition that selects which branch to use. A MethodDef and a CommentDef carry the derivation and comment text a reviewer reads instead of the Statistical Analysis System, or SAS, program. A reviewer uses this structure to ask why one parameter was derived differently, and the answer began as a spec row, not as code.
Metadata Is Build Output, Not Authoring
Metadata Is Build Output, Not Authoring
▪ One source of truth regenerates datasets, define.xml, and define.pdf together.
Origin column → ItemDef origin attribute
CT column → CodeList reference
Transformation column → MethodDef
Parameter-dependent → ValueList + WhereClause
• Hand-authored define.xml severs the chain
• define.pdf mismatch signals a broken generator
Explains spec-driven generation: origin, CT, and transformation columns become ItemDef, CodeList, and MethodDef elements; hand-editing severs the chain.
Speaker notes
Here is the mindset shift for this scene: in a spec-driven statistical computing environment, one source of truth regenerates your datasets, your Define-XML file, and your define.pdf together. That means each mapping column drives a specific metadata piece: the origin column becomes the ItemDef origin attribute, the controlled terminology column becomes a CodeList reference, and the transformation column becomes a MethodDef. If a transformation varies by parameter, you use a ValueList plus a WhereClause. Hand-authoring define.xml severs that chain, because a program edit with no regeneration means your package describes a study that no longer exists. And define.pdf is just the human-readable rendering of the same metadata from the same source, so if the two disagree, your generation process is broken.
Checkpoint: Element-to-Spec-to-Reviewer Map
1 In a define.xml file for a submission, which element carries the 'origin' attribute for one variable?
2 Which structures in define.xml are used to implement value-level metadata? (Select all that apply.) (select all that apply, then Check)
3 Which two properties make define.xml a contract rather than documentation? (Select all that apply.) (select all that apply, then Check)
Speaker notes
Time for a checkpoint quiz on how Define-XML, which is built on Extensible Markup Language, or XML, maps elements to specifications and reviewers. First, which element carries the origin attribute for one variable in a define dot XML file? The correct answer is A, ItemDef, because ItemDef defines a single variable and carries its origin metadata, such as whether it is collected or derived. Next, which structures implement value-level metadata? The correct answers are A, ValueListDef, and B, WhereClauseDef, because ValueListDef lists the value-level conditions, and WhereClauseDef defines when a value-level metadata block applies. Finally, which two properties make define dot XML a contract rather than documentation? The correct answers are A, it is machine-readable, and B, it asserts that the metadata reflects the actual submitted data, because these properties enable automated validation and give reviewers confidence that the metadata maps to the real datasets.
Walkthrough: SUPPDM Rows for Study 043-18101
SUPPDM Rows: Study 043-18101
| USUBJID | QNAM | QLABEL | QVAL | QORIG | QEVAL |
|---|---|---|---|---|---|
| 043-18101-74001-001 | ETTB | End of Treatment Tumor Biopsy? | Y | CRF | INVESTIGATOR |
| 043-18101-74001-001 | FRBS | Future Research on Biological Samples? | Y | CRF | INVESTIGATOR |
| 043-18101-74001-001 | PSCRNFAL | Previous Screen Fail? | N | CRF | INVESTIGATOR |
| 043-18101-74001-002 | ETTB | End of Treatment Tumor Biopsy? | Y | CRF | INVESTIGATOR |
| 043-18101-74001-002 | FRBS | Future Research on Biological Samples? | Y | CRF | INVESTIGATOR |
• Long format: one row per QNAM; QVAL/QORIG/QEVAL on it
• IDVAR/IDVARVAL blank ⇒ USUBJID-level, not a DM row
• PSCRNFAL absent on 002 ⇒ absence is not an implied N
• QORIG = CRF → the value came from the CRF
• QEVAL = INVESTIGATOR → per the investigator
Real-data walkthrough using only the supplied SUPPDM extract: reading long-format supplemental qualifier rows and the origin metadata each row carries.
Speaker notes
Here we trace Supplemental Qualifiers for Demographics, or SUPPDM, rows for this study. For the first subject, ETTB has label End of Treatment Tumor Biopsy? and value Y; FRBS has label Future Research on Biological Samples? and value Y; PSCRNFAL has label Previous Screen Fail? and value N. The second subject has only ETTB and FRBS, each value Y. Every row has origin Case Report Form and evaluator INVESTIGATOR, so each qualifier says the value came from the Case Report Form per the investigator. The identifying variable and its value are blank, so these are subject-level qualifiers attached to the unique subject identifier, not to a demographics row. PSCRNFAL is absent for the second subject, so absence is not an implied N; these rows drive the metadata obligations in define.xml and the reviewer's guide.
Hands-On: Element to Spec to Reviewer Question
Hands-on interactive — if it does not load, open the paired article and try the exercise there.
Speaker notes
This one is a hands-on exercise, and you run it in the browser on the website at jaimeyan.com/learn, not here in the video. You will match the Define-XML structures — ItemGroupDef, ItemDef, CodeList, ValueListDef with a WhereClauseDef, and MethodDef or CommentDef — to the spec columns and the reviewer questions they answer, using the Supplemental Qualifiers for Demographics, or SUPPDM, extract rows, where ETTB, FRBS, and PSCRNFAL each show QORIG of CRF and QEVAL of INVESTIGATOR, and you have to decide what the define metadata is allowed to claim. So after the video, go try it yourself, and hold on to the safe response to data change: update the spec row, regenerate define.xml, and revalidate — never hand-edit the XML.
The ADRG: A Human Preface, Not a Duplicate
The ADRG: A Human Preface,
Not a Duplicate
What reviewers read in an ADRG — in a fixed order
| Section | Why the reviewer reads it |
|---|---|
| 1 · Orientation | Dataset inventory; trace back to SDTM |
| 2 · Special derivations | Composite endpoints · custom windows · unusual imputations |
| 3 · Population filters | Each denominator is tied to a flag and an analysis |
| 4 · Known issues | Caveats surfaced up-front — not discovered late |
define.xml says what is there; the ADRG says why — don't make reviewers reread define.xml line by line.
Introduces the Analysis Data Reviewer's Guide as the human preface to define.xml and the four things reviewers actually read in it.
Speaker notes
The Analysis Data Reviewer's Guide, or ADRG, is the human preface to the machine-readable contract, and reviewers read it in a fixed order. First, orientation gives the dataset inventory and traces back to the Study Data Tabulation Model, or SDTM. Second, special derivations explain the analyses that break the standard pattern: composite endpoints, custom windows, and unusual imputations. Third, population filters tie each denominator to a flag and an analysis, so every table denominator traces back to this section. Fourth, known issues surface caveats up front, rather than letting reviewers discover them late. Remember the division of labour: define.xml says what is there, while the ADRG says why it is like that, and restating define.xml line by line just wastes the reviewer's attention.
The Defect Gallery and the Pre-Ship Sweep
The Defect Gallery and the Pre-Ship Sweep
Missing origins — blank/vague spec fails conformance
Value-level gaps — derivation logic lives only in SAS
CT drift — declared CT/version stale; fix: regenerate per cycle
Unexplained custom domains — no mapping or origin cited
Pre-Ship Sweep — merge actual variables vs spec;
flags: "in data, missing from spec" / "in spec, not built."
L3 agents (asOf 2026-09-01) — draft ADRG prose, diff define.xml vs datasets;
every claim must cite origin attr + ItemGroupDef.
Presents the four recurring define.xml and ADRG defects, the variable-versus-spec sweep that catches drift, and where modern workflow layers apply.
Speaker notes
Now let's walk through the defect gallery, the four problems that cause most findings. Missing origins, value-level gaps, controlled terminology (CT) drift, and unexplained custom domains. Blank or vague origins fail conformance checks outright, and derivations that live only in SAS become value-level gaps that conformance engines may miss. CT drift means define.xml declares a codelist or version your datasets stopped matching at an amendment; regenerate and revalidate every data cycle, not once before filing. The pre-ship sweep merges actual dataset variables against your spec and flags 'in data, missing from spec' or 'in spec, not built' before conformance runs. Level 3 agents, as of 2026-09-01, can draft Analysis Data Reviewer's Guide (ADRG) prose and diff define.xml versus datasets, but their failure mode is fluent false claims; require every claim to cite the origin attribute and ItemGroupDef.
Final Check: Contract to Package
1 If define.pdf disagrees with define.xml during a final package check, which explanation and response is most appropriate?
2 A study amendment changes a codelist that appears in the submitted metadata. What must happen as part of the update workflow? Select all that apply. (select all that apply, then Check)
3 List the four ADRG obligations a reviewer reads in order, and explain why the ADRG must not duplicate the content of define.xml. (reflect, then reveal)
Reveal analysis
Speaker notes
This is the final checkpoint quiz, tying the spec row all the way to define.xml, the Analysis Data Reviewers Guide, and the pre-ship sweep. If define.pdf and define.xml disagree during a final package check, the best answer is B: define.pdf is a stale generated view, so fix the generation step and regenerate define.pdf from the current define.xml. That is because the disagreement is a generation failure, not an isolated file mistake. When an amendment changes a codelist, select A, B, and C: update the codelist row in the source specification first, regenerate define.xml from that corrected source, and re-run validation against the exact amended data cut. The source specification is the contract, so fixing it first prevents stale metadata. For the short answer, the four Analysis Data Reviewers Guide, or ADRG, obligations in order are: orient the reviewer by stating the submission purpose and identifying the analysis datasets covered; explain analysis-specific methods such as major derivations and imputation rules; trace data and results by connecting source data, analysis datasets, and reported results; and point to define.xml for metadata instead of duplicating it, because define.xml is the single source of truth and duplication creates inconsistency.
Key Takeaways: From Spec Row to Front Door
Key Takeaways: From Spec Row to Front Door
1. Define-XML is the machine-readable contract of the SDTM/ADaM package; the ADRG is its human preface.
2. Generate define.xml from the spec — never hand-edit; regeneration keeps spec, programs, and metadata in lockstep.
3. Sweep the four recurring defects before conformance does: missing origins, value-level gaps, CT drift, custom domains.
4. Conformance engines check consistency, not truth; the MethodDef-versus-program match is a human diff built on generation.
5. Mirrors the tutorial 'Define-XML and the Reviewer's Guide, Explained for Programmers' on the Submission learning line.
Summary slide linking the course back to the source tutorial and the Submission learning line.
Speaker notes
Here's what I want you to remember from this lesson. Define-XML is the machine-readable contract for your Study Data Tabulation Model and Analysis Data Model package, while the Analysis Data Reviewer's Guide, or ADRG, is its human preface. Always generate define.xml from your specification and never hand-edit it; regeneration keeps spec, programs, and metadata in lockstep. Sweep the four recurring defects before conformance does: missing origins, value-level gaps, Controlled Terminology drift, and unexplained custom domains. Conformance engines check consistency, not truth, so the MethodDef-versus-program match remains a human diff built on generation. This mirrors the tutorial 'Define-XML and the Reviewer's Guide, Explained for Programmers' on the Submission learning line.