Start Here: The Startup Package
Start Here: The Startup Package
New package on the drive: protocol + SAP + email flagging two key sentences.
Nobody flags the sentence that fixes the early-visit efficacy read.
Extraction, not reading — the habit that sets study-startup pace.
Deliverables: analysis datasets & TLFs, not document summaries.
Reviewer
reads for compliance
Statistician
reads for intent
Programmer
reads clauses into code
Open with the concrete work situation a junior statistical programmer faces: a new study lands with a protocol, an SAP, and a statistician's email, and somewhere inside is the rule that decides an analysis question nobody has pointed to.
Speaker notes
Welcome to the startup package — this is where your work on a new study really begins. You get the protocol, the Statistical Analysis Plan, an email that flags two sentences, and a shared drive letter. Somewhere in those pages is the rule that decides whether an early visit belongs in the efficacy analysis, and no one has told you which sentence it is. That is extraction work, not reading. Your deliverables are analysis datasets and tables, listings, and figures — not document summaries. A reviewer reads for compliance, a statistician reads for intent, but a programmer reads the clauses that will become code.
Read Like a Producer, Not a Reviewer
Read Like a Producer,
Not a Reviewer
| If the source says… | Build this artifact… |
|---|---|
| Any protocol or SAP rule sentence | Owns one named artifact: analysis dataset · ADSL flag · spec row · logged query |
| Monitor → compliance · Statistician → intent · Programmer → code | Read as a sentence to implement, not background prose |
| “Efficacy evaluated in full analysis set” | ADSL FAS flag + denominator decision |
| “AE onset on/after first dose” | TEAE flag convention in ADAE |
| Protocol explains why · SAP states what | Disagreement? SAP governs · discrepancy is a query |
Establish the core reading stance: every rule sentence in a protocol or SAP must be named for the artifact it owns — a dataset, a flag, a spec row, or a logged query.
Speaker notes
As a statistical programmer, you read study documents as a producer, not a reviewer. The monitor reads the protocol for compliance, the statistician reads the Statistical Analysis Plan (SAP) for analysis intent, and you read both for the sentence that becomes code. Every rule sentence must own a named artifact: an analysis dataset, a Subject-Level Analysis Dataset (ADSL) flag, a spec row, or a logged query. For example, 'efficacy evaluated in the full analysis set' becomes a Full Analysis Set (FAS) flag in ADSL plus a denominator decision. Similarly, 'adverse event onset on or after first dose' becomes a Treatment-Emergent Adverse Event (TEAE) flag convention in the Adverse Events Analysis Dataset (ADAE). Remember, the protocol explains why a rule exists, the SAP states what you implement, and when they disagree, the SAP governs and the discrepancy is a query.
The Extraction Pass, In Order
The Extraction Pass, In Order
Run the eight steps in order — each ends in an artifact or a routed query
1 Endpoints
2 Populations
3 Visits & Windowing
4 Dates & Exposure
5 AE Rules
6 Derivations
7 Interim & DMC Flows
8 Shells
Endpoints first
Owner + TLF home
Populations → ADSL
SAP cite (DS, EX)
Shells check map
Walk TLF shells vs map
No private guesses
Map, flag, spec, query
Walk the eight-step extraction sequence that converts documents into a programmer work list, stressing that each step ends in an artifact or a routed query.
Speaker notes
Run the extraction pass in order: endpoints, populations, visits and windowing, dates and exposure, adverse event (AE) rules, derivations, interim and Data Monitoring Committee (DMC) flows, shells. Endpoints first: every primary and key secondary endpoint needs a dataset owner and a tables, listings, and figures (TLF) home. An endpoint without an owner is scope finding for the startup meeting, not week-nine trivia. Turn each population line into a Subject-Level Analysis Dataset (ADSL) flag with its Statistical Analysis Plan (SAP) citation, sourced from randomization records in the Disposition (DS) domain and documented dose in the Exposure (EX) domain. Shells close the loop: walk the TLF shells against the endpoint map so every output traces to an endpoint or a safety requirement. Each step ends in an artifact — a map, flag, specification row — or a logged query, never a private guess.
Populations and the ADSL Flags They Own
Populations and the ADSL Flags They Own
| Flag | SAP criterion (words) | Source records | ADSL flag output |
|---|---|---|---|
| ITT | Full analysis set: all randomized subjects | Randomization/disposition in DS | ITT flag + SAP citation |
| SAF | Safety set: any documented dose | EX exposure records | SAF flag + SAP citation; dose rule met |
• ADSL merge code implements definitions; it never creates definitions.
• Keep population rules in one extraction worksheet; it feeds mapping specs and define.xml.
Show how SAP population criteria in words are translated into source records and ADSL analysis flags with citations, and why the code only implements definitions it never creates.
Speaker notes
Population criteria start as words in the Statistical Analysis Plan, or SAP, and your job is to translate each one into a source and a Subject-Level Analysis Dataset, or ADSL, flag. For the intention-to-treat, or ITT, flag, the full-analysis-set definition of all randomized subjects is fed by randomization and disposition records in the Disposition, or DS, domain. For the safety analysis set, or SAF, the definition of any documented dose is fed by Exposure, or EX, records, and the flag is set only when the SAP's dose rule is met. The ADSL merge code implements these SAP definitions; it never creates them. Keep the population rules in one extraction worksheet, because that same worksheet can feed mapping specifications and define.xml, the metadata definition file, downstream.
Visits, Windowing, and the Week-Twelve Argument
Visits, Windowing, and the
Week-Twelve Argument
1. Windowing decides which scheduled visits enter the efficacy analysis; extract at Week 0, not the Week-12 argument.
2. For each parameter class set baseline definition, window anchor, and whether unscheduled visits can fill a window.
3. Date-and-exposure rules become branches in the first-dose cascade; a fallback branch must cite the SAP.
4. EX administration outranks an assumed DM start — derive TRT01SDT (first exposure date) from EX, never DM.
5. Pin the MedDRA version and controlled terminology in one define.xml — version drift is invisible to rule engines.
Explain why windowing, baseline, and date-and-exposure sentences decide whether BDS analysis-visit logic is correct and must be extracted at week zero instead of during late-stage disputes.
Speaker notes
Windowing decides which scheduled visits enter the efficacy analysis, so extract the windowing sentences at week zero, not at the week-twelve argument. For each parameter class, pin down the baseline definition, the window anchor, and whether unscheduled visits can fill a window. Date and exposure rules become branches in the first-dose cascade, and any fallback branch must cite the Statistical Analysis Plan, or SAP, because fallbacks are where studies differ. A documented administration outranks an assumed dose: derive the first exposure date variable, TRT01SDT, from Exposure, or EX, records, never from a reference start date in Demographics, or DM. Pin the Medical Dictionary for Regulatory Activities, or MedDRA, version and the controlled terminology package in one Define-XML file, define.xml. Version drift between domains is invisible to rule engines.
Concept Check: Artifacts and Sources
1 Your task is to set the intent-to-treat (ITT) flag in the subject-level analysis dataset (ADSL). Which source artifact owns the rule that determines whether a subject is in the full analysis set?
2 Under the canonical ADSL derivation, each unique subject identifier (USUBJID) is assigned a treatment start date (TRT01SDT). Which source dataset should supply the record used to derive TRT01SDT?
3 Select the two true statements about analysis-set conflicts and ADSL population flags. (select all that apply, then Check)
Speaker notes
This checkpoint quiz has three questions on artifact naming and key ADSL variable sources before the code walkthrough. For the intent-to-treat (ITT) flag in the subject-level analysis dataset (ADSL), the rule for full analysis set membership is owned by option B, the clinical study protocol. The full analysis set follows the ITT principle and is controlled by the protocol; the Statistical Analysis Plan (SAP) may restate it and define.xml documents metadata, but neither may supersede the protocol. Under the canonical ADSL derivation, which source dataset supplies the record for each unique subject identifier (USUBJID)'s treatment start date (TRT01SDT)? The correct answer is option B, Exposure (EX), because TRT01SDT comes from the earliest study-drug exposure record in the EX domain. For analysis-set conflicts and ADSL population flags, the correct choices are option B and option D. Option B says that when the SAP and protocol disagree, suspend derivation and ask the biostatistician to resolve the discrepancy, and option D says to set the Safety Analysis Set (SAF) flag to 'Y' for a USUBJID only if there is at least one exposure record.
Walkthrough: ADSL Merge Discipline
ADSL Merge Discipline: Walkthrough
code-level • no sample subject rows
| # | Step | Canonical ADSL move | Result / rule |
|---|---|---|---|
| 1 | EX: pre-collapse to first dose | sort EX by USUBJID, EXSTDTC; keep FIRST.USUBJID | TRT01SDT / TRT01STM from the first EX row |
| 2 | DS: latest disposition | sort DS by USUBJID, DSSTDTC; keep LAST.USUBJID | EOSDT and DSDECOD from the final record |
| 3 | DM-driven assembly merge | merge DM(in=a) with EX-first and DS-last by USUBJID; if a | one row per subject |
| 4 | Treatment variables | TRT01P from DM planned arm; TRT01A = EXTRT when an EX row exists | else TRT01A = 'Not Treated' |
| 5 | SAP analysis flags | ITTFL='Y' for randomized; SAFFL='Y' when a first EX dose exists | COMPFL='Y' if SAFFL and EOSDT >= TRT01SDT |
| 6 | Row-count QC habit | compare ADSL row count with DM row count on every run | counts must match |
Step through the canonical ADSL merge snippet line by line: first dose from EX, latest disposition from DS, one row per subject, and the row-count QC habit. No patient-level rows were supplied in this lesson's material facts, so the walkthrough is code-level only and no subject data is invented.
Speaker notes
Now we anchor the population and date concepts in real canonical SAS logic. Input A pre-collapses the Exposure domain, EX, to one row per subject: sort by USUBJID and EXSTDTC, keep FIRST.USUBJID, then build TRT01SDT from input(EXSTDTC, yymmdd10.) and TRT01STM from input(EXSTTC ?? time5., time5.). Input B keeps the latest disposition: sort DS by USUBJID and DSSTDTC, keep LAST.USUBJID, and carry EOSDT and DSDECOD from the final record. The assembly merge drives from DM with if DM, joins EX-first and DS-last by USUBJID, and preserves one row per subject. TRT01P takes the planned arm from DM, while TRT01A takes the actual first-dose treatment EXTRT when an EX row exists and otherwise falls back to 'Not Treated'. The analysis flags implement the SAP definitions: ITTFL is 'Y' for all randomized subjects, SAFFL is 'Y' only when a first EX dose exists, and COMPFL requires SAFFL plus EOSDT greater than or equal to TRT01SDT. Finally, the canonical QC habit checks that the ADSL row count equals the DM row count on every run.
Hands-on: Fill In the Merge Discipline
Hands-on interactive — if it does not load, open the paired article and try the exercise there.
Speaker notes
This one is hands-on on the website at jaimeyan dot com slash learn, so open it in a browser tab alongside the video. You will rebuild the duplicate-safe ADSL merge yourself: choose the collapse markers that guard the EX input with FIRST.USUBJID, the Unique Subject Identifier, and the DS input with LAST.USUBJID, assign each variable to its owner across DM, EX, and DS, then write the fallback branch and the flag rules in dependency order, with the Intent-to-Treat flag set for all randomized subjects and the Safety Analysis Set flag set for a documented first dose. Give it a try right after the video, and check the QC invariant that the final ADSL row count equals the DM row count.
Query Discipline, Specs, and the Agentic Way
Query Discipline, Specs, and the Agentic Way
01 SAP silent? Log + route — section, options, impact; no private picks.
02 Code-only decision = “invented” to an auditor — put it in the spec with a citation.
03 Worksheet is the contract — quote source lines, mark “not stated”, name an owner.
04 Agents = draft only; a confident paraphrase is not proof — verify every rule by its section number.
05 Tool layer is volatile — re-verify capabilities before relying on them.
Cover what to do when documents go silent: log and route rather than choose, keep the extraction worksheet as the contract, and treat agent-assisted extraction as a fast first pass that always requires human verification with citations.
Speaker notes
When the statistical analysis plan is silent on a rule that matters, do not guess — log the ambiguity. Record the section, the options, and the operational impact, then route it to the statistician instead of making a private choice. A decision living only in someone's code looks invented to an auditor, so the disposition belongs in the specification — for example, Define-XML, the data definition file — with a citation. The extraction worksheet is the deliverable: quote source lines, mark "not stated in document" instead of guessing a value, and name a human owner. Agents give a fast first pass, but verify like an auditor — ask every extracted rule for its section number, because a confident paraphrase is the failure mode. Treat the agentic tooling layer as volatile; capabilities and availability change, so re-verify before relying on specifics.
Closing Quiz: Extraction Discipline
1 During construction of ADSL, when must windowing and baseline rules be extracted, and what happens if they are extracted late?
2 Select all that are true. In a canonical ADSL merge, the subject identifier USUBJID is the unique subject-level key, TRT01SDT is the date of first exposure to treatment, EOSDT is the end-of-study date, and the ITT and SAF flags identify the intention-to-treat and safety-analysis populations. Which statements correctly describe the ownership of these ADSL variables? (select all that apply, then Check)
3 In regulatory submissions, define.xml is the metadata file that describes data sets, variables, and their derivations. Explain in one or two sentences why an auditor treats a decision that lives only in code as invented. (reflect, then reveal)
Reveal analysis
Speaker notes
This checkpoint quiz confirms integrated understanding across document reading, extraction order, and merge discipline. Question one: during construction of the subject-level analysis dataset (ADSL), windowing and baseline rules must be extracted, and the correct answer is A: they must be established during ADSL construction; if extracted late, downstream analysis datasets may re-implement the same rules independently, producing inconsistent logic that cannot be traced to one approved ADSL derivation. Question two: in the canonical ADSL merge, the unique subject identifier (USUBJID) is the subject-level key, treatment start date (TRT01SDT) is the date of first exposure to treatment, end-of-study date (EOSDT) is the end-of-study date, and the intention-to-treat (ITT) and safety-analysis (SAF) flags identify their populations; the correct answers are A, B, and C. In the Study Data Tabulation Model (SDTM), SDTM.EX owns TRT01SDT, SDTM.DS owns EOSDT, and ITT and SAF are owned by the statistical analysis plan (SAP), because ADSL derives them by applying SAP inclusion and exclusion rules to subject-level data. Question three: define dot x m l, the metadata file that describes data sets, variables, and derivations, is where planned decisions must appear; the model answer is that a decision that lives only in code is not represented in define.xml or in any approved specification, so the auditor cannot confirm it was planned under the SAP and treats it as invented.
From Pages to a Work List
From Pages to a Work List
| The line | Producer reading + ordered extraction = one work list: endpoints, populations, windowing, dates, query discipline. |
|---|---|
| Discipline | No private guesses: every rule becomes a dataset, flag, spec row, or logged query, each with a citation. |
| Material gap | No patient-level rows were supplied; the real-data pass used canonical ADSL snippet logic only — no subject-level data was invented. |
| Arc next | Bootcamp: SDTM domain basics → mapping-spec walkthrough → ADSL derivation worksheet (consumes this extraction). |
| At your desk | Protocol–SAP–Extraction Checklist: the eight-step extraction pass + query discipline, ready for the next study start-up. |
Loop back: source article + protocol/SAP reading method → reusable extraction kit.
Summarize the learning line and loop back to the source article and the wider bootcamp arc. Note the material gap honestly: because no patient-level extracts were supplied, the real-data pass was taught from the canonical ADSL snippet logic alone.
Speaker notes
This pass becomes practical here: producer reading plus the ordered extraction turns a protocol and a Statistical Analysis Plan (SAP) into one programmer's work list of endpoints, populations, windowing, dates, and query discipline. No private guesses: every rule must land as a dataset, a flag, a specification row, or a logged query, and each one carries a citation. One material gap matters: no patient-level rows were supplied, so the real-data walkthrough used only Subject-Level Analysis Dataset (ADSL) snippet logic, and invented no subject data. Next up in bootcamp: Study Data Tabulation Model (SDTM) domain basics, the mapping-specification walkthrough, and the ADSL derivation worksheet that consumes this extraction. Keep the Protocol–SAP–Extraction Checklist nearby: it packages the eight-step pass with query discipline for your next study start-up. Loop back to the source article and protocol/SAP reading method for a reusable extraction kit.