SCE & Modern Workflow All lessons Scene 1 / 12

SCE & Modern Workflow · Interactive Lesson

SDTM-ADaM-TLF Pipeline as Code

12 scenes· ~21 min· pairs with the article

Step through the scenes, pass the checkpoint quizzes, and try the hands-on exercises. Progress saves locally in this browser — no account, no tracking.

Scene index · 12 scenes
  1. ConceptTuesday 17:00: The DM Reissue
  2. ConceptThe Chain Is a DAG
  3. ConceptThree Traps of Memory-Order Runs
  4. ConceptPipeline as Code: Three Commitments
  5. Hands-onWatch a DM Correction Propagate
  6. ConceptPin the Environment
  7. ConceptHash Every Output
  8. ConceptWhere the Pipeline Lives in a Cloud SCE
  9. ConceptFailure Isolation
  10. ConceptThe Agentic Way
  11. CheckpointKnowledge Check
  12. ConceptKey Takeaways and Material Gap Note
Concept1 / 12

Tuesday 17:00: The DM Reissue

Tuesday 17:00: The DM Reissue

17:00 · Tue

● DM (SDTM demographics) reissued —

● 3 corrected records

Tue evening

● ADaM (Analysis Data Model) — rerun by hand in order

● TLF (Tables, Listings, Figures) — 2 tables rerun

● Listings left as-is: judged not to touch DM

Wed · shipped

● Wed: shipment contains two ADaM vintages

● Not noticed until the safety appendix disagrees

● with the demographics table

Goal: executable dependency graph, not memory order

Opening page: set the concrete work situation that motivates the course — a small upstream correction cascades because execution order lives in a human head.

Speaker notes

It's 17:00 on Tuesday, and Data Management reissues DM, the Study Data Tabulation Model (SDTM) demographics domain, with three corrected records. The Analysis Data Model (ADaM) programmer reruns the ADaM datasets by hand, in what they believe is the right order. The Tables, Listings, and Figures (TLF) programmer reruns two TLF tables, decides the listings do not touch DM, and goes home. Wednesday's package then mixes two vintages of ADaM data. Nobody notices until the safety appendix disagrees with the SDTM demographics table. The lesson goal is to replace memory-driven reruns with an executable dependency graph.

Concept2 / 12

The Chain Is a DAG

The Chain Is a DAG

Under the folder view, this is a DAG.

Raw data

source exports

SDTM

domain datasets

ADaM

analysis datasets

TLFs

final outputs

Named edges

ADaM reads its SDTM inputs

each TLF reads its ADaM data

Cascade

Change upstream invalidates

everything downstream of it

Real either way

The dependency graph holds

whether or not anyone writes it

Layer L1, fundamentals: establish the dependency graph underneath the familiar folder view.

Speaker notes

Raw data source exports feed the Study Data Tabulation Model, SDTM, domain datasets; SDTM then feeds the Analysis Data Model, ADaM, analysis datasets; ADaM then feeds the Tables, Listings, and Figures, TLF, final outputs. That sounds like a simple chain, but the folder view hides the real structure: it is a directed acyclic graph, or DAG. The graph has named edges: each ADaM dataset reads its specific SDTM inputs, and each TLF reads its specific ADaM data. Because of those edges, a change anywhere upstream creates a cascade that invalidates everything downstream of it. And this dependency structure is real either way: the graph holds whether or not anyone writes it down.

Concept3 / 12

Three Traps of Memory-Order Runs

Three Traps of Memory-Order Runs

Vocabulary: SDTM (standard tabulations) · ADaM (analysis datasets) · ADSL (subject-level analysis dataset) · TLF (tables, listings, figures)

1 · Missed downstream runs

2 · Order drift

3 · Untracked knowledge

• SDTM refresh lands;

• downstream TLF never re-runs.

• Folder-listing order ≠ dependency order;

• ADaM rebuilt from stale inputs.

• ADSL timing lives in one person’s head;

• absence stalls the study or silently misruns.

These traps — not data quality — are what pipeline-as-code removes.

Layer L1, fundamentals: show what happens when the DAG exists only as a run sheet, a macro list, or a veteran's head.

Speaker notes

Let's name the three traps that memory-order runs create, using our vocabulary: SDTM, or standard tabulations, ADaM, or analysis datasets, ADSL, the subject-level analysis dataset, and TLF, tables, listings, and figures. Trap one is missed downstream runs: an SDTM refresh lands, but a downstream TLF never re-runs. Trap two is order drift: programs run in folder-listing order, not dependency order, so ADaM gets rebuilt from stale inputs. Trap three is untracked knowledge: ADSL timing lives in one person's head, and their absence stalls the study or causes silent misruns. These traps, not data quality, are the failure mode that pipeline-as-code removes.

Concept4 / 12

Pipeline as Code: Three Commitments

Pipeline as Code: Three Commitments

1   Dependencies exist in a file next to the code

       not in a document or in memory

2   One entry script rebuilds any output

       make output/tlf/t_14_1_1.rtf

3   The runner computes order and rerun scope

       automatic when an upstream input changes

data/sdtm/dm.xpt data/sdtm/ae.xpt: programs/sdtm/*.sas raw/*.csv

    sas programs/sdtm/build_sdtm.sas

data/adam/adsl.xpt: data/sdtm/dm.xpt data/sdtm/ex.xpt programs/adam/adsl.sas

output/tlf/t_14_1_1.rtf: data/adam/adsl.xpt data/adam/adae.xpt

Layer L1, fundamentals: define pipeline-as-code and show the make-style rule file used throughout the source material.

Speaker notes

Now we make the pipeline a first-class artifact. Commitment one says dependencies live in a file next to the code, not in a document or in someone's memory. For Study Data Tabulation Model (SDTM), the rule might declare data/sdtm/dm.xpt and data/sdtm/ae.xpt from programs/sdtm/*.sas and raw/*.csv, then run sas programs/sdtm/build_sdtm.sas. For Analysis Data Model (ADaM), data/adam/adsl.xpt depends on data/sdtm/dm.xpt, data/sdtm/ex.xpt, and programs/adam/adsl.sas. That Subject-Level Analysis Dataset (ADSL) feeds a Tables, Listings, and Figures (TLF) output, output/tlf/t_14_1_1.rtf, which also needs data/adam/adae.xpt. Commitment two: one entry script rebuilds any output, like make output/tlf/t_14_1_1.rtf and nothing else, and commitment three: the runner in your Statistical Computing Environment (SCE), not the human, computes order and rerun scope after an upstream change.

Hands-on5 / 12

Watch a DM Correction Propagate

Hands-on interactive — if it does not load, open the paired article and try the exercise there.

Speaker notes

This exercise is hands-on, on the website at jaimeyan.com/learn, so you will do it in the browser rather than in the video player. There you push the DM reissued action and watch which graph nodes go stale: the Study Data Tabulation Model (SDTM) files data/sdtm/dm.xpt and data/sdtm/ae.xpt, the Analysis Data Model (ADaM) files data/adam/adsl.xpt for the Subject-Level Analysis Dataset (ADSL) and data/adam/adae.xpt, and the Tables, Listings, and Figures (TLF) output output/tlf/t_14_1_1.rtf, all driven by the dependencies declared in the Statistical Computing Environment (SCE) rule file. That is the skill worth practicing: the runner, not the human, decides which outputs rebuild and in what order, and a failed validation on data/adam/adae.xpt isolates the suspect downstream outputs while everything else keeps its evidence, so try it yourself right after this video.

Concept6 / 12

Pin the Environment

REPRODUCIBLE WORKFLOWS — LAYER L2

Pin the Environment

Every run executes on a declared compute stack.

The Rule

• Runs that differ by machine are not a pipeline.

• “Works in my workspace” is not an excuse.

Pin per Stage

• SAS — maintenance release

• R — lockfile via renv

• Python — pinned requirements.txt

The Payoff

• Three-year-old output replays —

manifest names its engine,

and that stack is stood up again.

SCE = Statistical Computing Environment — compute environment is configuration, per project and stage.

Layer L2, modern workflow: each run executes on a declared compute stack so reproducibility is built into the pipeline.

Speaker notes

This layer, L2, is about pinning the environment. A pipeline that runs differently on different machines is not a pipeline, and “works in my workspace” stops being a category of excuse. Pin per stage: the SAS maintenance release, R with an renv lockfile, and Python with pinned requirements.txt. In a cloud Statistical Computing Environment (SCE), compute environment is configuration, managed per project and often per pipeline stage. That is the payoff: a three-year-old output can be replayed because the manifest names its engine, and the SCE can stand that same stack up again.

Concept7 / 12

Hash Every Output

Hash Every Output

● Every stage writes a manifest: inputs + outputs → hashes.

● QC compares manifests — timestamps prove nothing.

● Input changed, output unchanged → skipped rerun = mismatch.

● Read bytes → SHA-256: XPT, SAS7BDAT, parquet.

ADaM=Analysis Data Model · ADSL=subject-level dataset

{ "stage": "adam/adsl" }

{ "inputs": ["dm.xpt", "ex.xpt"] }

{ "output": "adsl.xpt" }

{ "env": "sas 9.4m8 / r 4.4-renv / python 3.11" }

{ "written_to": "runs/adsl_2026-08-30.json" }

Layer L2, modern workflow: make rerun claims checkable with per-stage manifests.

Speaker notes

Every stage in your pipeline writes a manifest that records the hash of each input it reads and each output it produces. For example, the Analysis Data Model (ADaM) stage for the subject-level dataset (ADSL), called adam/adsl, reads dm.xpt and ex.xpt, produces adsl.xpt, and records its environment as sas 9.4m8 / r 4.4-renv / python 3.11, all written to runs/adsl_2026-08-30.json. During quality control, you compare manifests instead of trusting timestamps. If an input hash changed but the output hash did not, that is a mechanical mismatch, not a mystery. To get that hash, read the raw bytes of XPT, SAS7BDAT, or parquet files, compute a Secure Hash Algorithm 256-bit (SHA-256) digest, and store it at run time.

Concept8 / 12

Where the Pipeline Lives in a Cloud SCE

Where the Pipeline Lives in a Cloud SCE

1 · Version with code

Pipeline file + entry script stay with programs; PR-reviewed.

2 · Execute in the SCE

Launcher runs entry script in pinned env on study data.

3 · Replay from run history

Every run stamped; old run replayed from record.

Optional CI: commit triggers via repo job, or SCE scheduler alone.

Start simple: launch in SCE; add a CI server only when needed.

Layer L2, modern workflow: version the pipeline with study code and execute it in the SCE with run history.

Speaker notes

In this scene, we look at where the pipeline lives inside a Statistical Computing Environment, or SCE. The pipeline file and entry script stay in the same repository as your Study Data Tabulation Model, or SDTM, Analysis Data Model, or ADaM, Subject-Level Analysis Dataset, or ADSL, and Tables, Listings, and Figures, or TLF, programs. They go through pull requests like any other code, so review happens before anything runs. The SCE launcher runs the entry script in a pinned environment on study data; every run is stamped in run history, so you can replay an old run from its record. For commit-triggered rebuilds or cross-study dashboards, connect the repository to a continuous integration job, but the SCE scheduler often does the same without an external system. Start with the SCE launcher; add a continuous integration server only when needed.

Concept9 / 12

Failure Isolation

Failure Isolation

Validation check trips in ADAE — an ADaM analysis dataset — the DAG returns two lists:

SUSPECT — outputs downstream of the failed stage only

STANDS — everything else keeps its evidence

• Fix and rerun one stage — no cascade elsewhere

• Monolithic master macro reruns everything — extra runs risk unrelated drift

• Legacy validated macro libraries stay safe behind the wrapper layer

Layer L2, modern workflow: the operational payoff of a DAG-run pipeline on bad days.

Speaker notes

Now we reach failure isolation: when a validation check trips in the Adverse Events analysis dataset, ADAE, an Analysis Data Model, ADaM, dataset, the directed acyclic graph, or DAG, gives exactly two lists. The first list is suspect: only outputs downstream of the failed stage are in doubt. The second list stands: everything else keeps its evidence. So you fix and rerun that one stage, with no cascade elsewhere. A monolithic master macro would rerun everything, and every extra execution risks folding in unrelated drift. Legacy validated macro libraries can stay untouched behind a wrapper layer that the pipeline drives.

Concept10 / 12

The Agentic Way

The Agentic Way

Volatile

Layer L3 · Agentic workflow in the process-DAG

Last verified 2026-08-30

Agents work inside the process-DAG — never silently edit the graph.

• Fit plumbing: draft make-style rule files, diagnose failed runs, propose rerun scope

• Failure mode: invented certainty — adding a dependency edge or relaxing a tripped check

• Ground truth: the graph; every agent proposal is a diff for human merge

• Gate strictness never drops without a recorded decision

• Volatile layer: re-verify before relying on tool specifics

Layer L3, agentic way, flagged as volatile: agents should operate inside the process-DAG but never silently edit the graph.

Speaker notes

Now let's talk about the agentic way. Agents are useful for the plumbing: drafting make-style rule files, diagnosing failed runs, and proposing the smallest rerun scope. But their failure mode is invented certainty, where an agent might add a dependency edge or relax a tripped check with full confidence. Remember, the process directed acyclic graph is the study's ground truth, and every agent-proposed change is a diff that a human merges. And a gate never loses strictness without a recorded decision. This layer is volatile, last verified on August thirtieth, twenty twenty-six, so re-verify before relying on tool specifics.

Checkpoint11 / 12

Knowledge Check

1 Which statement best describes the dependency direction in a reporting pipeline that produces TLFs (Tables, Listings, and Figures) from SDTM (Study Data Tabulation Model) and ADaM (Analysis Data Model)?

2 Which statements distinguish a declared-dependency pipeline from a manual run sheet? Select all that apply. (select all that apply, then Check)

3 You must reproduce an old TLF after six months. The original manifest records an unchanged SDTM source hash, ADaM code commit main@f7e2c1, SCE (Statistical Computing Environment) R 4.3.1, and output hash d5f9a41. Current repo state is ADaM code commit main@55a0d4 and SCE R 4.5.0. Which rerun plan correctly applies environment pinning and manifest reasoning?

4 Describe how failure isolation and human-reviewed agent diffs protect study outputs in a rerun after a metadata change. Include what you would do if the isolated ADaM build fails after an agent diff has been approved and before the TLF node runs. (reflect, then reveal)

Reveal analysis
Reference answer: isolation means build_adam can fail without automatically cascading into render_tlf; you hold the TLF until the ADaM output succeeds, and the old approved TLF is not replaced by a file generated from failed ADaM. Human review is a control over the agent's proposed code changes: it ensures the diff matches the protocol amendment and analysis decisions, and creates an auditable approval trail. In this scenario, do not render the TLF; investigate the ADaM failure, rerun the isolated build, and regenerate only after the output passes validation.
Speaker notes

This checkpoint checks your understanding of the Study Data Tabulation Model, or SDTM, the Analysis Data Model, or ADaM, the Subject-Level Analysis Dataset, or ADSL, Tables, Listings, and Figures, or TLF, and the Statistical Computing Environment, or SCE. In the first question, the dependency direction in a pipeline that produces TLFs from SDTM and ADaM is best described by option B: ADaM analysis datasets are derived from relevant SDTM domains and ADSL, and each TLF should consume the specific ADaM datasets it needs, so the flow is SDTM, then ADaM, then TLF. For the multiple-select question, the correct choices are A and B: a declared-dependency pipeline stores relationships such as 'TLF-X depends on ADaM-Y' in machine-readable form so a rerun planner can identify stale downstream work, while a manual run sheet records the order a programmer executed but is not a code-level declaration the pipeline can check. To reproduce an old TLF after six months when the original manifest records an unchanged SDTM source hash, ADaM code commit main@f7e2c1, SCE R 4.3.1, and output hash d5f9a41, while current state is ADaM code commit main@55a0d4 and SCE R 4.5.0, the correct plan is B: rebuild ADaM using the pinned ADaM commit main@f7e2c1 and R 4.3.1, regenerate the TLF, and compare the new output hash to d5f9a41. An unchanged SDTM hash is necessary but not sufficient because the ADaM code and SCE have changed. For the short-answer question, isolation means build_adam can fail without automatically cascading into render_tlf, so you hold the TLF until the ADaM output succeeds and the old approved TLF is not replaced by a file generated from failed ADaM; human review controls the agent's proposed code changes by ensuring the diff matches the protocol amendment and analysis decisions and creates an auditable approval record. If the isolated ADaM build fails after an agent diff has been approved and before the TLF node runs, stop and resolve the ADaM failure, re-review the diff if needed, and do not render the TLF from failed ADaM.

Concept12 / 12

Key Takeaways and Material Gap Note

Key Takeaways & Gap Note

● SDTM → ADaM → TLF is a DAG; rerun from file, not memory

● Pipeline-as-code = declared deps + entry script; runner computes order & rerun scope

● Pin env per stage; hash I/O to per-stage manifests; skipped reruns become manifest mismatches

● Isolation payoff: failed stage reruns alone; monolithic master macro reruns the whole world

Material gap: no patient-level rows were supplied in the source

No simulated patient data page was drawn; pipeline file facts only.

Summary slide linking back to the pipeline-as-code article and noting the real-data gap in this session.

Speaker notes

Let's close with the picture. The chain from Study Data Tabulation Model (SDTM) to Analysis Data Model (ADaM) to tables, listings, and figures (TLF) is a directed acyclic graph (DAG), and most rerun disasters happen when we execute it from memory instead of from a file. Pipeline as code means declared dependencies plus one entry script, with a runner that computes the order and the rerun scope. Pin the environment per stage and hash inputs and outputs into per-stage manifests, so a skipped rerun shows up as a mechanical manifest mismatch rather than a silent surprise. That isolation is the payoff: one failed stage reruns alone, while the old monolithic master macro reruns the world. One honest note: the source material supplied no patient-level rows, so we drew no simulated patient data page and used only pipeline file facts.

✓

Lesson complete

Nice work — every scene seen. Keep the momentum going.

← → Space to navigate · progress is saved locally in your browser

AI Tutor

Ask the tutor