Opening: The Reviewer's Question
LESSON OPENING · SETTING THE SCENE
The Reviewer’s Question
“If a reviewer asks how this table was
produced, what exactly do we show them?”
— Sponsor QA lead, kickoff meeting
• Capability: settled. Task ahead: governance.
• Regulated AI adoption: assurance, not capability.
• Draft answer is not yet audit-shaped.
• Roadmap: validated → GAMP 5/CSA → SCE
draft–gate–sign → Monday policy.
Opening page (L1 / framing). Place the learner in the concrete work situation the article opens with: a sponsor QA lead asking what evidence sits behind an AI-assisted table.
Speaker notes
In the kickoff meeting, the sponsor's quality assurance lead asked: if a reviewer asks how this table was produced, what exactly do we show them? That question is about assurance, not capability, and it sits inside the Good Practice, or GxP, boundary. The honest draft answer—an assistant drafted the program, a programmer touched it up, quality control passed—is not yet shaped like an answer an auditor accepts. So today we close that governance gap. We will define validated, walk through Good Automated Manufacturing Practice, or GAMP 5, and Computer Software Assurance, or CSA, risk-based assurance, then see the System of Concern, or SCE, draft-gate-sign pattern, and finish with a three-line policy you can adopt on Monday.
What "Validated" Means in Practice
What “Validated” Means in Practice
Computer system validation (CSV) is documented evidence that a system does what it is supposed to do — reliably, and under control.
Five duties are stable across decades: requirements written down, testing against those requirements, defect handling, change control, and reconstructable state (version · config · data · who touched).
• Validated environment: the duties above are met for systems that produce regulated content.
• GAMP (Good Automated Manufacturing Practice): grades systems by risk and customisation → proportional assurance — a spreadsheet template is not a bespoke clinical database.
Concept page (L1). Establish the durable fundamentals of computer system validation before any AI discussion.
Speaker notes
Computer system validation, or CSV, is documented evidence that a system does what it is supposed to do, reliably and under control. In practice, five duties stay stable across decades: requirements written down, testing against those requirements, defect handling, change control, and a reconstructable state covering version, configuration, data, and who touched each. A validated environment is where those duties are met for the systems that produce regulated content. The Good Automated Manufacturing Practice, or GAMP, framework gives us a vocabulary for grading systems by risk and by how configurable or bespoke they are. The core principle is proportionality: a spreadsheet template and a custom clinical database do not warrant the same assurance dossier.
Risk-Based Assurance: GAMP 5 Second Edition + CSA
Risk-Based Assurance
GAMP 5 Second Edition + CSA
Regulatory Framing
Ask What Can Go Wrong
• GAMP 5 (2nd ed.) + FDA CSA
• Assurance effort scales with risk
• Risk = patient safety & product quality
• Lower-risk features → exploratory tests
• Record and justify what you did
• Category error: “Is the LLM validated?”
• Ask: “What can go wrong if this component misbehaves?”
• Evidence we need that it did not
• Statistical programming teams can answer
• Criterion by criterion, output by output
Concept page (L1). Show how current regulatory thinking frames assurance for AI-assisted work.
Speaker notes
Good Automated Manufacturing Practice (GAMP) 5 second edition and the U.S. Food and Drug Administration (FDA) computer software assurance (CSA) direction make proportionality explicit. Assurance effort scales with risk to patient safety and product quality. For lower-risk features, CSA pushes unscripted and exploratory testing, with the rationale documented, so you test less formulaically but record and justify what you did. Asking whether the large language model (LLM) is validated is a category error. The question that works is, what can go wrong if this component misbehaves, and what evidence do we need that it did not? A statistical programming team can answer that criterion by criterion, output by output.
The Assistant/Artifact Split and the Accountability Boundary
The Assistant/Artifact Split
and the Accountability Boundary
Validate the artifact and process — never the tool.
LLMs fail quietly — assurance at deterministic gates.
| Component | Role | Deterministic? | How assured |
|---|---|---|---|
| LLM assistant | Generates draft code from prompts | No — output varies | Scoped out of the regulated record |
| Generated program | Artifact to submit (SAS, dataset, TLF) | Yes — fixed when saved | Human review; validated-environment run |
| Pipeline & gates | Re-runnable QC and verification steps | Yes — by design | Process validation + audit logs |
| Run record | Evidence of what actually ran | Yes | Retained and reconciled |
| Human sign-off | Accountability for release | No — human judgment | Named signature after evidence review |
Fixed boundary: named human sign-off per artifact
Concept page (L1). Introduce the distinction that resolves most of the panic, with the article's component table, and the human signature boundary.
Speaker notes
After a large language model drafts a SAS program, two things exist. The assistant is a nondeterministic tool in the means of production, and the artifact is the program, dataset, or TLF that enters the regulated record. You validate the artifact and the process that produces it, never the tool itself. Unlike compilers, which fail loudly and deterministically, LLMs fail quietly and plausibly, so the assurance weight shifts onto deterministic gates and independent human review. The table walks through the split: the LLM assistant, the generated program, the pipeline and gates, the run record, and the human sign-off. The boundary that does not move is this: every artifact has a named person behind it, who reviewed the evidence and signed the validation record, and no capability advance moves that signature.
Knowledge Check: Validation Concepts
1 Which statement best defines the validation of a regulated statistical computing environment (SCE)?
2 In the validated statistical computing environment (SCE), which statements correctly apply the assistant/artifact distinction when a large language model (LLM) supports work that produces tables, listings, and figures (TLF)? Select all that apply. (select all that apply, then Check)
3 A programmer says, "Before we let the LLM help with programming, we need to validate the LLM." Which response best rejects the category error in that statement?
Speaker notes
Time for a knowledge check on the validated statistical computing environment (SCE) and the assistant/artifact split. For validation of a regulated SCE, the best definition is option B: collecting documented evidence that the environment consistently meets its predetermined requirements through testing, change control, and a reconstructable state. That's because under computer software assurance (CSA) expectations, validation is broader than a one-time run or a vendor certificate. For the validated SCE, when a large language model (LLM) supports work that produces tables, listings, and figures (TLF), the correct answers are A and C. The SCE's accountability boundary separates deterministic artifacts—like a final TLF or a quality control (QC) passed SAS program—from assistant output, so the LLM's suggestions do not become regulated artifacts just because they influenced code. The correct answer is C: the validation obligation belongs to the SCE workflow—including the deterministic gate and human accountability boundary—while the LLM is an assistant, so 'validate the LLM' is a category error.
Where AI Sits in SCE Platforms
L2 · MODERN WORKFLOW
Where AI Sits in SCE Platforms
Domino-class SCEs — the governed cloud platforms in this series —
deliver AI as a platform feature — not a side channel.
01 · Governed assistant access
• Study data stays inside the boundary
• No default hop to external endpoints
02 · Traceable AI contributions
• Session tied to project + user
• Diffs capture prompt → suggested → accepted
• Reviewed like any other change
03 · Auditable sessions
• Prompts + outputs captured
• Same audit trail as runs
• QA can inspect after the fact
Volatile layer: category description — not product certification.
Verify what your actual platform enforces, logs, and configures.
Modern-workflow page (L2). Describe the current SCE category from the article, with the volatility caveat.
Speaker notes
In a Statistical Computing Environment — an SCE — like Domino, AI shows up as a platform feature, not a side channel. The first capability is governed assistant access: study data stays inside the boundary, with no default hop to external endpoints. Second, traceable AI contributions: each session ties to a project and a user. Every diff records who prompted, what was suggested, and what was accepted — then it is reviewed like any other change. Third, auditable sessions: prompts and outputs land in the same audit trail as runs, so Quality Assurance can inspect them after the fact. This is a category description, not product certification — so verify what your actual platform enforces, logs, and configures. That diligence is exactly what a Computer Software Assurance — CSA — minded team should run inside the Good Practice — GxP — boundary.
The Pattern: Draft, Gate, Review, Sign
The Pattern: Draft, Gate, Review, Sign
Unchanged since drafting tools existed — enforced now by a statistical computing environment (SCE).
Draft
Gate
Review
Sign
• Assistant proposes code / fix inside the platform
• Prompt + proposal recorded with the session
• Deterministic checks only: compile, assert, QC compare
• No model sits inside the gate
• Human reads the diff with gate evidence attached
• Approve or reject in the pull request
• Validation record names the owner
• Archive: prompt, output, model version, gate log
Minimal Study XYZ gate check
Compare /adam/adsl.xpt (primary ADaM) with /qc/qc_adsl.xpt (independent QC)
assert equal shapes → “structure mismatch”; assert identical sorted USUBJID → “subject mismatch”
print “gate PASS — independent QC agrees”
Modern-workflow walkthrough page (L2). Step through the article's end-to-end pattern and its exact minimal gate check — no invented numbers or code.
Speaker notes
The working pattern remains draft, gate, review, sign, and a statistical computing environment, or SCE, makes every step enforceable. In the draft step, the assistant proposes code or a fix inside the platform, and the session records the prompt and the proposal. The gate is deterministic only: compilation, assertions, and independent quality control comparison, with no model inside. For the minimal Study XYZ check, read the primary ADaM output /adam/adsl.xpt and the independent quality control output /qc/qc_adsl.xpt, assert equal shapes for a structure mismatch, assert identical sorted USUBJID for a subject mismatch, then print that the gate passes when independent quality control agrees. Then a human reads the diff with gate evidence attached, approves or rejects in the pull request, and the validation record names the owner with the archived prompt, raw output, model version, and gate log.
Reproducibility and the Three-Line Policy
Reproducibility & Three-Line Policy
• Versions drift — identical calls can vary
• Temperature 0 is greedy sampling, not stability
• Replayability is pipeline property, not sampler
• Hash and version everything the model touched
Policy at a Glance
| # | Policy line | Assurance effect |
|---|---|---|
| 1 | Explore freely on synthetic and public data | Sandbox risk → 0 |
| 2 | Study data: approved assistants only, logged, in-platform | Data risk contained |
| 3 | Humans own the artifacts | Accountability fixed |
Modern-workflow page (L2). Connect reproducibility outside the model to a practical, adoptable governance policy.
Speaker notes
Inside the Good Practice, or GxP, boundary, sponsors often ask whether it will produce the same thing next time, and for a hosted large language model, or LLM, the honest answer is no. Not across model versions, and sometimes not even across identical calls. Temperature zero buys you greedy sampling, not stability, because replayability is a property of the pipeline, never of the sampler. So the architectural consequence is clear: hash and version everything the model touched, so drift is detected rather than silently absorbed. The three-line policy keeps this practical: explore freely on synthetic and public data; on study data, use approved assistants only, inside the platform, logged; and humans own the artifacts. This maps to assurance logic: line one drives sandbox risk toward zero, line two contains data risk, and line three fixes accountability.
Hands-On: Build the Validation Record
Hands-on interactive — if it does not load, open the paired article and try the exercise there.
Speaker notes
This exercise is hands-on on the website at jaimeyan.com/learn, so keep a browser tab open and work through it yourself. You will practice building a reviewer-ready evidence trail by sorting inspection-evidence cards: which items belong inside the validation record, and which stay outside. You will answer the reviewer's four process questions, learn to spot the trap cards, and rehearse the assistant and artifact split, so give it a try right after this video.
The Agentic Way
The Agentic Way
Time-sensitive layer
as-of 2026-08-30
Draft
Gate
Sign
Shift — agents run draft → gate → sign end-to-end
Safe — fixed DAG, deterministic gates, artifacts
Risk — unreviewed momentum; actions outrun reviews
Guard — audit events + attribution; named human
The agentic era makes the environment, not the model,
the decisive choice.
Agentic-layer page (L3), flagged as time-sensitive with the article's volatile-asOf framing.
Speaker notes
Now we shift to the agentic way, where agents execute the whole draft, gate, and sign loop. They rerun pipelines, triage failed runs, propose fixes, and open pull requests; inside a fixed directed acyclic graph, or DAG, with deterministic gates, this is tractable because every agent action lands on a reviewable artifact, and the gates do not care who produced the diff. The new failure mode is unreviewed momentum, where a chain of plausible agent actions can move faster than your review cadence. The countermeasures are a fixed workflow, deterministic gates, agent actions logged as first-class audit events with attribution, and a human named on the record. The agentic era makes the environment, not the model, the decisive choice; this layer is time-sensitive, last verified on August 30, 2026, so re-verify before relying on tool specifics.
Closing Quiz: Assurance Under Control
1 A clinical programmer wants to draft SAS code for a regulatory TLF using a large language model (LLM). Which workflow is consistent with the three-line policy in a validated statistical computing environment?
2 An inspector asks how an AI-assisted TLF was controlled. Under a computer software assurance (CSA) approach, which evidence items would directly answer the inspector’s process questions? Select all that apply. (select all that apply, then Check)
3 A study statistician asks, “Is the LLM validated? If we set the model temperature to zero, is the pipeline reproducible?” Write the model answer a clinical programmer should give. In your response, (a) state whether the LLM carries validation status, (b) explain where assurance actually lands, and (c) describe what reproducibility depends on, using the ideas of hashing/versioning, the deterministic review gate, and the accountability boundary. (reflect, then reveal)
Reveal analysis
Speaker notes
This checkpoint quiz has three questions, and each one tests how you apply the lesson's assurance reasoning to a realistic work situation. First, a clinical programmer wants to draft SAS code for a regulatory TLF using a large language model, and you must pick the workflow that fits the three-line policy in a validated statistical computing environment. The answer is B: use the approved enterprise LLM assistant on synthetic data to create a draft program, then run that draft through the usual deterministic review and validation, because the approved assistant keeps a recognized accountability boundary and the output is still treated as a draft under the same deterministic gate. Second, an inspector asks how an AI-assisted TLF was controlled, and under a computer software assurance approach you select every evidence item that directly answers the inspector's process questions. The correct selections are A, B, C, and D: the diff between the LLM-generated draft and the final SAS program, the gate log with the human approval decision and date, the archived prompt, model version, and configuration settings, and the named reviewer who accepts accountability for the final artifact with evidence of review and sign-off, because together they document the AI-assisted generation, the human control point, and the accountability boundary. Third, a study statistician asks whether the LLM is validated and whether setting the model temperature to zero makes the pipeline reproducible. The model answer is that the LLM is not validated and no single setting can make it a validated component of the analysis, because assurance lands in the controlled workflow around the model, where the approved assistant drafts only, the prompt, version, and configuration are archived, the draft is reviewed and edited in the version-controlled SAS environment, and reproducibility comes from hashing and versioning everything the model touched plus the deterministic review gate and the named accountability boundary.
Summary: What You Show the Reviewer
Summary: What You Show the Reviewer
Answer as evidence — process and records, not demos.
● Evidence — process and records, not demos.
● Assurance — GAMP 5 (2nd Ed.) + CSA, risk-scaled.
● Governance — three-line policy, human sign-off.
● Run record — diff, gate log, prompt/model, sign-off.
▍ Deepen — read the canonical article “AI in Validated Environments” · series #14 · SCE (Statistical Computing Environment) & Modern Workflow.
▍ Material note — this source has no patient dataset or SAS extracts; only the Study XYZ gate snippet was shown — nothing fabricated.
Summary slide. Reconnect to the paired article and note the material gap: this lesson's source contains no patient dataset or SAS row extracts, so no patient data was shown and no row-level walkthrough was attempted.
Speaker notes
Let's consolidate what you show the reviewer. In Good Practice (GxP) terms, the sponsor's question is evidentiary: answer with process and records, not demos, because acceptance moved from whether to under what controls. Good Automated Manufacturing Practice (GAMP) 5, second edition, and Computer Software Assurance (CSA) support risk-based assurance, scaling validation to risk; validating the LLM puts the obligation on the wrong object. Three-line policy is minimum viable governance: explore on synthetic data, use approved assistants on study data, let humans own the artifacts. The reviewer answer is the run record: diff, gate log, archived prompt and model version, named human sign-off. Read the article at /blog/ai-in-validated-environments.html, series 14 on the Statistical Computing Environment (SCE) and Modern Workflow line; this source has no patient dataset or SAS extracts, so the only artifact shown was the Study XYZ gate snippet.