Statistics All lessons Scene 1 / 12

Statistics · Interactive Lesson

P-Values & Clinical SAS Tests

12 scenes· ~21 min· pairs with the article

Step through the scenes, pass the checkpoint quizzes, and try the hands-on exercises. Progress saves locally in this browser — no account, no tracking.

Scene index · 12 scenes
  1. ConceptThe SAP Says Wilcoxon. Why?
  2. ConceptThe p-value, In One Honest Paragraph
  3. ConceptWhy Every Estimate Gets Its CI
  4. Conceptt-test or Wilcoxon: Read the Switch
  5. ConceptChi-square or Fisher: Watch the Expected Cells
  6. ConceptMultiplicity in Plain Words
  7. ConceptANCOVA and Why Longitudinal SAPs Moved to MMRM
  8. ConceptLog-rank, Tied Straight to ADTTE
  9. ConceptShift Tables and the CMH Row Mean Scores Engine
  10. Hands-onQC Craft Drill: Protect the P-Value Cell
  11. CheckpointFinal Knowledge Check: Explain, Program, Prove
  12. ConceptKey Takeaways: The Same Week, Every Week
Concept1 / 12

The SAP Says Wilcoxon. Why?

L1 · OPENING SCENE — STATISTICAL LITERACY FOR CLINICAL PROGRAMMERS

The SAP Says Wilcoxon. Why?

The desk reality

• SAP: Wilcoxon rank-sum — primary endpoint

• Clean merge ≠ ready to ship the table

• Explain · program · QC every output

• Not a stats exam: read SAP → see intent

Course route

1 · p-value + confidence interval (CI)

2 · test families: Wilcoxon rank-sum, Fisher's exact, MMRM, log-rank, CMH

3 · multiplicity

4 · QC craft behind every cell

Programmer rule · You do not choose the test — you explain, program, and prove the output right.

Opening scene (L1). Puts the learner in the real clinical programming situation the source article opens with: reading a SAP that specifies Wilcoxon for a primary endpoint and being expected to explain it, then QC the table that follows.

Speaker notes

Picture the interview desk: the Statistical Analysis Plan, or SAP, names Wilcoxon rank-sum for the primary endpoint, and a candidate who can merge datasets cleanly still freezes. That moment is not a statistics exam. It tests whether you can read the SAP, recognise what the statistician chose and roughly why, then quality control, or QC, the table that comes out of it. As programmers, we do not choose the tests; we explain them, program them, and prove the outputs are right. Over this course we walk the route: the p-value and its confidence interval, or CI, pairing, then each test family in turn — Wilcoxon rank-sum, Fisher's exact, mixed model for repeated measures, or MMRM, log-rank, and Cochran-Mantel-Haenszel, or CMH, row mean scores — and finally multiplicity and the QC craft behind every cell.

Concept2 / 12

The p-value, In One Honest Paragraph

The p-value, in one honest paragraph

Working definition: the probability of observing data at least as extreme as what was observed, computed under a model in which the null hypothesis holds.

Small p-value: awkward with the null. Large p-value: no awkwardness. That is the entire claim.

✗ Not the probability that the null is true

✗ Not the probability of a chance result

✗ Not a measure of effect size or importance

Programmer duty: right test · right records · right N · table transcribed without damage

Concept page (L1). Gives the working definition of a p-value exactly as the material states it, then removes the three common over-readings and reframes the p-value as one protected TLF cell.

Speaker notes

Let's make the p-value honest and plain. A p-value is the probability of seeing data at least as extreme as what was observed, assuming a model where the null hypothesis holds. So when the p-value is small, the data sit awkwardly with that null; when it is large, they do not. That is the whole claim, and we should not stretch it. Please remember what a p-value is not: it is not the probability the null is true, not the probability the result came from chance, and not a measure of effect size or importance. Your programmer duty is practical: run the right test on the right records with the right N, then transcribe the result into the table without damage.

Concept3 / 12

Why Every Estimate Gets Its CI

Why Every Estimate Gets Its CI

1 · What the CI says

The confidence interval (CI) is the range of

parameter values consistent with the data,

under the model: 95% default in the SAP.

2 · Signal vs. precision

Estimate: where the signal points.

CI: how precisely it points.

Precision = separate axis from signal.

3 · Why p-value alone fails

A tiny p-value can ride on an interval

so wide that no clinically meaningful

conclusion survives. The pair shows it.

4 · What QC must check

Check the CI with the estimate:

same population, same N as the estimate;

CI centered on the exact estimate shown.

Concept page (L1). Explains why TLFs pair every estimate with its confidence interval and why QC checks the pair against a single population and N.

Speaker notes

An estimate on its own only tells you where the signal points. The confidence interval, or CI, tells you how precisely that estimate points, and it is a separate axis from the signal itself. At the level stated in the statistical analysis plan, the SAP, usually 95 percent, the CI gives the range of parameter values compatible with the observed data under the model. That matters because a tiny p-value can sit on an interval so wide that no clinically meaningful conclusion survives it, and the estimate-plus-CI pair is what lets a reviewer see that. In quality control, or QC, you verify the interval together with the estimate: same population, same sample size, N, and the interval built around the exact estimate printed beside it.

Concept4 / 12

t-test or Wilcoxon: Read the Switch

t-test or Wilcoxon: Read the Switch

t-test

Compares two group means

Assumes well-behaved data

Wilcoxon rank-sum

Compares two distributions

Asks less of the data

SAP (Statistical Analysis Plan) switch triggers:

Skewed endpoint (lab deltas, pain scores) · ordinal

Small samples — no theory rescues the mean

Answer: comparisons + switch trigger — no formula

SAS: PROC TTEST · PROC NPAR1WAY WILCOXON

Concept page (L1). Covers the first two-group family: what the t-test and Wilcoxon rank-sum test compare, when SAPs switch, and which procedures run them.

Speaker notes

On this slide, we are choosing between the t-test and the Wilcoxon rank-sum test. The t-test compares two group means, and it assumes the data behave well enough for that comparison to be meaningful. The Wilcoxon rank-sum test compares two distributions using ranks, so it asks far less of the data. Your Statistical Analysis Plan, or SAP, switches to Wilcoxon when the endpoint is skewed, like lab deltas or pain scores, when it is ordinal, or when samples are small enough that no theory rescues the mean. In an interview, give two parts: name what each test compares, then name the condition that made the statistician switch; no one wants a formula. In SAS, which stands for Statistical Analysis System, use PROC TTEST for the t-test and PROC NPAR1WAY, the nonparametric one-way procedure, with the WILCOXON option for Wilcoxon.

Concept5 / 12

Chi-square or Fisher: Watch the Expected Cells

Chi-Square or Fisher: Watch the Expected Cells

Any expected cell below 5 → use Fisher's exact test

Chi-square — asymptotic

• p approximated from asymptotic theory

• Approximation weakens as counts shrink

• PROC FREQ: CHISQ option requests it

Fisher's exact — direct p

• p computed exactly from the table

• Use when any expected cell is below 5

• EXACT statement: cost grows as N climbs

PROC FREQ output lists both tests — no automatic match to the SAP.

Checkpoint: report the SAP-named test · verify expected cells were checked.

Concept page (L1). Explains the expected-cell rule that decides chi-square versus Fisher's exact test and the PROC FREQ trap where both tests appear in the output.

Speaker notes

Now you are looking at the choice between a chi-square test and Fisher's exact test, and the thing to watch is the expected cell counts. Chi-square approximates the null distribution of the discrepancy between observed and expected counts, and that approximation degrades when counts get small. The working rule is simple: when any expected cell count falls below about five, use Fisher's exact test, which computes the p-value directly instead of approximating it. In PROC FREQ, the CHISQ option gives the asymptotic tests, while the EXACT statement requests exact computation, and that cost grows as the sample size climbs. Because the output contains both tests, your duty is to confirm the table reports the one your statistical analysis plan, or SAP, names. And do not skip the checkpoint: verify that expected cells were checked before you trust the result.

Concept6 / 12

Multiplicity in Plain Words

Multiplicity in Plain Words

• Every 0.05-level test under the null carries a 5% chance of a false alarm.

• Run a family of tests — those chances compound into real risk of at least one spurious win.

If an interviewer asks what multiplicity means — the two sentences above are the whole answer.

Statistical Analysis Plan (SAP) pre-specifies:

• Testing hierarchy

• Alpha split: endpoints / interim looks

• Adjusted p-values and confidence intervals

Programmer duty: transcribe with fidelity

• Reproduce adjusted values the SAP names

• Footnote the method the SAP names

• Adjust nothing on your own initiative

Concept page (L1). Gives the two-sentence interview answer for multiplicity, then states the programmer's fidelity duty: reproduce what the SAP named and adjust nothing on initiative.

Speaker notes

Your Statistical Analysis Plan, or SAP, pre-specifies how to handle multiplicity. Every test at the 0.05 level carries a five percent chance of a false alarm under the null. Run a family of tests, and those chances compound into a real probability of at least one spurious win. So the SAP responds before the fact with a testing hierarchy, an alpha split across endpoints or interim looks, and adjusted p-values and confidence intervals. Your job is transcription with fidelity: reproduce the adjusted values the SAP names, footnote the method the SAP names, and adjust nothing on your own initiative. If an interviewer asks what multiplicity means, those two sentences about false alarms and compounding are the whole answer.

Concept7 / 12

ANCOVA and Why Longitudinal SAPs Moved to MMRM

ANCOVA and Why Longitudinal SAPs Moved to MMRM

ANOVA

• Compares means, 2+ groups

• One time point analysis

• No covariate adjustment

ANCOVA

• Adds covariates to ANOVA

• Classic covariate: baseline of endpoint

• Trims noise, corrects imbalance

LOCF (legacy)

• Last observation carried forward

• Imputes missing visits

• Analyzes the filled-in data

• Bias when dropout is informative

MMRM (current)

• Mixed model for repeated measures

• All observed values — no imputation

• Assumes missing at random (MAR)

• Why longitudinal SAPs adopted it

In SAS: PROC MIXED with REPEATED & unstructured covariance · confirm fit by denominator df

Concept page (L1 core with L2 workflow context). Explains ANOVA, ANCOVA, and the LOCF-to-MMRM migration that now defines longitudinal analysis in clinical reporting.

Speaker notes

ANOVA, analysis of variance, compares means across more than two groups at one time point, with no covariate adjustment. ANCOVA, analysis of covariance, adds covariates to ANOVA, classically the baseline value of the endpoint, which trims noise and corrects baseline imbalance. The older longitudinal approach, LOCF, last observation carried forward, imputed missing visits and analyzed the filled-in data, but that manufactures flat trajectories and biases results when dropout is informative. MMRM, the mixed model for repeated measures, analyzes all observed values under a missing-at-random assumption, MAR, with no imputation, which is why longitudinal statistical analysis plans, SAPs, moved to it. In SAS, the Statistical Analysis System, MMRM is PROC MIXED with a repeated effect and an unstructured covariance; the denominator degrees of freedom in the output confirm the model was fit as specified.

Concept8 / 12

Log-rank, Tied Straight to ADTTE

Log-rank, Tied Straight to ADTTE

ADTTE = time-to-event analysis dataset · CNSR = censor variable · KM = Kaplan-Meier · QC = quality control

1 · Log-Rank Test

• Compare survival over time

• Uses each event time

• Respects censoring

2 · ADTTE Variables

• AVAL = duration

• CNSR: 0 = event, 1 = censored

• KM plot pairs with test

3 · SAS Implementation

• PROC LIFETEST;

• STRATA treatment;

• Event counts in output

QC tie exact: log-rank event count = ADTTE records with CNSR = 0

Censored count reconciles the same way — mismatch with ADTTE is wrong before statistics

Concept page (L1/L2). Connects the log-rank test to the programmer's own ADTTE dataset and to the event-count reconciliation that makes the QC pass exact.

Speaker notes

Now we connect the log-rank test straight to the ADTTE, the time-to-event analysis dataset, which compares survival experience between groups across the whole follow-up time. It uses each event time and respects censoring. In ADTTE, AVAL carries the durations, and the censor variable, CNSR, tells the procedure which records are events: zero means event, one means censored. In SAS, you run PROC LIFETEST with a STRATA statement, and the Kaplan-Meier, KM, plot pairs with the test. The quality control, QC, tie is exact: the log-rank event count must equal the number of records where CNSR equals zero in the analysis population, and the censored count must reconcile the same way. If the log-rank table disagrees with its own ADTTE, it is wrong before any statistics get checked.

Concept9 / 12

Shift Tables and the CMH Row Mean Scores Engine

Shift Tables and the CMH Row Mean Scores Engine

What is a shift table?

• Cross-tabs baseline result × post-baseline result

• Tracks how each subject moves between categories

• Asks if shifts differ by treatment arm

Row mean scores statistic

• Treats ordinal categories as scores

• Tests average score shift across arms

• Stratifies by site or analysis pool

Enabling it in PROC FREQ

• Strata listed first in TABLES request

• CMH option enabled in TABLES statement

• Row mean scores section appears in output

Two QC checks

• Matches row mean scores statistic in shell

• Score order follows SAP categories

• Built on ADLB and ADVS BDS outputs

Concept page (L2). Explains the Cochran-Mantel-Haenszel row mean scores statistic that powers shift tables and the two checks that protect its p-value.

Speaker notes

A shift table crosses baseline category against post-baseline category, tracking each subject's movement and asking whether those shifts differ by treatment arm. The Cochran-Mantel-Haenszel (CMH) row mean scores statistic treats ordinal categories as scores, then tests whether the average score shifts across arms, stratified by site or analysis pool. In PROC FREQ, you enable this by listing the strata variables first in the TABLES request and setting the CMH options, so the row mean scores section appears in output. For quality control, confirm the table prints the specific CMH statistic the shell specifies, and confirm the score assignment follows the Statistical Analysis Plan (SAP) category ordering. These tables are built from Basic Data Structure (BDS) outputs such as the Laboratory Analysis Dataset (ADLB) and Vital Signs Analysis Dataset (ADVS).

Hands-on10 / 12

QC Craft Drill: Protect the P-Value Cell

Hands-on interactive — if it does not load, open the paired article and try the exercise there.

Speaker notes

This scene is a hands-on quality control drill, so you run it in the browser at jaimeyan.com/learn rather than watching it in the video. The skill you practice is the hand-verify-one-cell method on a responder comparison table from Study XYZ: you build it in PROC FREQ, the SAS Statistical Analysis System frequency procedure, using the chi-square test the plan expects and Fisher's exact test, cross-foot the N by summing the intent-to-treat flag ITTFL in the subject-level analysis dataset ADSL against the distinct count of the unique subject identifier USUBJID, and confirm the test statistic and degrees of freedom against the procedure's Output Delivery System ODS output instead of the rendered table. The drill also plants near-neighbor traps, such as a t-test where the statistical analysis plan, SAP, says Wilcoxon rank-sum, or an asymptotic chi-square where the SAP says exact, so try it right after the video and see whether you catch them.

Checkpoint11 / 12

Final Knowledge Check: Explain, Program, Prove

1 A table from a pre-specified Fisher exact test shows a p-value of 0.049. Which interpretation is correct when the Statistical Analysis Plan (SAP) declares a two-sided alpha of 0.05?

2 After the SAP-defined multiplicity control has been applied to the primary and key secondary endpoints, which QC statements are necessary before a table row is released? (select all that apply, then Check)

3 A 2x2 table for a binary endpoint has one expected cell count below 5. The SAP pre-specified: use the chi-square test when all expected counts are at least 5; otherwise use the Fisher exact test. Which choice correctly matches that SAP switching rule?

4 For a continuous repeated-measures primary endpoint, the SAP chooses MMRM (mixed model for repeated measures) rather than LOCF (last observation carried forward). Which statement best describes the MMRM approach?

5 The SAP (Statistical Analysis Plan) specifies the log-rank test for a time-to-event comparison, and the table will show a log-rank p-value computed from ADTTE (the analysis dataset for time-to-event). List three QC reconciliations that must be performed before the table is released. In your answer, address the event count/denominator and the SAS output that supports the p-value. (reflect, then reveal)

Reveal analysis
Reference answer: Confirm ADTTE contains one analysis record per subject and that the event/censor flags are internally consistent with ADSL. Count observed events by treatment group and verify that every denominator presented is an event count or risk set, not a randomized-patient count. Re-run or verify PROC LIFETEST with STRATA treatment; the log-rank test compares observed and expected event times across treatment groups. Reconcile the chi-square, degrees of freedom, and p-value with the ODS output, and make sure the table footnote names the same log-rank test as the SAP.
Speaker notes

This is the final knowledge check: explain, program, and prove each p-value and clinical Statistical Analysis System (SAS) test decision. When a pre-specified Fisher exact test shows a p-value of 0.049 under a Statistical Analysis Plan (SAP)-declared two-sided alpha of 0.05, the correct interpretation is B: assuming no association exists, the probability of a table as extreme as or more extreme than the observed table is 0.049, because a p-value is a conditional probability under the null hypothesis, not the probability that the null is true. After SAP-defined multiplicity control, the necessary quality control (QC) statements before releasing a table row are A and B: a row may be marked significant only if the endpoint passed its pre-specified hierarchical gate, and the estimate, confidence interval, and p-value must come from the same SAP-specified statistical model and analysis population. For a 2x2 binary table with one expected cell count below 5, when the SAP says use chi-square only if all expected counts are at least 5 and otherwise use Fisher's exact, the correct match is B: run PROC FREQ with CHISQ and EXACT FISHER, then report the Fisher exact p-value, because the chi-square approximation is unreliable with small expected counts. For a continuous repeated-measures primary endpoint, when the SAP chooses mixed model for repeated measures (MMRM) rather than last observation carried forward (LOCF), the correct description is A: MMRM estimates treatment-group mean differences over time from all available postbaseline data and models within-subject correlation, so it does not impute missing values, unlike LOCF. For the log-rank time-to-event comparison, the model answer is to reconcile the analysis dataset for time-to-event (ADTTE) event count and denominator against the SAP and clinical database, confirm censoring and risk sets, and verify the log-rank p-value against the SAS output from PROC LIFETEST or the Output Delivery System (ODS) table before release.

Concept12 / 12

Key Takeaways: The Same Week, Every Week

Key Takeaways

The Same Week, Every Week — durable rules

Statistical dialect

• Each test: compares what, SAP switch call, PROC.

• Precision ≠ signal: CI beside estimate; same pop & N.

• Informative dropout flattened LOCF; SAPs moved to MMRM; check PROC MIXED df.

QC craft & sources

• QC: hand-check one cell per block; tie every N to ADSL.

• Confirm rounding vs shell before trusting any p-value.

• Paired reading: p-values/tests article, ADTTE, Part 17 (LOCF), Part 3 (shells).

No participant rows; source SAS only, none fabricated.

Summary slide (L1/L2). Recaps the durable rules of the lesson, links back to the paired source article and its companion readings, and names the material-data gap honestly.

Speaker notes

For each test, say what it compares, when the Statistical Analysis Plan (SAP) switches to it, and which procedure runs it. Confidence intervals (CIs) ride beside estimates because precision is separate from signal; verify the pair uses one population and one N. Longitudinal SAPs moved to Mixed Model for Repeated Measures (MMRM) because Last Observation Carried Forward (LOCF) manufactured flat data under informative dropout; check the denominator degrees of freedom in PROC MIXED. Quality Control (QC): hand-verify one cell per block, tie every N to the Subject-Level Analysis Dataset (ADSL) flag, and check rounding against the shell before trusting a p-value. Pair reading: the p-values and tests article, the Time-to-Event Analysis Dataset (ADTTE) companion for log-rank, Part 17 for LOCF, and Part 3 for shells. This lesson used only canonical code snippets; no participant rows, so no fabricated patient numbers.

✓

Lesson complete

Nice work — every scene seen. Keep the momentum going.

← → Space to navigate · progress is saved locally in your browser

AI Tutor

Ask the tutor