P21 / define.xml
Deterministic parsing, source text preserved
Clinical programming / an inspectable draft
admiralagent.
From one specification to code with evidence for every line.
Turn ADaM derivation descriptions into a reviewable Layer IR, then generate admiral R drafts from deterministic templates. See the reasoning, the validation failures, and the decisions that must be made by a human.
The working notebook
See the results, and the judgment too.
Follow one date specification through rule classification, IR validation, code generation, and comparison against a public oracle.
Specification source
Why this classification
rule: impute DM RFSTDTC directly (pilot ADSL convention)
Source text preserved verbatim. Linking shows provenance; it does not mean CHECK has proven the full semantics of the spec.
Reviewable layer pipeline
Rule confidence 1.00
Rule-match metadata, not a probability of correctness
View full IR JSON
{
"dataset": "ADSL",
"variable": "TRTSDTM",
"steps": [
{
"layer": "impute_dtc",
"args": {
"target": "TRTSDTM",
"dtc": "RFSTDTC",
"output_class": "dtm",
"highest_imputation": "M",
"date_imputation": "first"
}
}
],
"spec_origin": "Date of first study treatment, imputed from DM RFSTDTC",
"confidence": 1,
"needs_human": false,
"rationale": "rule: impute DM RFSTDTC directly (pilot ADSL convention)"
}admiral · R
imputation_flagPASS not_all_naLoading interactive snapshot…
Architecture / bounded by design
Rules generate; humans judge.
Redrawn from the DESIGN.md pipeline. The public demo uses rules; the engine's optional LLM backend also only produces IR, constrained by the same validation gates and templates.
rules → closed vocabulary
Inexpressible → needs_human
Shape, parameter, and semantic checks
Invalid IR is refused rendering
admiral R + CHECK
DISCLAIMER + JSON sidecar
run_validation() + Oracle
Reviewer edits IR → regenerate
The 14 layers actually registered in aa_layers()
assignmerge_varlookup_joinimpute_dtcdtm_to_dtdurationdate_shiftcompute_paramsummary_recordextreme_flagcodelist_varobs_numbercategorizecompute_varEvidence / no hidden failures
Real specs, the complete matrix.
Of 49 ADSL variables, 16 have comparable results and 14 reach 100%. REVIEW, ERROR, and non-comparable items are shown as well.
| Comparison basis / notes | |||||
|---|---|---|---|---|---|
| STUDYID | EXECUTED | 100.0% | 100.0% | 254 | char exact |
| USUBJID | EXECUTED | 83.0% | 100.0% | 306 | set overlap (join key) |
| SUBJID | EXECUTED | 100.0% | 100.0% | 254 | char exact |
| SITEID | EXECUTED | 100.0% | 100.0% | 254 | char exact |
| SITEGR1 | REVIEW | — | — | 0 | not in generated ADSL |
| ARM | EXECUTED | 100.0% | 100.0% | 254 | char exact |
| TRT01P | EXECUTED | 100.0% | 100.0% | 254 | char exact |
| TRT01PN | REVIEW | — | — | 0 | not in generated ADSL |
| TRT01A | EXECUTED | 100.0% | 100.0% | 254 | char exact |
| TRT01AN | REVIEW | — | — | 0 | not in generated ADSL |
| TRTSDT | REVIEW | — | — | 0 | not in generated ADSL |
| TRTEDT | REVIEW | — | — | 0 | not in generated ADSL |
| TRTDURD | REVIEW | — | — | 0 | not in generated ADSL |
| AVGDD | REVIEW | — | — | 0 | not in generated ADSL |
| CUMDOSE | REVIEW | — | — | 0 | not in generated ADSL |
| AGE | EXECUTED | 100.0% | 100.0% | 254 | num tol 0.5 | mean|diff|=0.000 |
| AGEGR1 | REVIEW | — | — | 0 | not in generated ADSL |
| AGEGR1N | REVIEW | — | — | 0 | not in generated ADSL |
| AGEU | EXECUTED | 100.0% | 100.0% | 254 | char exact |
| RACE | EXECUTED | 100.0% | 100.0% | 254 | char exact |
| RACEN | REVIEW | — | — | 0 | not in generated ADSL |
| SEX | EXECUTED | 100.0% | 100.0% | 254 | char exact |
| ETHNIC | EXECUTED | 100.0% | 100.0% | 254 | char exact |
| SAFFL | REVIEW | — | — | 0 | not in generated ADSL |
| ITTFL | REVIEW | — | — | 0 | not in generated ADSL |
| EFFFL | REVIEW | — | — | 0 | not in generated ADSL |
| COMP8FL | REVIEW | — | — | 0 | not in generated ADSL |
| COMP16FL | REVIEW | — | — | 0 | not in generated ADSL |
| COMP24FL | REVIEW | — | — | 0 | not in generated ADSL |
| DISCONFL | REVIEW | — | — | 0 | not in generated ADSL |
| DSRAEFL | REVIEW | — | — | 0 | not in generated ADSL |
| DTHFL | EXECUTED | 1.2% | 1.2% | 254 | char exact |
| BMIBL | REVIEW | — | — | 0 | not in generated ADSL |
| BMIBLGR1 | ERROR | — | — | 0 | not in generated ADSL |
| HEIGHTBL | REVIEW | — | — | 0 | not in generated ADSL |
| WEIGHTBL | REVIEW | — | — | 0 | not in generated ADSL |
| EDUCLVL | REVIEW | — | — | 0 | not in generated ADSL |
| DISONSDT | REVIEW | — | — | 0 | not in generated ADSL |
| DURDIS | REVIEW | — | — | 0 | not in generated ADSL |
| DURDSGR1 | REVIEW | — | — | 0 | not in generated ADSL |
| VISIT1DT | REVIEW | — | — | 0 | not in generated ADSL |
| RFSTDTC | EXECUTED | 100.0% | 100.0% | 254 | char exact |
| RFENDTC | EXECUTED | 100.0% | 100.0% | 254 | char exact |
| VISNUMEN | REVIEW | — | — | 0 | not in generated ADSL |
| RFENDT | EXECUTED | 100.0% | 100.0% | 254 | num tol 0.5 | mean|diff|=0.000 |
| DCDECOD | REVIEW | — | — | 0 | not in generated ADSL |
| EOSSTT | REVIEW | — | — | 0 | not in generated ADSL |
| DCSREAS | REVIEW | — | — | 0 | not in generated ADSL |
| MMSETOT | REVIEW | — | — | 0 | not in generated ADSL |
Currently in original CSV order.
Numeric variables use a tolerance of |difference| < 0.5; date variables are compared via as.Date; character values must match exactly. Pairs missing on both sides do not count toward the numeric match rate; missing-status agreement is listed separately. USUBJID is a set-overlap rate. Execution status and match rate are shown separately: values already present in the base data may still be comparable, which does not establish a successful derivation. This matrix is not a CDISC compliance certification.
A few important distinctions
About this draft generator.
Stating the boundaries clearly is part of reviewability.
Is this "AI automatically writing submission code"?
No. It is a draft generator + human-in-the-loop. Every output must be reviewed line by line by qualified personnel, independently QC'd, and validated at the study level. A passing CHECK only proves that the implemented checks passed — it does not mean the spec is complete, clinically correct, or submission-ready.
Does this page call an LLM or execute R?
No. All cases are precomputed once locally in R by the rules backend. The browser only loads static JSON and links the views; it receives no study data and needs no API key.
For a version that actually executes online: Try the platform live submits real jobs to the tfl-platform API (synthetic / CDISC pilot data, 5 anonymous runs per day).
What is the difference between confidence, PASS, and Oracle 100%?
confidence is classifier metadata, not a calibrated probability of correctness. PASS is the result of a specific validation check. The Oracle match rate depends on the sample, missing values, and tolerance basis. TRTSDTM's 100% refers only to the 254 pairs non-missing on both sides; across all 306 records the basis is 83.0%.
Why does BMIBL still need human confirmation after revision?
Changing BASELINE to SCREENING 1 provides HEIGHT and WEIGHT, letting the not-all-missing and uniqueness checks pass. Whether that is an appropriate study baseline still requires confirmation. Its pilot1 Oracle match rate is 85.8%, not 100%.
Where are the public data and source code?
The examples use CDISC pilot public data from pharmaversesdtm / pharmaverseadam, plus local pilot1 public submission specs and oracle. The site publishes only specs, code, IR, and aggregate metrics — no subject records.
GitHub repository (placeholder, link pending public release) · Data and provenance snapshot