skills/clinical-decision-support/references/study_reporting.md
Checked 2026-07-23.
Choose by study purpose, design, and evaluation stage. A reporting guideline states what to report; it does not prove that the study was well designed, unbiased, clinically useful, safe, or effective.
Use risk-of-bias/applicability tools separately and preserve human judgments.
| Study or artifact | Primary framework | Important companion |
|---|---|---|
| Observational cohort/case-control/cross-sectional | STROBE | RECORD for routinely collected data |
| Prediction model development or evaluation | TRIPOD+AI | PROBAST+AI |
| Tumor prognostic marker | REMARK | Appropriate risk-of-bias and assay guidance |
| Diagnostic accuracy | STARD | STARD-AI when the index test uses AI |
| Randomized AI intervention protocol | Current SPIRIT base | SPIRIT-AI extension |
| Randomized AI intervention report | Current CONSORT base | CONSORT-AI extension |
| Early-stage live AI support evaluation | DECIDE-AI | Design-specific guideline |
| Evidence profile | GRADE | Design-specific risk-of-bias tools |
STROBE addresses reporting of cohort, case-control, and cross-sectional studies. Use the design-specific checklist and explanation material from the STROBE site.
RECORD extends STROBE for routinely collected health data such as administrative, EHR, primary-care surveillance, and registry data. It emphasizes code lists/algorithms, database linkage, selection, cleaning, and data-access transparency. See the RECORD site.
Neither framework authorizes this skill to read EHR rows.
TRIPOD+AI, published April 16, 2024, updates reporting guidance for development and evaluation of clinical prediction models using regression or machine-learning methods. It primarily targets non-generative models.
Report, at minimum:
PROBAST+AI, published March 24, 2025, replaces the original PROBAST for broad prediction-model assessment. It has two distinct parts:
Both parts use four domains:
Applicability is assessed for participants/data sources, predictors, and outcome. Do not average signaling questions into a score. Domain and overall judgments require knowledgeable assessors and rationale.
Use the FDA-NIH BEST Resource for terminology. Distinguish:
A biomarker is not itself a measure of how a person feels, functions, or survives. Analytical validation, clinical validation, and clinical utility are distinct.
For tumor prognostic markers, use REMARK. Report specimen handling, assay methods, prespecified hypotheses/cut points, participant flow, missing data, analysis, effect estimates, and validation.
STARD-AI was published September 15, 2025. It adds AI-specific or modified items to STARD 2015 for diagnostic-accuracy studies, including:
Use STARD-AI with STARD. Do not use it for a prognostic prediction model merely because the model returns a class.
SPIRIT-AI and CONSORT-AI were published September 9, 2020.
AI extensions emphasize the intervention version, input acquisition/quality handling, human-AI interaction, integration requirements, errors/failures, and analysis of performance.
DECIDE-AI, published May 18, 2022, covers early-stage live clinical evaluation of AI-based decision-support systems and includes human factors, workflow, safety, and iterative change reporting.
Live evaluation affects real care and is outside this skill's execution boundary. Use the guideline only to understand documentation requirements. Such work requires an approved protocol, qualified investigators, safety oversight, validated systems, applicable authorization, and institutional governance.
Always disclose:
Never state “reported according to” as proof of adherence without a completed checklist and human verification.