skills/hypothesis-generation/references/experimental_design_patterns.md
For each test, link:
candidate → mechanism → prediction → observable → operationalization → design → analysis → interpretation
Choose the design that can distinguish candidates under realistic uncertainty. Do not select a design merely because it is familiar or available.
Where applicable, address:
Apply only the elements relevant to the science and explain omissions.
Every design should state:
Useful for causal contrasts when intervention, allocation, and ethics permit.
Check:
Randomization does not solve measurement bias, nonadherence, post-randomization selection, interference, or poor external validity.
Useful for multiple interventions and interactions.
Check:
Useful when effects are reversible and carryover can be controlled.
Check:
Useful for mechanistic candidates when ethically and technically appropriate.
Check:
A rescue can still be explained by compensatory or non-specific effects.
Useful when candidates predict different ordering or dynamics.
Check:
Can estimate prevalence and associations at a defined time. It usually cannot establish temporal direction. Explicitly consider selection, reverse causation, survival/prevalence bias, and common-method measurement.
Can establish measured temporal ordering and incidence. It does not eliminate confounding.
Check:
Efficient for some rare outcomes.
Check:
Can strengthen causal identification when an assignment mechanism or discontinuity is credible.
Check:
The design label alone does not establish identification.
Use to test implications of assumptions, estimator behavior, or model dynamics.
Record:
Simulation can show consequences within a model, not that the model describes nature.
Separate prediction from causation.
Check:
Record provenance, data-generating process, inclusion, missingness, transformations, version, and prior analysis exposure. Avoid using the same data to generate and confirm a hypothesis without transparent separation or independent validation.
A condition expected to produce a known response. It checks whether the system and measurement can detect a relevant effect.
Matches handling, timing, delivery, or processing without the target active component.
Should not operate through the target mechanism but should share relevant bias pathways. State assumptions and failure interpretation.
A comparator representing no intervention or no difference may be useful, but it is not equivalent to a negative control and may not isolate placebo, handling, expectancy, or background trends.
Before collecting target outcomes:
A precise measure can be precisely wrong. Technical replicates do not replace independent biological, participant, site, or experimental units.
Do not use universal minima.
Base planning on:
Report assumptions and sensitivity to them. Pilot data may be too unstable for definitive effect-size planning; use external evidence, conservative ranges, or precision-based goals where appropriate.
Inventory:
Prespecify a strategy appropriate to the inferential goal, such as:
Do not report only favorable analyses. Threshold crossing is not a quality score or probability that a candidate is true.
Define:
Complete-case analysis is not automatically unbiased. Record deviations without overwriting the original plan.
Following the National Academies terminology:
Plan:
Open sharing remains subject to consent, privacy, community governance, intellectual property, biosecurity, and other controls.
For randomized intervention work:
These are reporting guidelines. They do not replace ethics review, trial registration rules, statistical expertise, or regulatory requirements.