skills/hypothesis-generation/references/hypothesis_quality_criteria.md
These criteria structure expert review. Do not sum them, assign weights, calculate a “quality score,” rank candidates automatically, or select a winner. Trade-offs and domain assumptions are not commensurable numbers.
For each criterion record:
Treat novelty as a separate evidence claim:
Use “not located within the documented search boundary” when that is all the evidence supports. Absence from a quick search is not evidence of novelty.
FINER—Feasible, Interesting, Novel, Ethical, Relevant—is a mnemonic for refining a question, not a pass/fail instrument. The earliest source located in this refresh is the first edition of Designing Clinical Research (Hulley and Cummings, 1988); later editions and current methodological articles present the mnemonic. The dated search did not establish that the 1988 edition was the first printed use, so do not claim coinage without checking the primary text.
A null result can be uninformative because of low precision, failed manipulation, insensitive measurement, missingness, or assumption failure. Record these possibilities before calling a result falsifying.
Mechanistic detail is not evidence. A more elaborate story can be less testable.
Separate:
State which assumptions are testable, partially diagnosable, or fundamentally untestable with available data.
For every source:
A prediction should identify:
Do not invent numerical effect sizes merely to appear specific. If magnitude is unknown, prespecify the direction, smallest effect of scientific interest, precision target, or a range of plausible values with rationale.
For each construct ask:
Require:
Predictive performance does not identify a causal effect. Adjustment does not guarantee exchangeability. Conditioning on a mediator or collider can introduce bias.
Check:
End human review with one of:
revise_before_test;ready_for_preregistration_review;blocked_by_safety_or_ethics_gate;blocked_by_measurement_or_feasibility;retain_as_exploratory_candidate;requires_specialist_review.These are workflow states, not scientific truth judgments and not outputs of a score.