skills/peer-review/references/common_issues.md
Use this reference as a prompt for inquiry, not a defect checklist. A possible issue becomes a review comment only when it is relevant to the study and supported by a manuscript location, supplied artifact, or applicable method principle.
Do not infer misconduct, poor quality, or manuscript merit from a missing reporting item. Separate:
Check each central claim against the design, analysis, result, and uncertainty that support it.
Common mismatches:
Constructive response:
Use scripts/validate_claim_evidence.py for a local identifier-based matrix. Its report never echoes claim text.
Check whether the population, intervention or exposure, comparator, outcomes, timing, and target quantity align from objectives through interpretation. For trials, identify the estimand when relevant. For prediction, distinguish model development from performance evaluation. For diagnostic studies, distinguish diagnostic accuracy from clinical utility.
Potential issues include:
Request a clear definition of the unit, nesting, repeated measures, and analysis that reflects dependence. Do not assume a mixed model is always the correct remedy; the model must match the design and question.
Assess, as applicable:
Avoid treating baseline significance tests as proof of successful randomization. Focus on chance imbalance, clinically important imbalance, prespecified adjustment, and departures from the randomized comparison.
For causal claims, ask:
Do not demand a specific causal method without showing why it fits the data-generating process.
Avoid fixed heuristics such as “n < 30 is too small” or “three replicates are sufficient.” Adequacy depends on the target effect or precision, variability, design effect, event count, model complexity, multiplicity, attrition, and decision context.
Check:
Observed or post hoc power calculated from the observed effect generally adds little beyond the estimate and its interval. Request effect estimates and uncertainty rather than “achieved power.”
Check whether the analysis respects:
Do not prescribe “parametric” or “non-parametric” methods from sample size alone.
The relevant assumptions depend on the estimand and model. A standalone normality test is not a universal gatekeeper and can be uninformative in very small or large samples. Look for design-aware diagnostics, residual behavior, influential observations, functional form, calibration, proportional hazards where applicable, and sensitivity to reasonable alternatives.
Comments should identify the assumption at risk and why it matters. “Check normality” without specifying the modeled quantity or consequence is not actionable.
Flag:
Prefer estimates, uncertainty, assumptions, and context. The ASA p-value principles and SAMPL reporting guidance are indexed in assets/source_ledger.csv.
Assess:
Not every collection of analyses requires the same correction. Ask authors to state the inferential family and rationale instead of automatically demanding Bonferroni adjustment.
Check:
Do not require a test that data are “missing completely at random”; missingness assumptions are not generally established by a single diagnostic test.
Check whether exclusions, transformations, winsorization, detection-limit handling, and influential-observation rules were prespecified or transparently justified. Request sensitivity analyses when conclusions depend materially on discretionary handling. Do not demand deletion merely because a value is extreme.
Look for prespecification, adequate interaction analysis, multiplicity, uncertainty, biological or clinical rationale, and consistency of direction. Within-group significance and between-group non-significance do not establish subgroup differences.
Check:
TRIPOD+AI applies to regression and machine-learning prediction models; STARD-AI applies when diagnostic accuracy is the primary evaluation target.
Check whether another qualified researcher could understand and, where permissions allow, repeat the work:
“Available on request” is not automatically invalid, and open release is not always ethical or lawful. Evaluate whether the access route is specific, feasible, and consistent with governance.
Do not claim to have reproduced an analysis unless it was actually run with documented inputs, environment, commands, and outputs.
Assess the supplied artifact directly; do not infer manipulation from low-resolution rendering alone.
Check:
Possible duplication or manipulation should be documented neutrally by location and referred to the editor under the journal’s image-integrity process. Do not accuse authors of fabrication.
Check what is applicable:
If a concern cannot safely be raised with authors, use the confidential editor channel. State the evidence and uncertainty; do not investigate people, contact institutions, or reveal the manuscript outside the authorized process.
Check:
The local scripts/audit_citations.py checks Pandoc-style keys such as [@ref-id] against a CSV. It does not verify source existence or support and must not be described as doing so.
For each major or minor comment, include:
Prefer: “At Methods, paragraph 3, the experimental unit is unclear. Because three measurements appear to come from each participant, please define the unit and explain how within-participant dependence was handled.”
Avoid: “The statistics are bad.”
Requests for new experiments should be necessary to support an existing central claim, ethically and practically proportionate, and distinguished from optional future work. Often the appropriate remedy is to narrow a claim, add a limitation, provide missing analysis detail, or share an existing artifact.