skills/scholar-evaluation/references/evaluation_framework.md
The paper currently referenced by this skill is Moussa et al., ScholarEval: Research Idea Evaluation Grounded in Literature, arXiv:2510.16234v2, revised 2026-02-28.
The preprint describes an experimental retrieval-augmented system for evaluating research ideas on:
It reports a 117-idea, four-discipline dataset, coverage comparisons with expert-annotated review points, and a user study. Those studies evaluate that framework. They do not validate the generalized rubric in this skill as a psychometric instrument, establish stable score meaning across disciplines, or authorize use in consequential decisions.
As of the dated source review, arXiv and the official project repository were
the verified primary publication records. No peer-reviewed publication status
was verified. Cite it as an arXiv preprint unless a later primary record is
checked. See references/source_ledger.md.
This skill borrows the useful discipline of:
It does not reproduce ScholarEval's model pipeline, prompts, retrieval system, dataset, or reported metrics. The bundled scripts do not call ScholarEval, search the web, invoke a model, or evaluate private documents.
The template is a locally governed developmental rubric. Its default construct is:
Traceable support for a scholarly work's claims and methods: the degree to which a work states a bounded question, situates its contribution, uses fit-for-purpose methods, aligns analysis with claims, and documents transparent and responsible practices using traceable evidence.
This construct must be reviewed and adapted by relevant disciplinary experts.
Review:
Do not treat fashionable topics, institutional affiliation, or venue expectations as evidence of significance.
Review:
Failure to find prior work does not establish novelty. Search coverage varies by database, language, date, indexing, terminology, discipline, and access.
Review:
Use discipline-specific reporting and methods standards. Do not reward complexity for its own sake.
Review:
The rubric's uncertainty value is a rater-supplied bounded judgment range. It
is not a sampling confidence interval, posterior interval, or standard error.
Review:
Open practice is not an absolute requirement when privacy, consent, safety, security, Indigenous data governance, commercial constraints, or other legitimate restrictions apply. Assess whether restrictions are justified and whether safe access or metadata alternatives are provided.
The template uses an ordinal 0–4 scale:
These are evidence anchors, not labels of a person or universal levels of research quality. The rubric defines criterion-specific anchors. Raters must use the anchor text, not intuition about what a number “usually means.”
Do not convert the score to:
Each criterion has exactly one status:
rated: score, uncertainty, evidence identifiers, and rationale reference
are required;missing: evidence needed for assessment is absent or unavailable; score and
uncertainty are null; ornot_applicable: the criterion does not apply to this work under a documented
rationale; score and uncertainty are null.Do not encode missing or not-applicable as zero.
For rated criteria (R), score (s_i), and predeclared weight (w_i):
\frac{\sum_{i \in R} w_i s_i} {\sum_{i \in R} w_i} ]
The score report separately provides:
The uncertainty aggregation is not a confidence interval. Normalization does not make incomplete evaluations comparable. Review missingness before looking at any score.
Before replacing content_validity_status: not_established, document:
Rubric provenance must identify the version, owner role, source identifiers, review date, and content-evidence reference.
At minimum:
The bundled agreement script reports exact agreement, within-one-step agreement,
and mean absolute difference. Those summaries do not replace a
design-appropriate reliability analysis. The rubric therefore separately
records inter_rater_reliability_status and
inter_rater_reliability_ref; the template leaves reliability not established.
Each rated criterion must point to one or more entries in the evidence manifest. Each entry records:
Never place an excerpt or raw private document in the manifest. Keep source content in the authorized source system.
Weights are value judgments. Predeclare and justify them. Run
scripts/weight_sensitivity.py before interpreting a composite.
The script increases and decreases one weight at a time and renormalizes the weights. It reports score ranges and whether pairwise ordinal relationships among scholarly works change. Instability is evidence that an apparent order depends on contestable weights.
The output must not be used to rank people or decide a high-impact outcome. Even stable ordering does not establish validity.
For each criterion, qualified reviewers should record:
Conclude with construct, provenance, coverage, agreement, sensitivity, bias, privacy, accessibility, and validity limitations—not a decision recommendation.