skills/scholar-evaluation/references/responsible_assessment.md
This skill is for developmental review of scholarly works and for auditing a low-stakes assessment process. It must not automate, recommend, materially influence, or provide a score for:
Do not rank people. Do not convert several judgments into a single composite person score. Do not infer a person's ability, character, integrity, future performance, protected characteristics, or institutional worth.
This boundary remains in force when a human is nominally “in the loop.” Checklists, committees, or disclaimers do not make a prohibited workflow safe. If a request crosses the boundary, stop and offer one of these alternatives:
Examples of allowed uses, subject to authorization and data protection:
“Publication readiness” and “accept/reject” recommendations are excluded. Describe evidence, limitations, and improvement options instead.
For any organizational use, an accountable committee must own the process. It must include relevant disciplinary and assessment-methods expertise and must:
No script in this skill is the accountable reviewer. Script output is a descriptive record for qualified human interpretation.
Start with the values and construct, not the data that happen to be available. Use evidence that directly bears on the criterion:
Ask what is missing, inaccessible, contested, or not applicable. A missing item is not a zero. A not-applicable item is not evidence of deficiency.
Do not score or infer quality from:
The bundled rubric validator rejects common proxy-measure criteria.
If a qualified reviewer has a legitimate, predeclared reason to mention a quantitative indicator descriptively outside the bundled scoring tools, record:
Never use an indicator merely because it is available. Never hide several different indicators inside an opaque composite.
A rubric is a measurement claim. Before operational use, record:
Do not call a rubric “validated” because experts reviewed it once, raters agreed, or scores correlated with another judgment. Validity concerns the evidence for a specific interpretation and use. Reliability or agreement alone is not validity.
The provided rubric explicitly records content_validity_status as
not_established. Replace that status only when a qualified team has documented
appropriate evidence for the exact discipline, population, language, and use.
Perform the fairness review outside these scripts in an authorized environment. Do not place protected-attribute records in rubric, evaluation, evidence, or ratings files.
A qualified review should examine, where lawful and appropriate:
Report sample limitations and uncertainty. Do not expose small cells or attempt to infer sensitive characteristics.
Use only public scholarly works, synthetic records, or deidentified low-stakes records processed under an approved local purpose. Do not put raw private applications, CVs, recommendation letters, reviewer identities, contact details, protected attributes, or source-document text into tool inputs or outputs.
The bundled formats contain only:
Keep source documents in the authorized records system. Use local references to them. Apply least privilege, retention limits, deletion, incident handling, and any stricter local law or policy.
Provide the rubric, evidence requirements, notices, feedback, and appeal process in accessible formats. Do not penalize an accommodation, assistive technology, language variant, or accessible presentation choice. Confirm that the rubric measures the intended construct rather than fluency with an inaccessible interface or format.
Lead with qualitative, traceable findings. For every criterion:
Do not label a person, declare a work “top-tier,” predict success, recommend a decision, or conceal uncertainty behind a decimal.