Back to Claude Scientific Skills

Responsible Assessment and Safety Boundary

skills/scholar-evaluation/references/responsible_assessment.md

2.55.08.9 KB
Original Source

Responsible Assessment and Safety Boundary

Non-negotiable boundary

This skill is for developmental review of scholarly works and for auditing a low-stakes assessment process. It must not automate, recommend, materially influence, or provide a score for:

  • hiring, promotion, or tenure;
  • admissions;
  • grant or other funding decisions;
  • prizes, honors, or awards;
  • discipline, dismissal, or sanctions; or
  • any other high-impact personnel decision.

Do not rank people. Do not convert several judgments into a single composite person score. Do not infer a person's ability, character, integrity, future performance, protected characteristics, or institutional worth.

This boundary remains in force when a human is nominally “in the loop.” Checklists, committees, or disclaimers do not make a prohibited workflow safe. If a request crosses the boundary, stop and offer one of these alternatives:

  1. developmental comments on a public or authorized scholarly work, with no person comparison or decision advice;
  2. an audit of whether an existing assessment process follows responsible assessment principles, without processing applications or recommending an outcome; or
  3. neutral documentation of criteria for review by the organization's legal, privacy, accessibility, labor, ethics, and disciplinary experts.

Allowed scope

Examples of allowed uses, subject to authorization and data protection:

  • feedback on a draft paper, protocol, research idea, or literature synthesis;
  • a retrospective methods or reporting review;
  • calibration exercises using synthetic or public scholarly works;
  • checking evidence traceability;
  • describing the sensitivity of work-level scores to rubric weights; and
  • auditing a low-stakes evaluation design for missing governance controls.

“Publication readiness” and “accept/reject” recommendations are excluded. Describe evidence, limitations, and improvement options instead.

Accountable human process

For any organizational use, an accountable committee must own the process. It must include relevant disciplinary and assessment-methods expertise and must:

  • publish the construct, intended use, rubric, weights, evidence requirements, and interpretation limits before reviewing;
  • record member qualifications, training, calibration, and drift checks;
  • disclose conflicts, require recusal, and maintain a conflict record;
  • provide an understandable notice and a meaningful appeal or correction route;
  • provide accessible materials and reasonable accommodations;
  • define lawful purpose, access controls, minimization, retention, and deletion;
  • review disciplinary, language, career-path, disability, and subgroup effects;
  • document disagreements and uncertainty rather than force consensus; and
  • periodically evaluate and revise the evaluation.

No script in this skill is the accountable reviewer. Script output is a descriptive record for qualified human interpretation.

Qualitative-first evidence

Start with the values and construct, not the data that happen to be available. Use evidence that directly bears on the criterion:

  • the work's questions, methods, analyses, outputs, limitations, and provenance;
  • datasets, software, protocols, materials, registrations, replications, and negative or null findings where relevant;
  • transparent records of responsible practices and justified restrictions;
  • contribution records, including CRediT roles when useful, without treating a role as proof of quality; and
  • influence on policy, practice, communities, teaching, infrastructure, or knowledge, when this is within the stated construct and supported by evidence.

Ask what is missing, inaccessible, contested, or not applicable. A missing item is not a zero. A not-applicable item is not evidence of deficiency.

Prohibited proxies and contextual indicators

Do not score or infer quality from:

  • Journal Impact Factor or other journal-level measures;
  • h-index, i10-index, publication counts, or citation counts;
  • altmetrics or attention counts;
  • journal, conference, institutional, geographic, or employer prestige;
  • venue identity or ranking; or
  • author affiliation, career path, network, or reputation.

The bundled rubric validator rejects common proxy-measure criteria.

If a qualified reviewer has a legitimate, predeclared reason to mention a quantitative indicator descriptively outside the bundled scoring tools, record:

  1. the exact construct and purpose;
  2. why the indicator bears on that construct at the correct unit of analysis;
  3. source, version, query date, coverage, exclusions, and data quality;
  4. field, language, output-type, career-stage, and time-window effects;
  5. uncertainty, missingness, gaming risks, and known biases;
  6. why qualitative evidence is insufficient by itself; and
  7. a statement that the indicator is not a direct measure of quality.

Never use an indicator merely because it is available. Never hide several different indicators inside an opaque composite.

Rubric evidence and psychometric caution

A rubric is a measurement claim. Before operational use, record:

  • construct: what is and is not being assessed;
  • intended interpretation and use: the exact meaning claimed for scores;
  • provenance: who designed and approved criteria, anchors, and weights;
  • content evidence: disciplinary expert and stakeholder review of coverage;
  • response process: how raters interpret anchors and use evidence;
  • rater protocol: selection, training, calibration, qualification, and drift;
  • agreement/reliability: a design-appropriate analysis and its uncertainty;
  • fairness: accessibility, subgroup, language, and disciplinary review;
  • traceability: stable evidence references for each rating;
  • missing/not applicable: explicit statuses and rationales;
  • weight sensitivity: whether plausible weights change descriptive results;
  • consequences: gaming, burden, goal displacement, and other effects; and
  • revision: review date, owner, change record, and retirement criteria.

Do not call a rubric “validated” because experts reviewed it once, raters agreed, or scores correlated with another judgment. Validity concerns the evidence for a specific interpretation and use. Reliability or agreement alone is not validity.

The provided rubric explicitly records content_validity_status as not_established. Replace that status only when a qualified team has documented appropriate evidence for the exact discipline, population, language, and use.

Bias and subgroup review

Perform the fairness review outside these scripts in an authorized environment. Do not place protected-attribute records in rubric, evaluation, evidence, or ratings files.

A qualified review should examine, where lawful and appropriate:

  • access to the measured construct and accommodation effectiveness;
  • differential missingness and evidence availability;
  • criteria that privilege particular languages, methods, fields, institutions, career patterns, or resource levels;
  • rater severity, drift, and disagreement patterns;
  • differential effects of weights and not-applicable decisions;
  • false precision and threshold effects;
  • burden, gaming, and chilling of collaboration or risky research; and
  • whether the evaluation should be redesigned or stopped.

Report sample limitations and uncertainty. Do not expose small cells or attempt to infer sensitive characteristics.

Privacy and data protection

Use only public scholarly works, synthetic records, or deidentified low-stakes records processed under an approved local purpose. Do not put raw private applications, CVs, recommendation letters, reviewer identities, contact details, protected attributes, or source-document text into tool inputs or outputs.

The bundled formats contain only:

  • pseudonymous work, evaluation, rater, criterion, and evidence identifiers;
  • bounded scores, statuses, and uncertainty;
  • local stable references; and
  • minimized control attestations and aggregate summaries.

Keep source documents in the authorized records system. Use local references to them. Apply least privilege, retention limits, deletion, incident handling, and any stricter local law or policy.

Accessibility

Provide the rubric, evidence requirements, notices, feedback, and appeal process in accessible formats. Do not penalize an accommodation, assistive technology, language variant, or accessible presentation choice. Confirm that the rubric measures the intended construct rather than fluency with an inaccessible interface or format.

Communicating results

Lead with qualitative, traceable findings. For every criterion:

  1. cite local evidence references;
  2. state the rating status;
  3. distinguish observed evidence from interpretation;
  4. report uncertainty and disagreement;
  5. state missing or not-applicable evidence;
  6. explain context and limitations; and
  7. offer non-prescriptive improvement options.

Do not label a person, declare a work “top-tier,” predict success, recommend a decision, or conceal uncertainty behind a decimal.