human-eval-handoff-repair

Solid

Use when validating, repairing, or mapping human-evaluation handoff packages, filled annotation CSVs, rebuttal annotation UIs, or reviewer annotation returns across package versions. Applies to checking row alignment, detecting cross-snapshot contamination, converting old labels to current schemas, generating refill/import CSVs, and deciding whether filled annotations are safe for formal aggregation.

AI & Automation 41 stars 7 forks Updated yesterday MIT

Install

View on GitHub

Quality Score: 87/100

Stars 20%
54
Recency 20%
100
Frontmatter 20%
70
Documentation 15%
100
Issue Health 10%
50
License 10%
100
Description 5%
100

Skill Content

# Human Eval Handoff Repair Use this skill when the user asks to inspect, repair, migrate, or validate human-evaluation packages or filled annotation CSVs, especially when multiple package versions exist. ## First Principle Never trust row number, filename, or visible ID alone. Treat a filled CSV as usable only after it is aligned to the target public package by stable task-specific keys and its editable labels pass schema checks. Do not fix labels by guessing. If a row cannot be matched safely, leave its annotation fields blank and produce a refill/import template plus an unmatched reference file. ## Inputs To Locate - Target public handoff package folder or ZIP. - Filled CSVs from annotators. - Any older package claimed as the source version. - Current task schemas from the target UI data or target task CSVs. - If relevant, QC reports from the target package. Prefer a user-specified output directory. If none is specified, write generated reports and repaired files under `codex_outputs/` in the current workspace. Avoid synced personal document folders unless the user explicitly asks for them. ## Task Types And Stable Keys Use these stable keys before copying any labels: - Claim Warrant Audit: `item_id`, `tradition`, `medium`, `period`, `layer`, `dimension_id`, `claim_text`. - Release Suitability Audit: `item_id`, `source_bucket`, `pipeline_mode`, `option_a_source`, `option_b_source`, `governed_option`, `option_a_text`, `option_b_text`. - Card Faithfulness Spot Audi...

Details

Author
yha9806
Repository
yha9806/academic-writing-toolkit
Created
6 months ago
Last Updated
yesterday
Language
Python
License
MIT

Similar Skills

Semantically similar based on skill content — not just same category

AI & Automation Listed

canvas-humanizer-surgical

Use after canvas-humanizer when a local draft has specific meaning, grammar, citation, fluency, voice, or rubric defects. Repair only named segments, preserve safe text, and return a verified local artifact without submitting.

125 Updated 1 months ago
X-isdoingreat
AI & Automation Listed

crm-export-auditor

Audit a CRM CSV export for the data faults that make every report built on it wrong: exact and near-duplicate accounts and contacts, close dates already in the past on open deals, deals with no activity in the last N days, blank required fields by stage, single-threaded deals carrying one contact, records with no owner or a departed one, stage and amount combinations that contradict each other, and currency, date and country formats that drift between rows. Returns a fix list ordered by how much each fault distorts a reported number, with the record IDs to correct and which fixes are safe to apply in bulk. Use when a pipeline number is disputed, when a migration or a dedupe is being planned, or when nobody has checked the export the forecast is built on.

1 Updated 2 weeks ago
imtiazrayhan
AI & Automation Listed

session-repair

Transform a legacy layout into the work root, or reconcile a work root that has drifted. Surveys read-only, refuses unsafe ground, prints the per-item plan, applies only on explicit confirmation, then verifies the result entry by entry. Triggers on "/session-repair" or when user says "migrate the backlog to the work root", "the sequence and the records disagree", "reconcile the work root", or "repair session-flow state".

3 Updated 2 weeks ago
matshoppenbrouwers