anti-benchmark

Featured

Challenge industry best practices' hidden assumptions. Deconstruct benchmarks to reveal unexamined constraints.

AI & Automation 501 stars 41 forks Updated yesterday Apache-2.0

Install

View on GitHub

Quality Score: 93/100

Stars 20%
90
Recency 20%
100
Frontmatter 20%
70
Documentation 15%
100
Issue Health 10%
80
License 10%
100
Description 5%
100

Skill Content

# Anti-Benchmark Challenge industry best practices' hidden assumptions. ## State Ledger | Resource | Target | Current | % | |----------|--------|---------|---| | web-search | 25 | 0 | 0% | | web-research | 10 | 0 | 0% | | paper-overview | 25 | 0 | 0% | | paper-search | 15 | 0 | 0% | | paper-research | 5 | 0 | 0% | ## HARD-GATE Cannot exit strategy until ≥80% of each budget line is consumed OR yield targets are met with justification for remaining budget. ## Available Tactics | Tactic | Role | |--------|------| | assumption-enumeration | Surface assumptions hidden in benchmarks | ## Available SOPs | SOP | Role | |-----|------| | benchmark-challenge | Identify and negate benchmark assumptions | | sacred-cow-identification | Find unquestioned beliefs behind best practices | | assumption-perturbation | Test what happens when benchmark assumptions fail | | constructive-rebellion | Build alternatives that violate benchmarks constructively | | destruction-synthesis | Synthesize anti-benchmark outputs | ## Execution Guidance 1. **Identify benchmarks**: Catalog the industry best practices and standards in the domain 2. **Deconstruct**: Use benchmark-challenge to expose hidden assumptions in each 3. **Surface sacred cows**: Use sacred-cow-identification for deeper unquestioned beliefs 4. **Perturb**: Use assumption-perturbation to test "what if this standard is wrong?" 5. **Research alternatives**: Search for domains that succeed WITHOUT these benchmarks 6. **Build**: Use co...

Details

Author
yogsoth-ai
Repository
yogsoth-ai/de-anthropocentric-research-engine
Created
7 months ago
Last Updated
yesterday
Language
HTML
License
Apache-2.0

Similar Skills

Semantically similar based on skill content — not just same category

AI & Automation Listed

benchmark-analysis

Takes one of the user's own metrics and checks it against a stated benchmark source, returning a clear over/under read and what that gap actually means. Use when the user wants to know if a number (churn rate, CAC, conversion rate, NPS) is good or bad relative to a real reference point, not just the number in isolation. Boundary: this skill does not have a live connection to any benchmark database. It compares against whatever source the user supplies, or discloses plainly when it is using general public knowledge instead.

1 Updated today
sidchaudhary
AI & Automation Listed

benchmark-curator

Design, curate, version, and quality-control benchmark or evaluation corpora for Agent Skills and quality workflows, including task taxonomies, discovery/forced/negative controls, adversarial and regression cases, difficulty strata, holdout isolation, provenance, duplication, contamination risk, coverage balance, and immutable benchmark hashes. Use when the user asks to build or maintain an eval dataset, golden set, benchmark suite, regression corpus, holdout, challenge set, or representative test cases. Do not use to execute the model experiment or claim lift (use skill-evaluator), to design the grading rubric (use rubric-designer), to edit the candidate skill (use skill-creator), or to treat leaked/known cases as a clean holdout.

1 Updated today
CometWeb-io
AI & Automation Listed

benchmark-methodology

Rules for measuring or quoting any benchmark number in this repo. Load before running the bench tier, writing a delta into a PR or CHANGELOG, or touching anything under benchmarks/.

2,715 Updated today
jgravelle