bench
FeaturedPlan real experiments and analyze supplied continuous measurements in Lab Bench. Use for defining factors and independent runs, randomized collection sheets, measurement CSV imports, experimental uncertainty and diagnostics, comparisons and follow-up confirmation runs with a reproducible report.
Install
Quality Score: 93/100
Skill Content
Details
- Author
- autonomous-ai
- Repository
- autonomous-ai/openharness
- Created
- 1 months ago
- Last Updated
- today
- Language
- C
- License
- MIT
Similar Skills
Semantically similar based on skill content — not just same category
run-ab-experiments
Use this skill when a user wants to plan, review, launch, manage, intervene in, ramp, monitor, analyze, or interpret an online controlled experiment or A/B test; choose hypotheses, units, variants, exposure, metrics/OEC, guardrails, sample size, duration, stopping rules, SRM checks, or ship/iterate/stop decisions. Manage the full lifecycle in three modes—plan, manage, and interpret—with explicit cognitive-load reviews and concise, non-repetitive outputs grounded in Kohavi, Tang, and Xu. Always begin with a mandatory question-first interview and context-and-assumptions confirmation before work. Do not use it for observational-only causal claims unless deciding whether an experiment is feasible.
benchmark-methodology
Rules for measuring or quoting any benchmark number in this repo. Load before running the bench tier, writing a delta into a PR or CHANGELOG, or touching anything under benchmarks/.
benchforge
Open Scientific Evidence Infrastructure for the Agentic Era (BenchForge v6.0 Final Specification & MVP Execution Plan). Features BDL v6.0, 2-Tier Dual Reporting Standard (BENCHMARK_SUMMARY.md for GitHub README & BENCHMARK.md for Research Deep-Dive), Multi-Variable Workload Taxonomy, Threat Model Validation, Immutable Hash Chain Evidence Ledger, Complete Agent Composition, Sequential Bayesian Adaptive Sampling, Blind Human Review Protocol, and Scientific Artifact Triad.