bench

Featured

Plan real experiments and analyze supplied continuous measurements in Lab Bench. Use for defining factors and independent runs, randomized collection sheets, measurement CSV imports, experimental uncertainty and diagnostics, comparisons and follow-up confirmation runs with a reproducible report.

AI & Automation 957 stars 79 forks Updated today MIT

Install

View on GitHub

Quality Score: 93/100

Stars 20%
99
Recency 20%
100
Frontmatter 20%
70
Documentation 15%
100
Issue Health 10%
50
License 10%
100
Description 5%
100

Skill Content

# bench Build a usable experiment from the person's question. Read [the project contract](references/project.md) when authoring or revising the source. Keep the procedure, independent unit, factors, levels, response unit and intended decision explicit. Use the studio to make the work inspectable. 1. Inspect `bench/project.json` and `bench/DESIGN.md`. Preserve existing collected runs, approved protocol, original files and change history. A new brief is open ended; the coffee plan is only an example. Record necessary assumptions and choose an achievable initial run budget. 2. Generate a balanced factorial with independent repetitions and complete blocks. If adding controls, bracket and intersperse them. A two-level design plus shared centers cannot separately identify multiple quadratic effects; inspect matrix rank before collection. 3. Build with `node tools/build.mjs` and give the person a real collection sheet. Do not invent observations. Use blank responses for unfinished runs. Labels, units and ids must survive CSV round trips. A source edit must leave the browser's source-save bridge operational. 4. Import original bytes with `prepareImport` / `commitImport`, or use `recordMeasurement` with a reason for a direct observation or correction. Do not quietly change units, sample identity, settings or response values. Exclude only with a defensible recorded reason; keep the value and compare the sensitivity fit using every recorded training measurem...

Details

Author
autonomous-ai
Repository
autonomous-ai/openharness
Created
1 months ago
Last Updated
today
Language
C
License
MIT

Similar Skills

Semantically similar based on skill content — not just same category

AI & Automation Listed

run-ab-experiments

Use this skill when a user wants to plan, review, launch, manage, intervene in, ramp, monitor, analyze, or interpret an online controlled experiment or A/B test; choose hypotheses, units, variants, exposure, metrics/OEC, guardrails, sample size, duration, stopping rules, SRM checks, or ship/iterate/stop decisions. Manage the full lifecycle in three modes—plan, manage, and interpret—with explicit cognitive-load reviews and concise, non-repetitive outputs grounded in Kohavi, Tang, and Xu. Always begin with a mandatory question-first interview and context-and-assumptions confirmation before work. Do not use it for observational-only causal claims unless deciding whether an experiment is feasible.

0 Updated 2 weeks ago
shahriarfarzadi
AI & Automation Listed

benchmark-methodology

Rules for measuring or quoting any benchmark number in this repo. Load before running the bench tier, writing a delta into a PR or CHANGELOG, or touching anything under benchmarks/.

2,715 Updated today
jgravelle
AI & Automation Listed

benchforge

Open Scientific Evidence Infrastructure for the Agentic Era (BenchForge v6.0 Final Specification & MVP Execution Plan). Features BDL v6.0, 2-Tier Dual Reporting Standard (BENCHMARK_SUMMARY.md for GitHub README & BENCHMARK.md for Research Deep-Dive), Multi-Variable Workload Taxonomy, Threat Model Validation, Immutable Hash Chain Evidence Ledger, Complete Agent Composition, Sequential Bayesian Adaptive Sampling, Blind Human Review Protocol, and Scientific Artifact Triad.

0 Updated 1 months ago
AxelS27