tangxiangru
UserAI handles execution, humans own the direction, and every run becomes an inspectable research artifact on disk.
Categories
Indexed Skills (71)
a-value-you-did-not-measure-still-has-a-source
Use at Stage 06 and Stage 07 when a deliverable the task named cannot be produced by this run at all — a wet-lab measurement, a synthesised material, a proprietary benchmark, hardware you do not have. Covers the difference between fabricating a number and citing one, where the cited value belongs, and why omitting the section is the worst of the three options.
the-attribution-is-the-deliverable
Use when the task statement names interpretability, explainability, feature importance, saliency or attribution among its outputs or objectives. The graded artifact is then the attribution map itself — per input unit, by the field's standard estimator, drawn as a figure — not a diagnostic about the model's internals and not an argument that the model is uninterpretable.
train-the-named-architecture
Use at study design, implementation and experimentation when the brief's deliverable is a model you have to build — it names an architecture family (graph network, autoencoder, diffusion module, surrogate net) or a training regime (pre-training, fine-tuning, self-supervised, inverse design). Covers why a cheaper model class scores near zero however well it performs, why a scaled-down run of the named architecture beats a released checkpoint on every architecture criterion, and what to ablate.
astronomy-error-budget-is-the-audit-trail
Use at analysis and writing when a result rests on a fit or a calibration chain and you are about to quote it with a single uncertainty. Covers itemising the error budget term by term, keeping the fit's own bookkeeping visible, and why the audit trail is the result a referee checks first.
astronomy-figure-is-the-unit-of-result
Use at study design when choosing the figure list, and again before writing, when a result is about to be reported as a pooled number or a table. Covers why the figure is the unit a result is delivered in here, which panels a paper of this kind is expected to carry, and what a pooled number hides.
chemistry-canonical-units-thresholds-incumbent
Use at study design and analysis when a chemistry result is about to be reported in the units your code happens to produce, or without the program the field already uses. Covers anchoring to the incumbent, converting to the canonical unit, and turning an error distribution into a threshold success rate.
chemistry-ranked-entities-and-property-curves
Use at analysis and figure planning when the computation ranks entities — molecules, poses, fragments, atoms — or sweeps a property along a coordinate. Covers printing the named ranked list and the property-versus-coordinate curve, the two artifacts most often computed here and least often reported.
citation-discipline
Use when adding, verifying or cleaning citations and BibTeX entries in Stage 07 (Writing), when a reference cannot be resolved cleanly from DBLP or CrossRef, when checking that a cited paper actually supports the claim attributed to it, or when filling citation_verification.json.
close-the-gap-to-the-published-number
Use at Stage 05 and Stage 06 the moment a reproduction lands materially off a number the source study published — a different order of magnitude, an inverted trend, a collapsed estimate. Covers why the gap is a defect in your pipeline until you have shown otherwise, how much of the remaining budget to spend closing it, and what to write when it will not close.
draw-the-source-figure-panel-for-panel
Use at study design when planning figures for a reproduction, replication or validation task, and again before the report is written. Covers deriving each panel's series list and axis ranges from the source's rendered figure, giving every source result a panel before your own hypotheses claim the slots, and printing the source's named constants as labelled values.
earth-comparator-set-lives-outside-the-supplied-archive
Use at study design when the supplied archive is about to become both the input and the thing you compare against. Covers where the comparison set has to come from, what a self-comparison cannot establish, and how to build a comparator from outside the shipped data.
earth-report-the-lattice-and-show-the-field
Use at analysis and figure planning when a geospatial or gridded result is about to be reported only as regional aggregates. Covers reporting the stratified lattice, showing the field the strata came from, and which map a study of this kind is expected to publish.
a-deliverable-is-not-an-instruction
Use at study design when listing the task's deliverables into report_plan.json, and again before writing when checking coverage. Covers how to tell a research deliverable from the harness's own operating instructions, why a padded list is worse than a short one, and what to write when a deliverable is genuinely out of reach.
a-detection-score-is-a-claim-about-its-population
Use at study design and at analysis whenever a detection or ranking score (area under a precision-recall or ROC curve, recall at a fixed precision) is about to be compared against another study's number, or when one arm detects items the other misses. Covers publishing the population beside the score, the prevalence ladder to run when the source never states its own, and the characterisation of the extra detections that needs no annotation.
a-model-you-can-audit-is-not-a-model-that-scores
Use at study design and implementation when choosing between a method you can validate quickly and a stronger one you are not sure you can afford. Covers pricing the expensive method with a measurement instead of an impression, the go/no-go that has to be written before the clock is spent, and why the safe choice is only safe on the axes nobody is grading.
a-null-test-bounds-the-instrument-not-the-answer
Use at analysis and again at writing whenever you run a permutation, shuffle, placebo, unforced-control or power test against your own headline result, especially when it comes back saying the result is not distinguishable from noise. Covers giving every condition the task names its own value line and stating the relation across them as a result, keeping the estimate and the bound as two results with two different subjects, and the sentence order that stops a bound replacing the answer.
a-refutation-banner-over-a-confirming-panel
Use at hypothesis freeze, and again after the last revision pass, when a run has found that the supplied data or your own reproduction disagrees with the source it names. Covers the branch-name test that stops an agreement being printed as a refutation, and the enumerate-and-search sweep that keeps fidelity verdicts and internal labels out of the title, the headings and the figure banners.
a-scoreable-file-in-the-first-hour
Use at the first stage of a run whose deliverable is a predictions file, and again at every stage when one still does not exist. Covers why a trivial submission written early dominates a good one written late, what the first version should contain, and how to improve it in place without ever leaving it invalid.
a-supplied-parameter-file-is-a-list-of-questions
Use at study design, analysis and writing when the task ships a small file of named constants, ranges, entity tables, case lists or run settings. Covers treating each entry as a question your run must answer in the file's own labels and units — including entries your own audit shows are wrong — and why an agreement count is not an answer.
assume-this-stage-is-the-last-one-you-get
Use at the first three stages of a run under a hard wall clock, when the plan defers the modelling to a later stage. Covers the measured probability that the later stages never execute, why deferring to the stage designed for the work is the most expensive available choice, and what each early stage should leave behind if it turns out to be the last one to run.
astronomy-sample-the-published-table-into-chains
Use at study design and implementation when the source's constraints reach you as a table of best-fit values with 1-sigma errors for two or more models and no posterior samples were released. Covers rebuilding the ensemble that table describes, writing it in the layout this field's posterior tools read, and what to state about the construction so it is evidence rather than decoration.
astronomy-the-caption-is-the-figure-specification
Use at literature survey the moment you have the source's full text, and again when the plotting code is written, whenever you are reproducing a figure whose rendering you cannot see. Covers mining each numbered caption for the panel order, the series and their colours, the reference model the residuals are taken against and how the data were normalised, and holding those fixed against a later stage that finds a better choice.
astronomy-the-joint-posterior-is-the-parameter-result
Use at study design when the figure slate is fixed, and again at analysis and writing, when the deliverable is constraints or posterior distributions on parameters for two or more competing models. Covers why one overlaid triangle plot over the source's own parameter list is the exhibit that answers it, which parameters get an axis, and why a row of one-dimensional error bars reads as the figure never having been drawn.
chemistry-a-cut-variant-takes-its-analyses-with-it
Use at literature survey, at study design, at every descope decision and again at writing, when the source's method is a family — the same module dropped into two or more backbones, or one architecture published in several named variants — and you are about to run only one of them. Covers listing which of the source's downstream analyses were produced from which variant before any of them is cut, shrinking a variant rather than deleting it, and what a saliency map, case study or ablation computed on the surviving variant is and is not evidence for.
chemistry-ablations-and-curves-without-an-accelerator
Use at study design after you have priced a scaled-down training arm and found the machine cannot carry it — no accelerator visible, or no wall clock for one arm. Covers the one-row-per-named-component table with the inference switch that removes each part, why an input ablation does not answer a component criterion, and the ladder of curves that still ships when nothing can be trained.
chemistry-accuracy-and-cost-for-every-module-you-swap-in
Use at study design, through experimentation and again at analysis when the method under test is a drop-in replacement for a standard layer — a different basis, kernel, activation family or transform — and the source claims the replacement is both more accurate and cheaper. Covers giving every alternative module a cell in the accuracy column and in the cost column, fixing one matching convention across both, and dividing the runtime by the invariant already sitting in your own results file before you publish a contradiction of the source's ratio.
chemistry-fill-every-row-of-the-comparator-table
Use at literature stage and study design when the source publishes performance broken out by class of system, and you are about to choose your evaluation panel with a filter written for throughput. Covers transcribing the table as rows, auditing the inclusion filter against those rows before it is frozen, and buying one target per row before a second target for any row.
chemistry-group-attribution-over-the-split-and-the-baseline-mode
Use at study design, experimentation and analysis when the deliverable includes which substructures, functional groups or motifs drive the model's predictions, once the attribution estimator is already chosen. Covers widening from the one molecule the source drew to the whole evaluation split with per-molecule normalisation, running the identical attribution on the comparator model so a claim of better interpretability becomes measurable, and treating a learned edge or subgraph mask as a first-class output.
chemistry-interaction-inventory-of-the-modelled-complex
Use at study design, analysis and writing when the result is a modelled or predicted molecular complex — a docked pose, a co-folded assembly, a binding interface — and RMSD, DockQ or lDDT is about to be the whole answer. Covers the reference-versus-prediction contact inventory per interaction class, the pocket-cropped figure with the interactions drawn, and the mechanism sentence.
chemistry-reproduce-the-scoring-path-before-you-replace-it
Use at implementation, experimentation and analysis when you are reproducing a published benchmark number and the source's scoring path is one you can read — which rows are scored, in what order, how many the loader drops, which epoch is reported, how tasks are pooled, over how many seeds. Covers implementing that path exactly before improving it, the one-row-per-step ladder from the published rule down to your own honest estimate, and why one un-replicated step makes the reproduction gap you report uninterpretable.
claims-before-harness-forensics
Use at hypothesis generation and study design on reproduction and method-evaluation tasks, once close reading of the release has turned up defects, ambiguities or under-specification, and again when ordering the report. Covers labelling every planned experiment as a test of a claim or a test of self-consistency, the count gate that follows, and where reproduction-fidelity statistics belong.
decide-the-input-or-the-deadline-decides-it
Use at study design, and again at every stage boundary after it, when the run's own notes still carry an open question about which file or which system one of the named experiments will run on - the shipped stand-in, the authors' release, or one you generate from the Methods. Covers writing the default outcome beside every open question, ranking the list by that default rather than by difficulty, and the three-route ladder for a system the task did not ship.
disclose-by-construction-not-by-absence
Use at analysis when figures are rendered and at writing when they are captioned, and whenever an internal review asks you to disclose something you could not do. Covers why a disclaimer drawn inside a figure's axes destroys the result it annotates, the single location a caveat is stated in and what counts as a second copy, and how to describe the substitute you built instead of the gap you had.
do-not-grade-your-own-result-down
Use when drafting limitations, the discussion or the abstract, and any time you are about to call your own result unimproved, inconclusive or unverifiable. Covers the hedge that contradicts the run's own decision record, and the check a caveat has to fail before it is published.
draw-the-system-not-your-study
Use at study design when allocating figure slots, and again at analysis and writing, on tasks where the source's own rendered figures are not available to copy. Covers the four slots reserved for the system before any hypothesis claims one, drawing the loaded arrays instead of their counts, why a panel that reports a shortfall is not the panel carrying the result, and why a deliverable marked covered_by an artifact path is not covered.
earth-a-verified-answer-key-does-not-change-the-question
Use at literature survey when you find that the supplied archive holds the source study's published output as well as its raw inputs, and again at hypothesis generation before anything is frozen. Covers what to do with the source's headline numbers in the hour you first recompute them, why a confirmed answer pulls a run into auditing the method that produced it, and the ordering rule that keeps the critique behind the delivered product.
earth-shape-of-the-record-and-share-of-the-budget
Use at analysis and again at writing when the deliverable is a multi-year record - an annual time series, a reconstruction, a trajectory - and you are about to report it as a period mean, a cumulative total and a validation residual against a reference version of itself. Covers the change over the record's own length, the extremes, the per-unit split, and the record's share of the budget it is one term of.
earth-the-technique-grid-needs-values-in-its-cells
Use at study design when the figure list is chosen, at implementation before any pipeline code is written, and again at analysis, when the supplied archive is stratified by measurement technique or instrument and you are about to describe that stratification. Covers reading the per-technique columns the archive already ships, drawing the technique grid with estimates in it rather than file counts, and why a no-peek rule about the reference product must not reach your figures.
earth-two-orderings-of-a-regional-decomposition
Use at study design when the figure and table plan is fixed, at analysis, and again at writing, whenever the deliverable splits a global or basin-wide total into per-region parts and you are about to report each part as an absolute rate. Covers the two normalisations every row owes and why their disagreement is the result, where the intensity denominator has to be captured before you need it, and what the spare cell of a small-multiple grid is for.
earth-window-mismatch-is-an-alignment-problem
Use at study design, and again when planning figures, whenever a comparator you hold - a model projection ensemble, a scenario run, a prior published assessment, a sibling record - is reported over a different period, baseline epoch, initial state or unit than your result, and you are deciding whether the comparison can be made at all. Covers re-baselining onto a common start date, plotting an ensemble that publishes only horizon endpoints, reading the crossing date, expressing prior assessments as revisions, and where a genuine refusal belongs.
the-unit-of-analysis
Use at analysis and figure planning when the brief names the units its data is grouped into — patients, cells, classes, labs, behaviours — and you are about to report one pooled number over all of them. Covers why the pooled number hides the result, which strata a study of this kind is expected to report, and when an aggregate is the right answer after all.
answer-the-why-not-only-the-what
Use when writing results and discussion, and when a task or a reviewer asks why an effect happens rather than whether it does. Covers the difference between reporting an effect and accounting for it, and what a mechanism claim needs behind it.
cover-what-the-task-named
Use at study design and again before writing, to check that every deliverable the task statement names has been produced. Covers how to enumerate what was asked for, why partial coverage scores worse than it feels, and what to do when a named deliverable is out of reach.
information-exhibit-the-intermediate-objects
Use at analysis and writing when a multi-stage pipeline is about to be reported by its end-to-end metric alone. Covers exhibiting each stage's intermediate object, and re-running the source's own demonstrations on the source's own inputs rather than on yours.
information-fill-the-whole-results-grid
Use at study design and again at writing when the source reports a grid — variants crossed with backbones, datasets or metrics — and you are about to fill part of it. Covers reproducing the whole grid at reduced N where you must, and why a labelled reduced-N cell beats an empty one.
latex-repair
Use when a LaTeX build fails or produces a broken PDF in Stage 07 (Writing) — undefined control sequences, missing style packages, unresolved citations or references, float placement blowing the page budget, or a build_log.txt full of errors you need to triage.
life-benchmark-against-the-incumbent
Use at study design when a life-science method result is about to be reported on its own numbers. Covers the head-to-head against the incumbent tool, the cost table that goes with it, and finding an orthogonal truth set the method was not fitted to.
life-full-study-skeleton-including-the-wet-lab-half
Use at study design and again when laying out the results section, to check every slot of a life-science study is filled. Covers the skeleton a paper of this kind carries, and what to put in the slots this run cannot compute rather than leaving them out.
material-as-specified-run-and-stage-diagnostics
Use at study design and implementation when a protocol is specified and you have found a reason to deviate, or when a pipeline stage is about to run without its conventional diagnostic. Covers running the protocol as specified as the foreground result, and leaving every stage's default panel behind you.
material-landmark-scalars-in-physical-units
Use at analysis when a materials result exists as a curve, a distribution or a trajectory and is about to be reported as one. Covers extracting the landmark scalar a reader compares — peak position, transition temperature, barrier height — in the property's physical unit, against a reference value.
math-canonical-curve-on-the-cost-counter
Use at figure planning when a convergence or performance curve is about to be drawn against wall-clock, or folded into a composite panel. Covers the field's plain two-curve figure, plotting against the algorithm's own cost counter, and why it comes before any richer diagnostic.
math-equal-effort-baselines-and-knob-sweeps
Use at study design when the source names competing algorithms and they are about to become a related-work paragraph instead of arms. Covers running every named baseline at equal tuning effort, and sweeping the parameter you claim credit for.
mine-the-papers-you-were-given
Use when the task ships PDFs in related_work/, at literature stage and before the study plan is costed. Covers reading those papers for the named tools, benchmarks, events and metrics the work will be judged against — as a work list rather than as background — and what to record for each one.
neuroscience-comparator-ladder-and-per-unit-predictions
Use at study design and analysis when a model is about to be compared against one alternative, or a fit reported without a negative control. Covers the two-sided comparator ladder, the control representation panel, and splitting per-unit predictions into the ones a measurement validates and the ones that stay predictions.
neuroscience-stratify-and-report-detection-metrics
Use at analysis when a detection or classification result is about to be reported as one accuracy over a pooled population. Covers per-group and per-class precision, recall and confusion matrices at a stated threshold, and sweeping the degradations the recording modality actually suffers.
paper-writing
Use when drafting, structuring or revising the manuscript or report in Stage 07 (Writing) — shaping the contribution into one story, writing the abstract and introduction, fixing prose that reads generic or templated, ordering sentences for clarity, or deciding what Figure 1 should show.
physics-discriminate-model-families-and-defend-the-fit
Use at analysis when a fit is about to be reported as the answer without a rival model being excluded. Covers naming the competing model families, showing which the data rules out, and treating the fit protocol — range, weighting, priors — as part of the result rather than as a setting.
physics-two-estimators-propagation-and-a-forward-model
Use at study design and analysis when a physical quantity is about to be reported from one estimator, or an uncertainty quoted without propagation. Covers measuring it a second independent way, propagating the error through the chain, and generating the observable forward from the fitted model to check it.
publish-what-the-run-already-computed
Use at Stage 06 and again before the report is finalised, when deciding which of the run's results enter the deliverable. Sweeps the run's own outputs for quantities it computed and never published, and covers the three shapes that sweep finds — the diagnostic never persisted, the column requested and dropped, the feasibility measurement discarded — and what to promote out of an appendix.
reproducibility-check
Use in Stage 08 (Dissemination) when assembling the release or submission bundle — auditing whether the run's code, data, results and figures are actually reproducible by someone else, writing the readiness checklist and threats-to-validity notes, or deciding what has to be disclosed as not verified.
result-table
Use when turning measured results into a table or figure for the paper — building a LaTeX or markdown results table from workspace/results/*.json, deciding what uncertainty to report, choosing which baselines and ablations belong in the main table, or writing a caption that stands alone.
run-the-conditions-the-source-ran
Use at study design, before any experiment of your own is costed, on reproduction and method-evaluation tasks. Covers enumerating the systems, scenarios, stress sweeps and case studies the source names, running each one by name, measuring the preconditions the method declares it needs, and what to do when one of them fails.
the-canonical-figure
Use when planning figures, at study design and again before writing. Covers the figures a paper in this field is expected to contain, why an original figure does not substitute for a standard one, and how to decide what to draw first.
the-reproduction-is-a-hypothesis
Use at Stage 02 and Stage 03 whenever the task is to reproduce, re-implement or verify a published study and the hypotheses you are drafting are all about something else. Covers how to write the reproduction itself as a falsifiable frozen commitment, why a self-invented question crowds it out, and how to budget between the two.
the-supplied-item-is-the-graded-unit
Use at study design whenever the task ships a specific named object in data/ — one paper, one structure, one instance — and again before writing. Covers reporting that item's own numbers under its own name, choosing the worked example by the task's pointer rather than by your result, and how to widen scope without dropping it.
use-the-sources-own-names
Use at Stage 06 and Stage 07 when writing up a reproduction, and any time you have given a reproduced quantity, equation, figure or sequence a name of your own. Covers why a correct reproduction under private names reads as a missing one, which names have to be carried, and where they have to appear.
venue-checklist
Use when the target venue's submission requirements matter — checking a draft against NeurIPS, ICML or ICLR expectations, deciding which required sections (checklist, broader impact, reproducibility, LLM disclosure) the paper needs, or running the Stage 08 submission-readiness review.
evidence-not-assertion
Use whenever a number, a comparison or a claim is about to enter a stage summary or the report — at analysis and writing, and any time you are tempted to state a value you have not computed in this run. Covers where a number must come from, what to do when the experiment did not run, and why an honest gap outscores a plausible sentence.
record-what-you-learned
Use when a run is finished and the report is written, after the report is written, to record one reusable lesson for the next run in this field. Covers what counts as a lesson worth passing on, what must never be passed on, and how to write it.
reproduce-then-extend
Use when the task is to reproduce, replicate or re-implement a published study, at design time and when reporting results. Covers what a reproduction must report, how to compare against the source study's numbers, and why the reproduction comes before any improvement.
run-the-requested-analysis
Use when the supplied data looks synthetic, degraded, incomplete or wrong, and whenever you are tempted to reframe the study around what you found about the inputs, the harness or the evaluation. Covers what to do with a real data problem without losing the study.
Bio shown is the top-scored skill's repo description as a fallback — real GitHub bios land in a future update.