Cookbook
From one spreadsheet of past decisions to an evidence report. Every step has a console equivalent; the commands are here because they are exact. For the package format, read the spec.
The commands and outputs below illustrate the research workflow. They are not published run results. Cloud training and the private enterprise engine are not connected to this hosted site.
- 1
Install the CLI and sign in
The CLI talks to the same API as the console. Sign-in opens your browser and stores a token in ~/.assay/config.toml with mode 600.
$ pipx install assay-cli$ assay login
signed in as jahaan (acme)
- 2
Create a project
A project is one model program: a base model, a data policy, suites, runs and evidence. Private projects only use open-weight models on your Modal account.
$ assay init acme/ap-automation --base hf:Qwen/Qwen3-4B-Instruct-2507 --policy privatewrote assay.toml · project acme/ap-automation
- 3
Bring a workflow
Upload labeled rows from a real process. The builder renders one task per row (or per batch), keeps the label as private ground truth, masks PII columns and plans a stratified train and held-out split.
$ assay data add ./invoices-q3.csv --kind raw$ assay builder create --dataset invoices-q3 --template three-way-match --label expected_action --split 240/60
environment invoice-three-way-match · 300 tasks · checks: 4 ok, 1 warn
- 4
Validate the environment
Every task is replayed in a Modal sandbox: oracle 8 times, no-op 8 times, verifier agreement, mutation testing and the 13-gram leak check against every eval suite.
$ assay env validate invoice-three-way-match --on modal-sandboxok oracle.reruns 300/300 scored 1.0 on 8 of 8 warn verifier.flake 1 task disagreed on 2 of 8 ok mutation.kill_rate 94.2% of perturbed outputs rejected
- 5
Profile difficulty and red-team
The base model attempts each task four times to find the learnable band. A frontier model then tries to earn reward without solving; anything that works goes to review.
$ assay env profile invoice-three-way-match --model hf:Qwen/Qwen3-4B-Instruct-2507 --attempts 4$ assay env redteam invoice-three-way-match --model anthropic:claude-sonnet-5-5
- 6
Set the suites
Train on some families, hold out unseen instances, hold out whole families for transfer, and keep an unrelated regression suite. Delta Grounds refuses overlapping suites.
$ assay suites set --train env:invoice-three-way-match:train --heldout env:invoice-three-way-match:heldout \ --transfer env:accrual-schedule,env:fx-revaluation --regression env:general-regression
- 7
Run a proof
A proof run evaluates the base model, trains the treatment and a matched control with the same recipe, evaluates both, and scores Δ with a paired bootstrap.
$ assay proof --recipe grpo --warm-start distill-sft --teacher anthropic:claude-sonnet-5-5 \ --control generic --steps 240 --seeds 3 --on modal-h100
run_8c2d41f0a6b3 queued · follow with: assay runs watch run_8c2d41f0a6b3
- 8
Read the evidence
The report grades each hallmark and states the verdict in a sentence. Share the link or export a PDF certificate.
$ assay evidence show --latestAS-2026-0417 proven L6 transfers +16.9 pp vs control (95% CI 8.4 to 25.3) · regression within noise
- 9
Enter the Arena (optional)
Submit the collection to a challenge. The organizer's sealed suite and recipe are fixed; your collection is ranked by Δ after review.
$ assay arena submit --challenge backoffice-4b --collection acme/ap-automation@v3