Skip to content
Delta Grounds

Task package spec

A task package is one problem a model can attempt, plus everything needed to grade it without a human. Delta Grounds reads the common open package layout (task, environment, verifier, reference solution) unchanged, so existing packages import as they are, and adds an optional assay block for enterprise metadata.

Package layout

invoice-match-0412/
├─ task.md                    frontmatter + "## prompt"
├─ environment/
│  ├─ Dockerfile              the sandbox image
│  └─ seed/inputs.json        copied into the workspace
├─ verifier/
│  ├─ test.sh                 runs pytest, writes reward.txt and reward.json
│  ├─ test_outputs.py         the checks
│  ├─ verifier.md             strategy and rubric
│  └─ data/expected.json      private ground truth, uploaded after the attempt
└─ oracle/
   └─ solve.sh                reference solution

The workspace is $BENCHFLOW_WORKSPACE (default /root); Delta Grounds also sets ASSAY_WORKSPACE to the same path. The verifier directory is never visible during the attempt.

task.md frontmatter

---
version: "1.0"
metadata:
  author_name: Acme Finance Ops
  author_email: finops@acme.example
  category: data-processing        # one of 18 standard categories
  license: Apache-2.0              # SPDX identifier
  origin: original                 # original | adapted | generated
  difficulty: medium               # easy | medium | hard
  tags: [accounts-payable, three-way-match]
agent:
  timeout_sec: 900
verifier:
  timeout_sec: 180
environment:
  build_timeout_sec: 600
  cpus: 1
  memory_mb: 2048
  storage_mb: 10240
  allow_internet: false
---

The assay block

assay:
  domain: finance-ops              # finance-ops, support-crm, legal-compliance, hr-people,
                                   # procurement-supply, data-bi, it-ops, sales-revops
  family: invoice-three-way-match  # tasks in a family share one skill
  harness: program                 # answer | program | agent
  workspace_files: [inputs.json]   # inlined into the prompt for single-turn harnesses
  outputs: [answer.json]           # files the verifier reads
  split_hint: train                # advisory; project suites decide

Families matter for proof. The transfer suite is made of families no training task belongs to. If every task were its own family, held-out improvement could only ever mean memorization.

Writing the prompt

Put the instructions under ## prompt. Name every input path and the exact output path and schema. The verifier is mechanical, so an ambiguous prompt produces ambiguous grades. Write the business rule the way the policy document states it, including ordering and tie-breaks.

Never put the answer, or anything that determines it without doing the work, in the prompt or in the environment.

Environment

  • Start from a small pinned base image and preinstall pytest==8.4.1 and pytest-json-ctrf==0.3.5.
  • Pin every verifier dependency, so grading works with the network off.
  • Copy seed data only. Never copy the verifier, expected outputs or the oracle into the image.
  • Generate seed data deterministically from a seed, so reruns start from the same state.

Verifier

test.sh runs test_outputs.py with pytest and writes /logs/verifier/reward.txt (1.0 if every test passes, else 0.0), reward.json and a CTRF report. Delta Grounds also records partial credit as the share of tests passed.

A good verifier:

  • recomputes the expected answer from its private copy of the seed data instead of hard-coding it;
  • checks values, not just that an output file exists;
  • is deterministic: no clocks, randomness, network or model calls;
  • scores the oracle 1.0 and an untouched workspace 0.0;
  • rejects plausible-but-wrong outputs. Delta Grounds measures this with mutation testing: it perturbs the oracle output (a changed amount, a flipped flag, a dropped row) and counts how often the verifier notices.

Oracle

oracle/solve.sh solves the task in the same sandbox the model gets. It must score 1.0 on all 8 reruns. Compute the answer from the seed data and keep it simple: it documents the intended method and seeds the mutation tests.

Harnesses

HarnessThe model receivesThe model returnsUse for
answerPrompt plus inlined workspace filesOne or more <file path=…> blocksShort structured answers
programPrompt plus inlined workspace filesOne Python program, run once in the sandboxTraining (default): single-turn and procedural
agentPrompt and tools: list, read, write, runTool calls over many turns, then submitEvaluating API models and long workflows

Quality gates

SeverityTriggersEffect
BlockTask name matches a sealed task; a prompt shares half its 13-grams with a sealed taskThe whole collection fails validation
RejectImage copies verifier data or answers; verifier reads answer-like files; any shared 13-gram with a sealed taskThe task is excluded
ControlNo working oracleEligible only if a no-op scores 0 on 8 reruns and the base model solves it at least once
ReviewExistence-only assertions; remote ADD; oracle downloads; caches or .git in the imageAdvisory, shown to the reviewer

Dynamic gates run in the sandbox: the image builds and the oracle scores 1.0 on 8 reruns, an untouched workspace scores 0.0 on 8 reruns, and the base model solves the task 1 to 3 times out of 4.

Collections

my-collection/
├─ submission.yaml     team_name, contact_email, track: environments
└─ envs/
   ├─ invoice-match-0412/
   └─ …                1 to 200 packages

Upload a collection as a tar.gz, point at a public Hugging Face dataset or GitHub repository, or create one from the Task Builder. Delta Grounds pins the revision, runs the gates, and then it can be trained against a challenge in the Arena.