The benchmark library
Find the right test for enterprise work, architecture and CAD.
One model. A fixed recipe. An unseen exam.
Measure what your data changes.
EXPLORE THE GROUNDS
Start with the public specifications,
or enter your private workspace.
Find the right test for enterprise work, architecture and CAD.
A goal, a working state and a verifier. See how a task is defined.
Make a precise revision. Keep every protected feature intact.
Enter the private workspace and inspect the grader on each attempt.
Compare a trained model with a control on unseen work.
Follow the research workflow, from a collection to its evidence report.
Δ =trained on your work−trained on a control
Fix the model, recipe and exam. Change the training data. Measure the difference in what the model can do.
See our recorded procurement checkpoint →Does every task build and reset the same way each time?
Is each task solvable, and is doing nothing worth zero?
Does high reward mean the task was really solved?
Is there something left for the model to learn?
Does reward climb during training?
Is the model better at held-out work than a control?
Our live CAD development exam asks for a precise revision to a generated mounting plate. Here is the first task, checked by the same grader used in the workspace.
Increase H1 to an 8 mm diameter. Preserve the plate and every other hole, including all centers. All holes must retain at least 2 mm of edge and inter-hole clearance.
| Grader check | Unchanged | Correct edit |
|---|---|---|
| Feature identity | Pass | Pass |
| Requested edit | Fail | Pass |
| Protected geometry | Pass | Pass |
| Analytical validity | Pass | Pass |