Shipworthy — the release gate for AI agents
Ship only what holds.
Shipworthy re-grades each change beyond its visible tests, audits the trajectory for reward hacking, and returns a ship / limit / block verdict.
A gate above your existing agent stack — not a rewrite of it.
Shipworthy — Agent Release Gate
Proof your agent actually improved — before it ships.
Any update can top the tests you can see. We verify the gain holds where you can't — and that it wasn't gamed.
RE-GRADE BEYOND VISIBLE TESTS
Run the candidate through Isomorphic Perturbation Testing (IPT) — unseen, equivalent variants of each task, no LLM judge involved. A win has to hold on the variant to count.
AUDIT THE AGENT'S TRAJECTORY
Scan the tool-call trace for documented reward-hack patterns — test tampering, grader patches, forced exits, gold-answer access. Deterministic rules, 0 model calls.
SEE WHAT FAILED
Every block returns machine reasons — the broken invariant and the gap that triggered it — not just an aggregate score.
SHIP ONLY WHAT HOLDS
A ship / limit / block verdict on every scan, with its error rate stated on the card.
ASSURANCE CARD
Demo
Catch a reward-hacker before it games your verifier.
Re-grading without an LLM judge catches what test-pass filters accept.
- No LLM judges — 0 model calls, $0 per scan, same verdict every run
- When the re-grade can't conclude, the verdict is LIMIT — held for human review, not a guess
- Catches hardcoded shortcuts that pass the visible tests — measured on harder cheat families too
- Its error rate is stated on every assurance card
A release workflow for every agent update.
REGISTER
Register the agent and point the gate at a baseline and a candidate version.
RE-GRADE
Re-check the candidate's wins on unseen variants of each task and scan its trajectory for tamper patterns — no LLM judge, 0 model calls, in your trust domain.
DIAGNOSE
See where the candidate breaks, with machine reasons — every applicable check composed on one redacted assurance card.
GATE
A ship / limit / block verdict on every scan, with its error rate stated on the card. A LIMIT on a high-risk agent routes to a named human for sign-off.
Built for teams shipping AI agents
Agent product teams
Gate every prompt, model, workflow, and guardrail change before it reaches users — with a clear ship or block call.
Watch demo →RL & platform teams
Drop a release gate with no LLM judge between your RLVR/RFT reward loop and production, so gamed checkpoints are caught before they promote.
Book a demo →Enterprise AI teams
A repeatable, auditable review for agent changes — without exposing private data, hidden evals, or gold answers.
Talk to us →Security
Private by default. Evidence when needed.
The gate runs in your trust domain — reference code, tests, and gold answers never leave it. Assurance cards carry the verdict and score deltas, never the material that produced them.
- Redacted assurance cards
- Private evaluation boundaries
- No customer data in public demos
- Designed for security review








