Private beta open for teams shipping AI agents.·See it catch a reward-hack

Trust & privacy

Review agent updates without exposing what's private.

The gate runs in your trust domain. Built for security review — redacted evidence and private evaluation boundaries by default.

Customer code in telemetry
0
Model calls per scan
0
Signed assurance certificates
Ed25519

The principle

Private by default. Evidence when needed.

Security review shouldn't mean handing over your data. The gate runs in your trust domain and reviews agent updates without exposing customer data, private evals, hidden cases, gold answers, raw traces, or secrets.

Redacted evidence

A reviewable record that reveals scores, not secrets

Every run produces an assurance card — the decision, machine reasons, and score deltas. The reference code, tests, and traces that produced them never appear.

  • Decision, reasons, and score deltas
  • Engine verdicts and the policy that was applied
  • Redacted by default on every run
  • Assurance certificates are Ed25519-signed and verifiable against our public JWKS

Assurance Card · example

decision
BLOCK
reasons
ood_regressed · hidden_regressed
public
0.740 → 0.910
hidden
0.732 → 0.611
ood
0.701 → 0.488
record
redacted · reviewable

✗ candidate not promoted

Runs in your trust domain

Reference code and tests stay in your trust domain

Audit telemetry posts only scalar counts — candidate code and graders stay with you. The IPT re-grade runs in your own loop, so reference code and tests never leave your trust domain. LangSmith import is a client-side push with no outbound fetch and no stored credentials. Private deployment paths are available for pilots.

  • No customer data in results or demos
  • Hidden cases and gold answers not exposed, by design
  • Imports are one-way: your client pushes to us; our systems make no outbound fetches into your network
Private evaluation boundary

What an assurance card discloses

  • The gate decision — ship, limit, or block
  • Machine reasons behind the decision
  • Score deltas (scores, not answers)
  • Contamination & anti-hack engine verdicts
  • The gate policy that was applied

What it never contains

  • Customer data
  • Hidden evaluation cases
  • Gold answers
  • Raw model traces
  • Secrets or credentials

No third-party certification yet — ask us for our current security posture.

Responsible disclosure

Report a vulnerability

Found a security issue — in our website, our API, or the evaluation pipeline itself (prompt injection, reward hacking, result tampering)? Email security@verifiable-labs.com. We review good-faith reports, acknowledge them, and coordinate a fix — please allow reasonable time before public disclosure.

security.txt

Improve what fails. Ship what holds.

Bring a baseline and a candidate.