New/zeroproof-ecommerce-1b, a real-time safety judge for disputes and refunds. Read the release

Behavioral Trust Infrastructure for AI

Define the behavior you want. We measure it, then train it in.

Zero Proof Labs reads your agent’s traces and scores them on metrics you define. Our framework finds which behaviors actually move your outcome, then simulates and trains until the model has them, before you spend a dollar on GPUs.

Safety judge for disputes and refunds

1B params

“There’s an $81.40 charge from Amazon I never made, I want it refunded.”

Dispute type
reverse / refund
Amount
$81.40
Grounded in evidence
yes
Confidence
0.85
Matches the requestsha256:7f3a…1b2
Built with the framework: simulated dispute cases, measured for signal before post-training, then trained on a 1B base. It judges every refund before it reaches the payment rails.

How it works

Measure, simulate, then train.

  1. 01Measure

    Send your traces. You name the behaviors that matter; we score every run against them.

  2. 02Simulate

    We rebuild the workflow as an environment and find which behaviors actually move your outcome.

  3. 03Train

    Only then do you spend a dollar building those behaviors into the model.

Proven on e-commerce. The safety judge below came out of this loop, with data, adapters, eval set, and serving cost public.

Worked example

How we built a safety judge with our framework.

We proved the framework on refunds and disputes, where a wrong decision costs a chargeback, and published every artifact. Rerun the evaluation against our numbers.

Eval and train set
17,933 train, 1,977 held out
Balanced, labeled on the decision
Post-training
Rank-16 QLoRA, two bases
A 1B and a 0.5B open model
Serving
One L4, vLLM
OpenAI-compatible endpoint
Result
75.3% macro accuracy
$0.18 per 1M output tokens

Accuracy against cost to serve

Near-frontier accuracy at about one hundredth of the cost.

1007550250$0.10$1$10$ per 1M output tokens, log scalebase, before post-trainingGPT-5Sonnet 5Opus 4.8zeroproof-ecommerce-1b$0.18 per 1M

Macro intent-type accuracy (%); zeroproof-ecommerce-1b on the full 1,977-row held-out set, frontier models on the balanced comparison subset. Frontier cost at published list prices; Zero Proof Labs cost from measured throughput on a single L4. The dashed line is one pass of post-training on the same base at the same serving cost.

The same method trains zeroproof-ecommerce-0.5b to within five points at half the size.