Define the behavior you want. We measure it, then train it in.
Zero Proof Labs reads your agent’s traces and scores them on metrics you define. Our framework finds which behaviors actually move your outcome, then simulates and trains until the model has them, before you spend a dollar on GPUs.
Safety judge for disputes and refunds
1B params
“There’s an $81.40 charge from Amazon I never made, I want it refunded.”
How it works
Send your traces. You name the behaviors that matter; we score every run against them.
We rebuild the workflow as an environment and find which behaviors actually move your outcome.
Only then do you spend a dollar building those behaviors into the model.
Proven on e-commerce. The safety judge below came out of this loop, with data, adapters, eval set, and serving cost public.
Worked example
We proved the framework on refunds and disputes, where a wrong decision costs a chargeback, and published every artifact. Rerun the evaluation against our numbers.
Accuracy against cost to serve
Near-frontier accuracy at about one hundredth of the cost.
Macro intent-type accuracy (%); zeroproof-ecommerce-1b on the full 1,977-row held-out set, frontier models on the balanced comparison subset. Frontier cost at published list prices; Zero Proof Labs cost from measured throughput on a single L4. The dashed line is one pass of post-training on the same base at the same serving cost.
The same method trains zeroproof-ecommerce-0.5b to within five points at half the size.