New/zeroproof-ecommerce-1b, a real-time safety judge for disputes and refunds. Read the release

Behavioral Trust Infrastructure for AI

Agents are only as trustworthy as the data and evals behind them.

Zero Proof Labs simulates your domain in milliseconds. Our optimization framework then tunes the simulation to maximize learning efficiency, cut GPU hours, and produce higher quality evals, before you spend a dollar on model training or harness optimization.

Safety judge for disputes and refunds

1B params

“There’s an $81.40 charge from Amazon I never made, I want it refunded.”

Dispute type
reverse / refund
Amount
$81.40
Grounded in evidence
yes
Confidence
0.85
Matches the requestsha256:7f3a…1b2
Built with the framework: simulated dispute cases, measured for signal before post-training, then trained on a 1B base. It judges every refund before it reaches the payment rails.

How it works

Simulate, optimize, then train.

  1. 01Simulate

    You bring the workflow. We build it as an environment and simulate it in milliseconds.

  2. 02Optimize

    Our optimization framework tunes the simulation to maximize learning efficiency, cut GPU hours, and produce higher quality evals.

  3. 03Train

    Only then do you spend a dollar on model training or harness optimization.

Proven on e-commerce. The safety judge below came out of this loop, with data, adapters, eval set, and serving cost public.

Worked example

How we built a safety judge with our framework.

We proved the framework on refunds and disputes, where a wrong decision costs a chargeback, and published every artifact. Rerun the evaluation against our numbers.

Eval and train set
17,933 train, 1,977 held out
Balanced, labeled on the decision
Post-training
Rank-16 QLoRA, two bases
A 1B and a 0.5B open model
Serving
One L4, vLLM
OpenAI-compatible endpoint
Result
75.3% macro accuracy
$0.18 per 1M output tokens

Accuracy against cost to serve

Near-frontier accuracy at about one hundredth of the cost.

1007550250$0.10$1$10$ per 1M output tokens, log scalebase, before post-trainingGPT-5Sonnet 5Opus 4.8zeroproof-ecommerce-1b$0.18 per 1M

Macro intent-type accuracy (%); zeroproof-ecommerce-1b on the full 1,977-row held-out set, frontier models on the balanced comparison subset. Frontier cost at published list prices; Zero Proof Labs cost from measured throughput on a single L4. The dashed line is one pass of post-training on the same base at the same serving cost.

The same method trains zeroproof-ecommerce-0.5b to within five points at half the size.