Skip to content
MouseCat
Platform · Evaluation and Tuning

Evaluation and Tuning

Backtest against your own history, score against your ground truth, and compare versions before and after you ship them.

Backtesting.

A backtest runs an investigation with the clock pinned to a point in the past. The cutoff is enforced everywhere — every agent and every tool sees only the data that existed at that moment — and everything else about the run is identical to production.

Point it at last quarter's cases to see what MouseCat would have done on your own book.

as ofwithheldearlierlater
agent toolthe cutoff is threaded through every one of them

Scoring.

Score agent output against your ground truth — analyst annotations, chargebacks, dispute outcomes — with precision, recall, and confusion matrices out of the box.

The metrics interface is extensible, so you can evaluate against whatever your team actually measures.

Version Comparison and A/B Testing.

labeled set3 versionsverdict
block review allow versions diverge
  1. 01

    Compare offline

    Run up to three versions of an agent against the same labeled set and compare them directly, including where they agree and disagree case by case.

  2. 02

    Take it to production

    Run multiple versions live, orchestrate traffic between them, and measure the difference on real cases. Rolling back is a one-step operation.

Get started.

We'll walk you through MouseCat live.

Book a demo