DriftbenchSign in

Regression testing for prompt & model changes

Know exactly which cases broke before you ship the new prompt.

Save your test cases once. Paste in a baseline and a candidate output for each one. Driftbench classifies every case as pass, changed, or failed — no more scrolling a spreadsheet to guess what got worse.

No credit card. No teams. Just a faster way to catch regressions.

Illustrative examplegpt-4o-mini → gpt-4.1-mini
  • Refund policy question

    Cites the 30-day window in both.

    Pass
  • Angry customer, order lost

    Candidate is warmer, check still passes.

    Changed
  • Password reset steps

    Candidate dropped the reset-link mention.

    Failed

The flow

Three steps between a prompt change and a confident answer.

01

Build a suite of real cases

Name the case, add the prompt or scenario, and optionally an expected check — a phrase the output must or must not contain.

02

Run baseline vs. candidate

Paste in what the current prompt/model produces and what the new one produces for each case, side by side.

03

Read the regression report

Every case is scored pass, changed, or failed with the exact diff. Every run is saved so you can reopen it later.

Stop trusting your memory of what the old prompt used to say.

Sign up, and Driftbench can seed one example suite for you so you can see a full comparison before you write your own cases.