Tutorials
    Chapter 1 of 3

    Your First Braintrust Evaluation

    24-second evaluation loop

    Turn chatbot failures into permanent tests

    A production chatbot failure becomes a versioned test case, receives deterministic and semantic scores, and enters a repeatable release gate.

    Read the visual transcript
    1. 01A support chatbot incorrectly promises that a customs fee is always refunded.
    2. 02Capture that production failure as one versioned dataset case with the expected policy behavior.
    3. 03Score hard rules with code and semantic quality with a calibrated judge or human review.
    4. 04Release on evidence, then make every meaningful production failure a permanent regression test.

    Begin with one small support bot

    Imagine a shoe-store chatbot with four rules: unworn items can be returned within 30 days, original shipping is not refundable, damaged items go to a specialist, and unanswered policy questions must not be guessed.

    An evaluation asks the bot the same five questions before and after a change. Each answer gets a clear pass or fail. That is all you need to understand before using Braintrust.

    Braintrust gives the first evaluation a small mental model: data is the questions, the task is your chatbot, and scorers check the answers.

    What Braintrust calls the parts

    A Braintrust Eval needs three things: data, a task, and scores. Keep the first scores binary—correct policy, no invented promise, and correct handoff—so every failure is easy to explain.

    The workflow

    Provide five data rows → call the chatbot as the task → run simple scorers → inspect each result → compare with a previous experiment.

    The five-step evaluation loop
    1
    Save
    Write five real customer questions and what a passing answer must say.
    2
    Run
    Ask the current chatbot all five questions and save its answers.
    3
    Check
    Mark each answer correct, invented, or missing a required handoff.
    4
    Change
    Change one prompt, model, or retrieval setting—not several at once.
    5
    Compare
    Run the same five questions and inspect every case that became worse.

    Interactive test / Braintrust

    Try three support-chatbot cases in Braintrust

    Provide the five questions and expected answers as the evaluation data.

    Answer Northstar Shoes customers using only the return policy.

    Input

    Can I return unworn shoes after 20 days?

    Expected behavior

    Say yes and mention the 30-day window

    Observed output

    Yes. Unworn shoes can be returned within 30 days of delivery.

    correct answerPass
    no made-up policyPass
    useful next stepPass
    Case passed3/3 checks
    Key Takeaways
    • -Start with five understandable questions
    • -Write the pass rule before running the chatbot
    • -Keep the questions the same when comparing versions
    Knowledge Check

    How should you compare the current chatbot with a new version?