Back to Blog
    RAG Evals

    How to Evaluate a Simple RAG Chatbot

    Test whether a RAG chatbot found the right passage, answered from that passage, cited it, and admitted when the answer was missing.

    Bhaulik Patel·Sep 7, 2026·5 min read

    Start with the simple chatbot evaluation guide if you have not created a small test set yet.

    A RAG chatbot searches your documents before it answers. That gives you two separate things to test:

    1. Did search find the right passage?
    2. Did the chatbot use that passage correctly?

    Keeping those questions separate is the difference between “the bot was wrong” and knowing what to fix.

    We will keep using the Northstar Shoes support bot. Its knowledge base contains three short documents:

    • Returns policy: unworn items, 30-day window, original shipping not refundable
    • Damaged items guide: send damaged-item cases to a specialist
    • International orders FAQ: duties and customs fees vary; contact support for order-specific help

    Step 1: Save the evidence each question needs

    Start with five cases:

    QuestionPassage that must be foundA passing answer must
    Can I return unworn shoes after 20 days?Returns policySay yes, mention 30 days, cite the policy
    Is original shipping refunded?Returns policySay no, cite the policy
    My shoes arrived damagedDamaged items guideOffer a specialist handoff
    Are customs fees refunded?International orders FAQSay it varies and offer support
    Do you repair shoes?NoneSay the documents do not answer this

    The expected answer does not need to be a perfect paragraph. Save the facts that must appear, the claims that must not appear, and the source the chatbot should cite.

    Step 2: Check retrieval first

    For each question, record the passages returned by search.

    Question: Is original shipping refunded?

    Expected passage: Returns policy

    Retrieved passages: Returns policy, Size guide, Store locations

    Retrieval passes because the required passage is present. If the returns policy is missing, fix search or document indexing first. Rewriting the answer prompt cannot help the chatbot use evidence it never received.

    For a beginner test, use one yes-or-no retrieval check:

    Did the search results include the passage needed to answer the question?

    Step 3: Check the answer second

    When the right passage was retrieved, mark four checks:

    1. Correct: Does the answer match the passage?
    2. Supported: Can every important claim be found in the retrieved text?
    3. Cited: Does the answer point to the right document?
    4. Honest when missing: Does it avoid guessing when no passage answers the question?

    Passing answer

    “No. The original shipping fee is not refundable. [Returns policy]”

    Failing answer

    “No, but the store will give you a credit for the shipping fee.”

    The second answer adds a store-credit promise that does not appear in the evidence. It fails even though its first word is correct.

    Step 4: Use the failure to find the right fix

    What happenedLikely causeFirst thing to fix
    Right passage was not retrievedSearch or document problemIndexing, query, or ranking
    Right passage was retrieved, but answer was wrongAnswer-generation problemPrompt or model behavior
    Answer was right, but citation was wrongCitation mapping problemSource IDs or formatting
    No passage answered, but bot guessedMissing-answer behaviorAdd an explicit “do not guess” rule

    This small table is the main reason to separate retrieval from answering. One overall quality score cannot tell you which part broke.

    Step 5: Run one missing-evidence test

    Remove the returns-policy passage for the shipping-fee question and run the chatbot again.

    Pass: “I cannot find that answer in the available documents. I can connect you with support.”

    Fail: The chatbot repeats a policy from memory or invents one.

    This test proves the bot depends on your documents instead of merely producing a plausible answer.

    Step 6: Compare one change on the same cases

    Run the five cases with your current search settings. Then change only one thing, such as the number of passages returned, and run the same cases again.

    VersionRetrieval passedAnswers passedUnsupported claims
    Top 3 passages4/53/51
    Top 5 passages5/55/50

    More passages are not always better. Extra irrelevant text can confuse the model, so inspect the rows instead of assuming a larger number wins.

    Step 7: Add a multi-step agent only when needed

    If the chatbot can also look up orders or change account data, keep the RAG checks above and add the tool checks from the tool-calling guide.

    For a request such as “Find the return policy, check whether order 1042 qualifies, and start a return,” test each step:

    1. Did it retrieve the right policy?
    2. Did it look up the correct order?
    3. Did it apply the policy correctly?
    4. Did it ask before creating the return?
    5. Did the final reply match the actual result?

    That is an agent evaluation in plain language. You are checking the steps and the outcome, not assigning one mysterious score to the whole run.

    A simple release rule

    For this RAG chatbot, release only when:

    • every critical question retrieves the required passage
    • every policy claim is supported by retrieved text
    • every answer cites the right document
    • every missing-answer case admits that the evidence is missing
    • the new version does not break a case the current version passes

    Use the interactive example below to see retrieval and answer failures separately. When the test set outgrows a spreadsheet, the Phoenix, Langfuse, and LangSmith courses show how to run the same loop in a platform.

    Interactive test

    Try three RAG chatbot test cases

    Find the right Northstar Shoes document and answer only from that evidence.

    Input

    Is original shipping refunded?

    Expected behavior

    Find Returns policy; answer no; cite it

    Observed output

    No. Original shipping fees are not refundable. [Returns policy]

    right evidencePass
    supported answerPass
    right citationPass
    Case passed3/3 checks
    Direct answers

    Frequently asked questions

    What should I evaluate in a RAG chatbot?

    Check retrieval and answering separately: did search find the needed passage, and did the chatbot answer only from that evidence with the right citation?

    What should a RAG chatbot do when the answer is not in its documents?

    It should say the available documents do not answer the question and offer a useful next step instead of guessing.

    RAG EvalsAgent EvalsLLM Evaluation
    Share
    BP

    Bhaulik Patel

    Forward deployed AI engineer and creator of Deployed Engineer.