Skip to content

Reduce unsupported answers with a better review process.

Start with the evidence an assistant can retrieve, then test what it says when that evidence is incomplete. A confident tone is not a substitute for support.

First-party product guidance. How we maintain these guides · Report a correction

Check the source before changing the model.

Pick an answer that matters to the business and identify the passage that should support it. Confirm that the source is public, readable, current, and included in the site’s knowledge. Check whether other pages contradict it.

A missing exception or an outdated price can lead to a wrong answer even when the assistant accurately repeats retrieved text. Decide which page is authoritative and resolve inconsistencies there. Review confirmed facts and saved answers as well, since they can outlive an edit to the original page.

Understand what Sairo’s controls do.

Sairo checks saved answers and retrieves site facts and passages. Generated responses receive instructions to use that context, report missing information, and offer follow-up when needed. Instruction-like phrases in retrieved text are filtered before entering the prompt.

A response guard checks unsupported currency and percentage figures against the retrieved material. That is a targeted check, not a complete factual verifier. It does not prove that every policy, date, implication, or recommendation in an answer is correct.

Write questions with known expectations.

Build a small test set from actual visitor needs. Include direct questions, natural paraphrases, questions with an extra condition, and plausible questions whose answer is absent. Write down the expected behavior before running the test.

Test the boundary around a saved answer. A response approved for one service, location, or customer type should not be treated as a universal rule for a nearby but different question.

Example reliability test set for a service website
Question typeExampleExpected behavior
Known factWhat are your published opening hours?Use the current confirmed detail or supporting page.
ConditionAre those hours the same on holidays?Use explicit holiday guidance or acknowledge the gap.
Absent detailCan you guarantee a visit tomorrow?Avoid inventing availability and offer the real next step.
ParaphraseWhen can someone take my call?Use relevant information without confusing it with appointment availability.
Human requestI need to speak to a person.Enter the site’s live chat queue when enabled, or explain the configured follow-up route.

Record the failure precisely.

Save the question, the response, the source you expected, and the specific unsupported statement. Distinguish a retrieval failure from a content problem or an overbroad interpretation. Each points to a different intervention.

A response can be partly correct and still fail the test. Check conditions and scope, not just whether it includes the right phone number or a relevant citation. A linked page must actually support the claim being made.

Make one change and retest the boundary.

Improve or exclude the source, update a confirmed fact, or write an approved answer. Ask the original question again after the source refresh finishes, then test nearby questions that should produce a different result.

If you change the model, repeat the regression set rather than assuming a more expensive model is always a better fit. Keep the source, configuration, and review criteria consistent enough to understand what changed.

For a specific Sairo reply, use the inbox’s “Review with Ask Sairo” action. Supply the confirmed answer, inspect the proposed change and check the original and nearby questions. This gives the correction a reviewable record without replacing source maintenance.

A useful review record

Question → expected source → answer received → unsupported detail → change made → retest result. Add the date and the person responsible for maintaining the underlying business information.

Keep a person responsible for exceptions.

Some questions cannot be answered from a public website because they require an individual decision. Treat human handoff as a deliberate outcome: use the site’s live chat queue when enabled, or provide a usable follow-up route with a team that knows how to respond.

Continue sampling real conversations after launch. A test set provides a repeatable check for known situations; it cannot cover every future question, content change, or visitor expectation.

Common questions

Can a source-grounded chatbot still give a wrong answer?

Yes. A source may be outdated or incomplete, retrieval may find the wrong passage, or the model may interpret the material incorrectly. Grounding makes evidence available; review is still necessary.

Should every missing answer become a saved answer?

Only after the business has established the information. Some requests should remain a human decision rather than becoming a general rule.

See what your website can answer.

Try a sample, review the result, and decide what comes next.