Grounded assistants that refuse well
Retrieval, guardrails and refusal design for an assistant that speaks to customers — including the failure modes we hit building the one on this site.
An assistant that talks to buyers is a liability the first time it invents a price, claims a certification the firm does not hold, or agrees to a contract term. It is also, done properly, the most patient salesperson a firm has. The difference between those outcomes is almost entirely in what happens when the assistant does not know something.
This article describes the architecture behind the assistant on this site, including two mistakes that made it worse before it got better.
Separate what the model decides from what it says
The most useful structural decision is to give the model exactly one job. In our case it writes answers from retrieved passages. It does not choose which qualifying question to ask, does not score a lead, does not decide where an enquiry is routed. Those are deterministic rules in a content file that a non-engineer can read and argue with.
- Two identical visitors produce identical briefs for the sales team — a model-scored lead cannot promise that.
- When the routing is wrong, you fix a rule rather than tune a prompt and hope.
- The parts that must be auditable are auditable, and the part that benefits from fluency gets fluency.
Lexical retrieval is often enough — and debuggable
A corpus of forty short passages does not need an embedding index. IDF-weighted lexical overlap with a tag boost retrieves correctly, costs nothing, adds no operational dependency, and — the property that matters most in production — can be explained. When it retrieves the wrong passage you can see which token did it. Reach for embeddings when the corpus is large enough that vocabulary mismatch genuinely hurts, not because it is the expected answer.
Refusal is a design problem, not a safety switch
Our first version refused anything the corpus did not cover word for word. It was safe and it was useless: a buyer asking "do you build mobile apps?" was told there was no published answer, which reads as a firm that does not recognise its own work. The fix is tiering, so that "I do not know" is reserved for the rare case that deserves it.
| Retrieval strength | Mode | What the assistant does |
|---|---|---|
| Strong | Grounded | Answers from the retrieved passages, cites the page |
| Weak, but in our domain | Capability | Answers from an approved capability brief, says specifics come from a person, offers the call |
| Nothing, out of domain | Redirect | Says what it does cover and offers a human — politely, once |
The threshold that separates the first two tiers should scale with question length. A fixed bar means a two-word question can never clear it, which is how "what does it cost" ends up in the generic tier while a rambling question about the same thing is answered properly.
Assert the constraints twice
A system prompt is an instruction, not a guarantee. Every hard constraint is stated in the prompt and then checked against the generated text before it is returned. If the check fires, the answer is discarded and a safe reply is sent instead. A guardrail that exists in one place is not a guardrail.
PRICE_RE = re.compile(r"(₹|\brs\.?\s*\d|\$\s*\d+|\blakhs?\b|\bcrores?\b)", re.I)
PROMISE_RE = re.compile(r"\b(guarantee[ds]?|we promise|100%\s+(uptime|secure))\b", re.I)
ISO_RE = re.compile(r"\biso\b|\b27001\b|\b9001\b", re.I)
ISO_OK_RE = re.compile(r"(do(es)? not hold|not certified|not yet|implementing)", re.I)
def violates_policy(text: str) -> str | None:
if PRICE_RE.search(text): return "price"
if PROMISE_RE.search(text): return "promise"
if ISO_RE.search(text) and not ISO_OK_RE.search(text):
return "certification"
return None- Log every block with the rule that fired. The blocked answers are the highest-value data you have about prompt weaknesses.
- Rate-limit per client. It caps both abuse and the bill, and it belongs in the application as well as at the gateway.
- Cap input length. A long input is either a paste of a whole document or an attempt at prompt injection; neither needs eight hundred tokens of context.
Degrade honestly, and never claim delivery
Two failure modes matter more than model quality, because they occur in front of a customer who is deciding whether to trust the firm.
Evaluating it like software
Every question a buyer asks becomes a test case with an expected tier and an expected refusal or grounding. The suite runs in seconds because retrieval is deterministic, and it catches the regressions that matter: a corpus edit that makes a common question fall out of the grounded tier, or a threshold change that lets an out-of-domain question through.
GROUND what services do you offer -> services-overview
GROUND who owns the code you write -> ownership
GROUND are you ISO 27001 certified -> iso-status
GROUND can you integrate razorpay -> integrations
CAPABLE can you build an inventory module -> capability brief
CAPABLE we need a school management system -> capability brief
REFUSE what is the capital of France -> redirect
REFUSE best biryani in hyderabad -> redirect