01 article

Cutting Production LLM Hallucination from 15% to 1.5%

One team cut their measured production LLM hallucination rate from 15% to 1.5% without touching the model. The fix lived in the platform layer: a schema that doubles as a lie detector, a three-class retry loop with hard caps, and prompts you can roll back without a redeploy.

Cutting Production LLM Hallucination from 15% to 1.5%

Cutting Production LLM Hallucination from 15% to 1.5%

Most teams that ship an LLM into production learn the same lesson a few weeks too late. The demo was fine. The model is fine. What is broken is the layer around the model, and that layer is the part you actually control. One team building an internal platform for inventory audits, replenishment review, and discrepancy triage wrapped their model in a retry and validation loop and cut their measured hallucination rate from 15% to 1.5%, without touching the foundation model at all. The model did not get smarter. The plumbing got honest.

The setup is deliberately boring. A single gateway is the only ingress. It authenticates the caller, issues a role-carrying JWT, exposes the agent execution endpoint, and serves per-user and per-team metrics. Behind it sits a coordinator agent, built on Google's ADK, that classifies intent on every turn and hands the request to one or more specialist agents. Each specialist owns its own prompt, its own output schema, and its own MCP server, partitioned by team or by usage. Nothing is shared in a way that lets one team's request bleed into another's.

The schema is a hallucination detector

The move that does the most work here is making the output schema double as a lie detector. Every specialist returns a typed object, not free text. The docs specialist's object carries a sources field, one entry per docs lookup, plus two flags. grounded may only be true if every claim is backed by what the tools returned, and out_of_scope is true when the request was never a docs question to begin with. Those two booleans are free to generate. They cost nothing. But they force the model to commit, in a machine-readable form, to whether its answer is actually supported.

A response that fails validation, is the wrong shape, is missing a required field, or is unparseable is the highest-precision hallucination signal a platform can get. The model did not just give a bad answer. It failed to give a conforming one. That is something you can count, retry, and alert on.

Grounding here is not a prompt trick. It is a pipeline. The tools fetch the evidence, the prompt restricts the model to that evidence, the schema forces a declaration, and validation turns a broken declaration into a counted, retryable failure.

The docs agent follows one strict rule: fetch the official documentation before answering any question about an API, a parameter, or version behavior. Never from training data, which is exactly where stale details hide. If the docs do not cover the question, the agent says so instead of filling the gap with plausible-sounding API names. And because the execution agent's claim about a deployment comes from the MCP tool's response rather than the model's memory, a confident sentence and a verified one finally become different things a dashboard can tell apart.

Three failure classes, three different retries

The retry loop is where a lot of teams lose money, because they collapse every failure into one generic "try again." That is the mistake. There are three classes and each needs a different answer.

  • Schema violations get re-prompted with the validation error appended, so the model sees exactly what broke. The grounded field came back as the string "yes" instead of a boolean. The required sources array is missing. Show it the shape it got wrong.
  • Hallucination signals get re-prompted with stronger grounding context. The docs agent returned grounded: false, or it cited a source for a library it was never given. Push the fetched evidence back in and ask again.
  • Infrastructure errors back off exponentially. A 429 rate limit or a 503 from the provider is not fixed by a reworded prompt, so you wait and retry with backoff rather than spending more tokens.

Every request gets a hard cap. The team uses three tries, on the theory that a structured-output failure which has not fixed itself by the third attempt almost never fixes itself on the fourth. You can instead cap the total tokens a request may spend, which controls actual cost since one retry on a long input can cost more than three on a short one, or cap wall-clock time when a fast answer matters more than a complete one. Pick one. What matters is that the limit exists and is enforced, so a single bad request cannot spin forever and quietly run up a bill. Each failed attempt is logged, and the final failure throws a hard exception, so callers never receive a silently broken payload. Tool-call failures work differently: when an MCP tool errors, the raw error is fed straight back into the model's context so it can self-correct and reattempt the call.

Prompts are versioned artifacts, not strings in a file

A prompt is executable text. Change one sentence and you change runtime behavior in a way no diff makes obvious. Two prompts that differ by a single line can have meaningfully different hallucination rates, latency, and token cost, and there is no compiler to warn you the new one is worse.

The platform treats prompts like deployed code. Every prompt has an ID, a version, a timestamp, and tags. Applications reference prompts by ID, never by pasting a string literal. Registering a new version does not overwrite the old one. It adds to the history and moves the "live" pointer forward, so you can roll back to v1.0.0 without redeploying anything. Tweaking prompt instructions without an audit trail is the fastest way to silently break your AI's behavior, and a history-preserving registry is the cheapest fix for it.

Authorize tools at the server, not the gateway

The last decision is the one people get wrong most often. Do not enforce tool authorization only at the API gateway. If the orchestrator has a single bug, a gateway-only check can expose every connected tool. Instead, each specialist's MCP server verifies the caller's token and applies a strict default-deny policy itself before any tool runs. The authority lives where the resource is, so a bug in the coordinator cannot hand a low-privilege user a high-privilege tool.

When this is overkill

If you are running one model behind one script with no team to share it with, a lot of this is scaffolding you do not need. The gateway, the per-team MCP partitioning, and the prompt registry all pay off when dozens of teams call the same platform and you need per-team cost attribution and rollback. For a single internal tool, the parts worth keeping are the structured output with a grounded flag and the three-class retry with a hard cap. Those two do most of the work, and the rest can wait.

The takeaway is not that a better prompt fixes a bad production system. It is that the platform around the model is where hallucination becomes a number you can drive down, and most of the levers cost nothing but discipline.

Comments