• Compute
  • Customers
  • Pricing
Sign In
More Blog Posts
XDiscordLinkedInYouTube

Products

  • GPUs
  • Inference
  • Studio

Developers

  • Model library
  • Documentation
  • Glossary

Company

  • About Us
  • Blog
  • Events
  • Partnership
  • Scale
  • Career
  • Ambassador program
  • Mission & Vision

Popular models

    Stay in the loop

    By submitting, you acknowledge that we may collect and use the information you provide, which may include personal information.

    XDiscordLinkedInYouTube

    Copyright ©2026 All rights reserved.

    Privacy PolicyTerms of UseLegal Documentation
    More Blog Posts
    Research

    Reducing Hallucination in Production LLM Apps: Grounding, Citations, and Faithfulness Scoring

    A production guide to reducing LLM hallucinations through grounding, citations, claim-level evaluation, retrieval quality checks, and layered monitoring, with special attention to multi-step agent workflows.

    October 04, 2026

    A mean groundedness score of 0.92 can hide a 20 percent per-claim hallucination rate. That single observation explains most of why teams ship systems they believe are reliable and then discover they are not: the measurement was taken at the wrong granularity. An answer-level score averages across every sentence in a response, so an answer with nine supported claims and one fabricated one scores 0.9 and looks excellent. The user reads the fabricated one. Reducing hallucination in production is partly a matter of grounding and citation discipline, and substantially a matter of measuring at claim level rather than answer level, because the aggregate is what makes a failing system look acceptable.

    • Hallucination rates vary by task shape more than by model. Benchmark data for 2026 puts extractive question answering at 3 to 8 percent of responses, open-ended generation at 15 to 25 percent, and multi-step agent workflows at 20 to 40 percent of tool-call chains.

    • Layered guardrails work, and the effect size is large. A meta-analysis across 12 production deployments found that combining grounding-discipline system prompts, RAG, and real-time monitoring cut hallucination rates by 71 to 89 percent relative to unguarded deployments.

    • GMI Cloud serves the judge layer as a distinct inference workload, which matters because running an LLM judge on every response across several detectors is the cost problem that causes teams to abandon evaluation in production.

    • Faithfulness, groundedness, and factuality are three different metrics. Faithfulness scores the whole answer against the retrieved context. Groundedness scores each sentence. Factuality scores against ground truth rather than against context. Using one where another is required is the most common measurement error.

    • RAG reduces hallucination substantially and does not eliminate it. Models still misinterpret correct sources, merge conflicting evidence without attribution, and fail to acknowledge when the retrieved documents do not contain the answer.

    • Structural citation checks miss the class that matters most. A fabricated citation that is well-formed passes format validation, which is why citation coverage has to be paired with claim-level entailment rather than treated as sufficient on its own.

    Three Metrics That Are Routinely Confused

    Getting these distinct is the precondition for everything else, because each catches a different failure and they are not interchangeable.

    Metric

    Granularity

    Reference

    Catches

    Faithfulness

    Whole answer

    Retrieved context

    Fabrication relative to provided evidence

    Groundedness

    Per sentence or claim

    Retrieved context

    Individual unsupported statements

    Factuality

    Whole answer or claim

    External ground truth

    Statements that are simply wrong

    Hallucination

    Final answer

    Reference-free

    Claims no retrieved chunk supports

    Faithfulness asks whether the generated answer is supported by the retrieved context with no fabricated claims. It is a whole-answer judgement and it is the right metric for summarisation against a source document.

    Groundedness is the per-sentence version. It scores each sentence in the answer against the retrieved evidence, which is what surfaces the case where most of an answer is supported and one sentence is not.

    Factuality compares against ground truth rather than against the context. An answer can be perfectly faithful to a retrieved document that is itself wrong, which faithfulness will score highly and factuality will not.

    Hallucination scoring acts as a reference-free safety net on the final answer, flagging claims that no retrieved chunk supports. Faithfulness should catch most of what it catches; it exists as a second line.

    The selection rule. For RAG question answering, groundedness and faithfulness against the retrieved context. For summarisation, faithfulness against the source document. For multi-step agents, faithfulness to tool outputs plus factuality of the final answer against a gold set, with human review on a sample.

    The Averaging Trap

    This deserves its own treatment because it is the measurement error that most reliably produces false confidence.

    The mechanism. An answer-level groundedness score averages across the claims in a response. A ten-claim answer with one unsupported claim scores 0.9. A dashboard showing mean groundedness of 0.92 across production traffic looks healthy and is consistent with one in five individual claims being unsupported.

    Why this matters more than the average suggests. Users do not experience averages. A user reading an answer encounters each claim individually, and the fabricated one is as visible as the nine supported ones. In a domain where a wrong claim has consequences, a 20 percent per-claim rate is a liability regardless of what the mean says.

    The fix is decomposition. Break the response into atomic claims and score each one against the evidence separately. The resulting metric is the proportion of claims that are supported, which is the number that corresponds to what a user encounters.

    What decomposition also enables. Once claims are separated, verification can be routed by claim type rather than applied uniformly. Research on financial hallucination detection found that existing detectors treating all claims the same way miss 43 percent of computational errors, because an arithmetic claim requires re-computation rather than entailment checking against text. A six-type taxonomy covering numerical, temporal, entity-attribute, comparative, regulatory, and computational claims, with type-routed verification including arithmetic re-computation, produced a 68 percent hallucination reduction over the strongest baseline when retrieval was held constant across systems.

    The generalisable lesson is that entailment is the right check for a descriptive claim and the wrong check for a calculation. A number that does not follow from the source figures is a hallucination that reads as perfectly entailed text.

    What RAG Fixes and What It Does Not

    Retrieval is the single highest-impact intervention and it is routinely oversold.

    What it fixes. The failure mode where the model has no relevant information and generates something plausible. Supplying the relevant passage removes the need to invent, and this is why retrieval with a strict citation contract produces the largest single-step reduction in fabrication for most production use cases.

    What it does not fix. Three failure modes survive retrieval. The model misinterprets a correct source, producing a claim the document does not actually support. It merges conflicting evidence from multiple retrieved chunks without flagging the conflict or attributing either position. And it fails to acknowledge that the retrieved documents do not contain the answer, generating one anyway rather than declining.

    The third is the most consequential in practice, because the retrieval step appears to have worked. Documents were retrieved, the answer cites them, and the answer is not in them.

    The upstream check that gets skipped. RAG pipelines need detection at two stages, not one. Faithfulness scoring compares the output to the retrieved documents, which catches generation failures. Context relevance scoring compares the retrieved documents to the query, which catches retrieval failures before generation runs.

    Checking only the output means a retrieval miss presents as a generation problem. The model was given irrelevant chunks and did its best, and the groundedness score penalises the generation step for a failure that occurred earlier. The RAG triad of context relevance, groundedness, and answer relevance exists precisely because these are distinct failures requiring distinct fixes.

    Citation Contracts and Their Limits

    Requiring citations changes model behaviour measurably, and the check on those citations has to go beyond format.

    What a citation contract does. Instructing the model to attach a source reference to each factual claim, and enforcing that structure in the output schema, constrains generation toward claims it can attribute. The constraint itself reduces fabrication, independently of whether anyone verifies the citations afterwards.

    Enforcement at the schema level. Structured output with a required citation field per claim means a response without attribution is not a valid response. This is the same principle as constrained decoding applied to the attribution requirement: making the invalid shape unproducible rather than detecting it afterwards.

    The class that structural checking misses. A citation can be well-formed, point at a real document in the retrieved set, and still not support the claim attached to it. The format validation passes. More seriously, in domains with recognisable citation conventions, a model can fabricate a reference that parses correctly and refers to nothing, which is the failure that produced several widely reported legal filing incidents.

    What catches it. Per-claim entailment scoring against the cited chunk specifically, rather than against the retrieved set as a whole. The question is not whether the claim is supported somewhere in the context, it is whether the cited source supports it. These come apart more often than they should.

    The three-part citation check. Structural validity, meaning the citation is well-formed and resolves to a retrieved chunk. Coverage, meaning the proportion of factual claims carrying a citation. And entailment, meaning the cited chunk actually supports the claim. All three are necessary and the third is the one most often omitted.

    The Cascade: Making Evaluation Affordable

    Running an LLM judge on every response across several detectors gets expensive quickly, and cost is why production evaluation gets switched off.

    The pattern that works. A cascade with increasing cost at each stage. Cheap heuristics first, a classifier-backed check next, and an LLM judge only for borderline cases.

    Stage one: deterministic checks. Schema validation, citation structure, known hallucination patterns, and format compliance. These cost effectively nothing, run in microseconds, and catch a meaningful share of failures outright.

    Stage two: classifier-backed scoring. A natural language inference classifier scoring claim-against-evidence entailment is far cheaper than a generative judge and handles clear cases well. Implementations that run an NLI classifier first and fall back to an LLM judge only on borderline scores capture most of the accuracy at a fraction of the cost.

    Stage three: LLM judge. Reserved for cases the classifier scores near the decision boundary. This is where the expense lives and where it is justified, because borderline cases are where a cheap check is unreliable.

    Sampling by stakes. Not every detector needs to run on every response. A documented production split runs grounding and factual checks at 100 percent while sampling reasoning and citation checks at 5 to 20 percent. The distinction is which failures are expensive: an ungrounded claim reaching a user is worse than an unmeasured reasoning step.

    Judge reliability. LLM-as-judge groundedness scoring reaches roughly 80 percent agreement with human evaluators, with one production implementation reporting 81.3 percent human correlation. That is good enough to drive thresholds and alerts and not good enough to be the only check in a high-stakes domain, which is why human review on a sample remains part of the design rather than a phase that ends.

    Thresholds and Runtime Blocking

    Scoring produces a number. Deciding what to do with it is a separate design question.

    Threshold guidance. Production groundedness thresholds above 0.85 are the published recommendation for critical domains, with roughly 85 percent appearing consistently as the blocking floor for high-stakes applications.

    Composite scoring. Rather than blocking on a single metric, combine groundedness, citation coverage, and semantic coherence into a weighted composite and block below a composite floor. This reduces the false positive rate relative to blocking on any single signal, because a response can score marginally on one dimension while being clearly acceptable overall.

    What blocking actually does. A response failing the threshold should not simply be suppressed. Three better options: regenerate with the failing claims flagged, return the response with the unsupported claims marked as unverified, or decline with an explanation that the available sources do not support a confident answer. Silent suppression produces a user staring at nothing, which is a worse experience than an honest declination.

    The regeneration option is underused. Hallucination-triggered regeneration, where a failing response is regenerated with the specific unsupported claims identified, converts a blocked response into a second attempt that is informed about what went wrong. It costs an additional generation and recovers responses that would otherwise be discarded.

    Where blocking is inappropriate. For latency-sensitive applications, a synchronous judge call before the response reaches the user adds its full latency to every request. The alternative is scoring asynchronously, surfacing results in monitoring rather than in the request path, and reserving synchronous blocking for the subset of requests where the stakes justify the latency.

    Agent Workflows: The Hardest Case

    The 20 to 40 percent hallucination rate on multi-step agent tool-call chains is the number that should reshape how agent evaluation is approached.

    Why the rate is so much higher. Errors compound across steps. A tool call constructed from a hallucinated parameter returns a plausible result, which the agent then reasons over as though it were correct. By the final answer, the fabrication is several steps upstream and the output reads as well-grounded in the tool results it was given.

    What to score. Faithfulness to tool outputs at each step, rather than only factuality of the final answer. An agent whose final answer is faithful to the tool results it received can still be wrong because the tool call itself was constructed incorrectly.

    The specific failure to look for. Silent continuation after a failed or partial tool result. An agent that receives an error, reasons around the missing information, and produces a confident complete-looking answer has hallucinated in a way that output-only scoring will not catch, because the output is internally coherent.

    Trace-level scoring rather than response-level. The unit of evaluation for an agent is the execution trace, not the final message. Step-by-step trace scoring, which checks each reasoning step and each tool interaction, is what locates where the chain broke.

    Production Monitoring

    Pre-deployment evaluation catches what the test set covers. Production monitoring catches the rest, and the dimensions it is tracked along determine whether it is actionable.

    Segment by feature, persona, and prompt version. An aggregate hallucination rate across all traffic hides the fact that one feature is responsible for most of it. Prompt version is the dimension most often missing, and without it a regression following a prompt change is indistinguishable from drift.

    Cluster failures by mode, not individually. Grouping bad traces by failure mode, such as citation invention, retrieval miss, or tool argument fabrication, converts a backlog of individual bad answers into a small set of patterns. Fixing a pattern addresses every instance; fixing an instance addresses one.

    Compare production scores against pre-deployment baselines. Scoring production responses against the same groundedness and faithfulness metrics used during development makes the two directly comparable, which is what turns a production score into a signal rather than a number without context.

    Track retrieval quality separately. Context relevance scoring on production traffic catches the retriever degrading as the corpus grows or as query patterns shift, which manifests downstream as rising hallucination that looks like a model problem.

    Infrastructure for the Evaluation Layer

    Hallucination detection is an inference workload, and it has a different profile from the application it monitors.

    The judge is a separate model serving a separate purpose. A small classifier handling the NLI stage of the cascade and a larger model handling the borderline cases are both distinct from the generation model. Running them on the same endpoint means evaluation competes with user-facing generation for the same capacity, which at load means either evaluation is throttled or latency rises.

    Asynchronous evaluation is batch inference. Scoring that happens after the response is delivered has no latency requirement, which makes it schedulable and batchable. This is the cheapest form of inference to serve and the most commonly run on the most expensive capacity.

    Synchronous blocking needs the latency profile of the hot path. Where a judge call gates the response, its time to first token adds directly to the user’s wait. GMI Prime Inference provides reserved capacity with warm weights for exactly this case, where a cold start on the judge is a cold start on the user’s request.

    Model selection for the judge is its own decision. A judge needs to be at least as capable as the model it evaluates on the dimension being judged, which is not the same as being the same model. The routing considerations that apply to selecting a model per task type, covered in Model Routing for AI Applications, apply to judge selection as directly as to generation.

    For agent deployments where trace-level scoring is the unit of evaluation, GMI Agentbox provides the per-session logging that step-by-step scoring attaches to.

    Conclusion

    The layered approach works and the effect is measurable: grounding-discipline prompts, retrieval, and runtime monitoring together cut hallucination rates by 71 to 89 percent against unguarded deployments across a dozen production systems. The components are not interchangeable and each addresses a different failure.

    The measurement decision matters more than any single technique. An answer-level groundedness mean of 0.92 is consistent with one claim in five being unsupported, which means the metric that looks healthy is the one that conceals the problem. Decomposing responses into atomic claims and scoring each against the evidence produces the number that corresponds to what a user actually encounters, and it enables routing verification by claim type, which matters because entailment checking is the wrong test for a calculation.

    Two practical constraints shape what is deployable. Evaluation cost is why production scoring gets switched off, so the cascade of deterministic check, then classifier, then LLM judge on borderline cases only, is what keeps it running. And in agent workflows, where tool-call chains hallucinate at 20 to 40 percent, the unit of evaluation is the execution trace rather than the final answer, because an output faithful to the tool results it received can still be wrong several steps upstream.

    Build AI Without Limits

    GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies

    FAQ

    They differ in granularity and in what they compare against. Faithfulness is a whole-answer score asking whether the generated answer is supported by the retrieved context with no fabricated claims, which makes it the right metric for summarisation against a source document. Groundedness is the per-sentence version, scoring each claim against the retrieved evidence, which is what surfaces an answer that is mostly supported with one unsupported sentence. Factuality compares against external ground truth rather than against the context, which catches the case where an answer is perfectly faithful to a retrieved document that is itself wrong. For RAG question answering the relevant pair is groundedness and faithfulness; for multi-step agents it is faithfulness to tool outputs plus factuality of the final answer against a gold set.

    Ready to build?

    Explore powerful AI models and launch your project in just a few clicks.

    Get Started