Grounding: Cut Hallucination with Web Search
Grounding anchors a model to evidence retrieved from the web instead of its frozen training data, which is how you stop it filling gaps with plausible-but-wrong detail. In VantagePrompt it is the web search toggle in the OpenRouter Advanced Controls, capped by a max-results value of 1 to 10, and it only has effect on search-enabled models.
A model only knows what it was trained on. Ask about a release that shipped last week, a price that changed yesterday, or a person who just changed jobs, and it will confidently fill the gap with something plausible and wrong. That gap-filling is hallucination. Grounding closes the gap by feeding the model live facts pulled from the web before it answers. In VantagePrompt, that switch is the web search control.
By Andrei Bădulescu, Founder at VantagePrompt·Updated
What does grounding actually do?
Grounding means anchoring a model's output to retrieved evidence instead of its frozen training data. When web search is enabled, the model can pull current results from the web and use them while it works. The output is built on facts that exist outside the model's memory, so it is far less likely to invent details.
This is not magic. Grounding does not make a model smarter — it gives it fresher inputs. The model still reasons over what it retrieves. Garbage in, garbage out still applies. But for anything time-sensitive or fact-heavy, fresh inputs beat a confident guess every time.
Where is the web search toggle, and what does max results do?
We explore a general-purpose fine-tuning recipe for retrieval-augmented generation (RAG) — models which combine pre-trained parametric and non-parametric memory for language generation.
Web search lives in the OpenRouter Advanced Controls panel on the optimizer, alongside routing strategy and ZDR. Flip it on and the request carries a web plugin to the provider. Next to the toggle is a max results control that caps how many search results feed the model — a value from 1 to 10.
More results means broader coverage but more noise and more tokens to process. Fewer results keeps it tight and cheap. A practical default is small: 3 to 5 results cover most factual lookups. Push toward 10 only when a topic is genuinely scattered across many sources and you need breadth.
Web search is a paid-tier OpenRouter control. It surfaces in the Advanced Controls panel, which is hidden on the free tier and only appears when OpenRouter is the active provider.
Which models support web search?
The toggle is only meaningful on models that support web search — specifically Anthropic and OpenAI search-enabled models routed through OpenRouter. Turn it on for a model that does not support search and it is effectively a no-op: the model answers from training data as if the switch were off. If grounding matters for your task, pick a model that supports it. See the model selection guide for how routing and model choice interact.
When does grounding help, and when does it not?
Grounding earns its latency and tokens when the answer depends on a fact that can change; it earns nothing when the answer is already in your input or in the model’s stable knowledge. The split is about whether there is anything to look up.
| Grounding helps | Grounding does not help |
|---|---|
| Current events, releases, or version numbers that postdate the model's training cutoff. | Pure reasoning, math, or logic — there is nothing to look up. |
| Prices, specs, availability, or anything that changes frequently. | Creative writing, brainstorming, or tone work — facts are not the constraint. |
| Named entities — people, companies, products — whose status may have shifted. | Code structure and refactoring, where the answer comes from your input. |
| Niche or fast-moving topics where the model's prior is thin or outdated. | Stable, well-known facts the model already holds reliably. |
| Anything where a wrong-but-confident answer is worse than a slower, sourced one. | Cases where extra results add noise and cost without changing the answer. |
Grounding adds latency and tokens. Every search result the model reads is more input to process. Leave it off by default and switch it on deliberately for fact-dependent prompts — do not treat it as a free accuracy boost.
What does grounding change on a real prompt?
Compare the same request with grounding off versus on. Off, the model leans on training data and may quote a stale figure. On, it retrieves current results and writes from them.
Write a short briefing on the current pricing tiers and context window
of the latest flagship models from the major LLM providers.With web search off, the optimized prompt is sound, but the eventual answer is only as fresh as the model's cutoff — pricing and limits drift fast, so it risks confidently stating last year's numbers. With web search on (max results 5) and a search-enabled model, the model pulls live data and the briefing reflects what is actually published now.
How do I write a prompt grounding can use?
Grounding works best when the prompt tells the model what to verify. Vague asks retrieve vague results. Be explicit about the facts that must be current and, where it helps, ask the model to lean on retrieved sources over assumptions.
Summarize the latest stable release of <library>, including its
version number and release date. Prioritize current, verifiable
information over prior assumptions; if a fact is uncertain, say so.That last clause matters. Telling the model to flag uncertainty instead of papering over it is itself an anti-hallucination move — it turns a silent guess into a visible gap you can chase down.
Does grounding eliminate hallucination?
Grounding lowers hallucination risk; it does not eliminate it. After a grounded run, sanity-check the facts that matter and read the quality score to gauge how tight the optimized prompt is. If results still feel thin, nudge max results up, tighten the phrasing, or confirm your model actually supports search — then run it again.
Frequently asked questions
- What does grounding actually do to a prompt?
- It anchors the output to retrieved evidence rather than the model's frozen training data. It does not make the model smarter — it gives it fresher inputs, and the model still reasons over whatever it retrieves. Garbage in, garbage out still applies.
- How many search results should I allow?
- Three to five covers most factual lookups. More results means broader coverage but more noise and more tokens to process; fewer keeps it tight and cheap. Push toward the maximum of ten only when a topic is genuinely scattered across many sources.
- Why does the web search toggle seem to do nothing?
- It is only meaningful on models that support web search — the Anthropic and OpenAI search-enabled models routed through OpenRouter. On a model without support it is a no-op: the answer comes from training data as if the switch were off. It is also a paid-tier control and lives in the Advanced Controls panel, which is hidden on the free tier and only appears when OpenRouter is the active provider.
- Does grounding eliminate hallucination?
- No, it lowers the risk. After a grounded run, sanity-check the facts that matter. If results still feel thin, raise max results, tighten the phrasing, or confirm the model actually supports search.
- How do I write a prompt that grounding can use well?
- Tell the model what to verify. Vague asks retrieve vague results. Name the facts that must be current, and add a clause asking the model to prioritise verifiable information and to say so when a fact is uncertain — turning a silent guess into a visible gap.
Sources
reasoning prompts to try
Browse all reasoning promptsPublished by the community and free to copy — worked examples of what this guide describes.
Put it into practice.
Run this technique in the optimizer.