Model Selection & Smart Routing
On the OpenRouter path you either name an exact model or let a routing strategy pick — price, throughput, latency, or exacto — and then bound the run with up to three fallback models, per-token and per-request price caps, and a provider allowlist or denylist. These four levers are OpenRouter-only; the Gemini and LM Studio drivers ignore them without error.
VantagePrompt routes through OpenRouter by default, which means you can name an exact model or hand the choice to a routing strategy — then bound the run with fallbacks, price caps, and provider lists so it stays fast, cheap, and on infrastructure you trust.
By Andrei Bădulescu, Founder at VantagePrompt·Updated
VantagePrompt runs your prompts through one of three backends: OpenRouter (the production default), Google Gemini, or a local LM Studio server. OpenRouter is a meta-provider — one connection, hundreds of models from OpenAI, Anthropic, Google, Meta, Mistral, and more. That breadth is the point: you pick the model, or you let routing pick for you, and you set guardrails so a run never costs more or lands on a provider you didn't want.
This guide covers the four levers that matter on the OpenRouter path: routing strategy, fallback models, max-price caps, and provider allow/deny lists. These live in the Advanced Controls and apply only when you route through OpenRouter. Gemini and LM Studio ignore them — the setters are no-ops on drivers that don't support the feature, so nothing breaks, it just has no effect.
Should I name a model or let routing pick one?
You can name a specific model and VantagePrompt sends the request straight to it. If you leave the model unset on the OpenRouter path, routing chooses one for you based on your use case and the live model catalog. That's the right default when you don't care which model runs — you care about price, speed, or reliability instead. Routing strategy is how you tell it which of those to optimize for.
What do the four routing strategies do?
The routing strategy picks an ordered preference list of models and serves the first one that's live. Four strategies are supported:
| Strategy | Orders providers by | Reach for it when |
|---|---|---|
| price | Cheapest first. | Bulk, low-stakes work where good-enough output beats a premium model. |
| throughput | Highest tokens-per-second first. | You want the full output back fast and do not mind who serves it. |
| latency | Lowest time-to-first-token first. | Interactive work, where the wait before streaming starts is what you feel. |
| exacto | Providers that honour every requested parameter. | The request leans on structured output, tools, or reasoning parameters. |
Strategy wins over the use-case default: set one, and it overrides the model list the use case would have chosen. Leave it unset and routing falls back to the use-case preference list.
How do fallback models work?
A fallback chain lets OpenRouter try your primary model first, then move down the list if the primary errors. You can add up to three fallback models. They're sent as an ordered list — primary first, fallbacks after — and tried in order on failure.
OpenRouter rejects a request with more than 3 entries in the model list. VantagePrompt enforces that cap for you, merging your chosen primary, any per-request fallbacks, and the configured default into a list of at most three — so you can't accidentally build an invalid chain.
This is separate from the automatic circuit breaker. After three consecutive transient failures (rate limits, 5xx, timeouts), VantagePrompt trips a breaker and routes to a fallback for a five-minute cooldown, then retries the primary. Fallback models are your explicit per-request preference; the breaker is the automatic safety net underneath it.
How do max-price caps stop a run from overspending?
Max price filters out any provider whose rate exceeds your cap before the request is sent. You set three optional ceilings:
| Cap | Unit | Filters out |
|---|---|---|
| prompt | USD per 1M input tokens | Providers charging more to read your prompt. |
| completion | USD per 1M output tokens | Providers charging more to generate the answer. |
| request | Flat USD for one request | Providers whose total for this call exceeds the ceiling. |
{
"prompt": 5.0,
"completion": 15.0,
"request": 0.005
}OpenRouter drops providers above any cap you set. Pair this with the price routing strategy when cost is the hard constraint: the strategy prefers cheap models, and the cap guarantees nothing expensive slips through even if a cheap provider is briefly unavailable.
How do I restrict which providers serve my request?
Many models are served by more than one upstream provider. These lists control which providers OpenRouter is allowed to use:
- Allowlist — restrict routing to only the providers you name (sent as an ordered preference).
- Denylist — route to anything except the providers you name.
Allowlist and denylist are mutually exclusive — set one or the other, not both. Invalid provider names are stripped against the list of known OpenRouter providers, so a typo silently drops rather than breaking the request. If your allowlist filters out every option, the request has nowhere to route.
What does a cost-controlled setup look like?
Say you're refining a marketing prompt and want it cheap, resilient, and never routed through a provider you've ruled out. A sensible Advanced Controls setup:
Routing strategy : price
Fallback models : openai/gpt-4o-mini, mistralai/mistral-small
Max price : { prompt: 2.0, completion: 6.0 }
Provider denylist : (any provider you want to avoid)Routing prefers the cheapest live model; the price caps guarantee nothing over your ceiling runs; the fallback chain keeps the run alive if the first pick errors; the denylist keeps it off providers you don't trust. Every lever is optional — set only the ones that match your constraint and leave the rest on their defaults.
Running the same controls across many prompts? Set them once and refine a whole set together with Batch Optimization. Want a provider to never log or train on your request? That's a separate toggle — see Privacy & Zero Data Retention.
Frequently asked questions
- What is the difference between fallback models and the circuit breaker?
- Fallback models are your explicit per-request preference: OpenRouter tries the primary first and moves down the list on error. The circuit breaker is the automatic safety net underneath — after three consecutive transient failures (rate limits, 5xx, timeouts) it trips and routes to a fallback for a five-minute cooldown, then retries the primary.
- Why am I limited to three fallback models?
- OpenRouter rejects a request with more than three entries in the model list. VantagePrompt enforces that cap by merging your chosen primary, any per-request fallbacks, and the configured default into a list of at most three, so you cannot build an invalid chain.
- Can I use a provider allowlist and a denylist together?
- No — they are mutually exclusive; set one or the other. Invalid provider names are stripped against the list of known OpenRouter providers, so a typo drops silently rather than breaking the request. Be careful: an allowlist that filters out every option leaves the request nowhere to route.
- Do these controls work with Gemini or a local LM Studio server?
- No. Routing strategy, fallbacks, max price, and provider lists apply only on the OpenRouter path. On the other drivers the setters are no-ops — nothing breaks, the settings simply have no effect.
- Should I use price routing or a max-price cap?
- Both, when cost is the hard constraint. The strategy prefers cheap models; the cap guarantees nothing above your ceiling runs even if a cheap provider is briefly unavailable.
Sources
Put it into practice.
Run this technique in the optimizer.