All guides
RoutingAdvanced·Advanced7 min

Model Selection & Smart Routing

On the OpenRouter path you either name an exact model or let a routing strategy pick — price, throughput, latency, or exacto — and then bound the run with up to three fallback models, per-token and per-request price caps, and a provider allowlist or denylist. These four levers are OpenRouter-only; the Gemini and LM Studio drivers ignore them without error.

VantagePrompt routes through OpenRouter by default, which means you can name an exact model or hand the choice to a routing strategy — then bound the run with fallbacks, price caps, and provider lists so it stays fast, cheap, and on infrastructure you trust.

By Andrei Bădulescu, Founder at VantagePrompt·Updated

VantagePrompt runs your prompts through one of three backends: OpenRouter (the production default), Google Gemini, or a local LM Studio server. OpenRouter is a meta-provider — one connection, hundreds of models from OpenAI, Anthropic, Google, Meta, Mistral, and more. That breadth is the point: you pick the model, or you let routing pick for you, and you set guardrails so a run never costs more or lands on a provider you didn't want.

This guide covers the four levers that matter on the OpenRouter path: routing strategy, fallback models, max-price caps, and provider allow/deny lists. These live in the Advanced Controls and apply only when you route through OpenRouter. Gemini and LM Studio ignore them — the setters are no-ops on drivers that don't support the feature, so nothing breaks, it just has no effect.

Should I name a model or let routing pick one?

You can name a specific model and VantagePrompt sends the request straight to it. If you leave the model unset on the OpenRouter path, routing chooses one for you based on your use case and the live model catalog. That's the right default when you don't care which model runs — you care about price, speed, or reliability instead. Routing strategy is how you tell it which of those to optimize for.

What do the four routing strategies do?

The routing strategy picks an ordered preference list of models and serves the first one that's live. Four strategies are supported:

StrategyOrders providers byReach for it when
priceCheapest first.Bulk, low-stakes work where good-enough output beats a premium model.
throughputHighest tokens-per-second first.You want the full output back fast and do not mind who serves it.
latencyLowest time-to-first-token first.Interactive work, where the wait before streaming starts is what you feel.
exactoProviders that honour every requested parameter.The request leans on structured output, tools, or reasoning parameters.
The four routing strategies. Set one and it overrides the use-case preference list.

Strategy wins over the use-case default: set one, and it overrides the model list the use case would have chosen. Leave it unset and routing falls back to the use-case preference list.

How do fallback models work?

A fallback chain lets OpenRouter try your primary model first, then move down the list if the primary errors. You can add up to three fallback models. They're sent as an ordered list — primary first, fallbacks after — and tried in order on failure.

OpenRouter rejects a request with more than 3 entries in the model list. VantagePrompt enforces that cap for you, merging your chosen primary, any per-request fallbacks, and the configured default into a list of at most three — so you can't accidentally build an invalid chain.

Fallback cascade and circuit breakerPer request, OpenRouter tries the primary model and moves down an ordered list of at most two fallbacks on error. Separately, three consecutive transient failures trip a circuit breaker that routes to a fallback for a five-minute cooldown before retrying the primary.Per request — your explicit preferencePrimary modelerrorFallback 1errorFallback 23 entries max — OpenRouter rejects a longer model listUnderneath — the automatic safety net3 transient failures429 / 5xx / timeoutBreaker tripsFallback model5-minute cooldownthen retry
Your fallback list is a per-request preference; the circuit breaker sits underneath it.

This is separate from the automatic circuit breaker. After three consecutive transient failures (rate limits, 5xx, timeouts), VantagePrompt trips a breaker and routes to a fallback for a five-minute cooldown, then retries the primary. Fallback models are your explicit per-request preference; the breaker is the automatic safety net underneath it.

How do max-price caps stop a run from overspending?

Max price filters out any provider whose rate exceeds your cap before the request is sent. You set three optional ceilings:

CapUnitFilters out
promptUSD per 1M input tokensProviders charging more to read your prompt.
completionUSD per 1M output tokensProviders charging more to generate the answer.
requestFlat USD for one requestProviders whose total for this call exceeds the ceiling.
Three independent ceilings — set any subset; unset caps do not filter.
{
  "prompt": 5.0,
  "completion": 15.0,
  "request": 0.005
}
Cap prompt and completion price per million tokens, plus a flat per-request ceiling

OpenRouter drops providers above any cap you set. Pair this with the price routing strategy when cost is the hard constraint: the strategy prefers cheap models, and the cap guarantees nothing expensive slips through even if a cheap provider is briefly unavailable.

How do I restrict which providers serve my request?

Many models are served by more than one upstream provider. These lists control which providers OpenRouter is allowed to use:

  • Allowlist — restrict routing to only the providers you name (sent as an ordered preference).
  • Denylist — route to anything except the providers you name.

Allowlist and denylist are mutually exclusive — set one or the other, not both. Invalid provider names are stripped against the list of known OpenRouter providers, so a typo silently drops rather than breaking the request. If your allowlist filters out every option, the request has nowhere to route.

What does a cost-controlled setup look like?

Say you're refining a marketing prompt and want it cheap, resilient, and never routed through a provider you've ruled out. A sensible Advanced Controls setup:

Routing strategy : price
Fallback models  : openai/gpt-4o-mini, mistralai/mistral-small
Max price        : { prompt: 2.0, completion: 6.0 }
Provider denylist : (any provider you want to avoid)
Cost-controlled, resilient OpenRouter setup

Routing prefers the cheapest live model; the price caps guarantee nothing over your ceiling runs; the fallback chain keeps the run alive if the first pick errors; the denylist keeps it off providers you don't trust. Every lever is optional — set only the ones that match your constraint and leave the rest on their defaults.

Running the same controls across many prompts? Set them once and refine a whole set together with Batch Optimization. Want a provider to never log or train on your request? That's a separate toggle — see Privacy & Zero Data Retention.

Frequently asked questions

What is the difference between fallback models and the circuit breaker?
Fallback models are your explicit per-request preference: OpenRouter tries the primary first and moves down the list on error. The circuit breaker is the automatic safety net underneath — after three consecutive transient failures (rate limits, 5xx, timeouts) it trips and routes to a fallback for a five-minute cooldown, then retries the primary.
Why am I limited to three fallback models?
OpenRouter rejects a request with more than three entries in the model list. VantagePrompt enforces that cap by merging your chosen primary, any per-request fallbacks, and the configured default into a list of at most three, so you cannot build an invalid chain.
Can I use a provider allowlist and a denylist together?
No — they are mutually exclusive; set one or the other. Invalid provider names are stripped against the list of known OpenRouter providers, so a typo drops silently rather than breaking the request. Be careful: an allowlist that filters out every option leaves the request nowhere to route.
Do these controls work with Gemini or a local LM Studio server?
No. Routing strategy, fallbacks, max price, and provider lists apply only on the OpenRouter path. On the other drivers the setters are no-ops — nothing breaks, the settings simply have no effect.
Should I use price routing or a max-price cap?
Both, when cost is the hard constraint. The strategy prefers cheap models; the cap guarantees nothing above your ceiling runs even if a cheap provider is briefly unavailable.

Sources

Put it into practice.

Run this technique in the optimizer.

Open the optimizer

Keep reading