Gateway monitors
An AI gateway can be up while every request from your app fails. A gateway monitor lists the models, asks each one you pick a tiny question, and can send it with the same settings your app sends, so a model that refuses one of them fails the check the way it fails your app.
What each run does
- The models list (
/v1/models, or the Anthropic and Google equivalents). It proves the key and the routing work, and it is free. - The gateway’s own health path, when it has one (LiteLLM
/health/liveliness, Portkey/v1/health). - A tiny streamed reply from each model you pick: “Reply with OK”, capped at 5 tokens. We time the first token and the tokens a second, and record which model really answered.
It works with any OpenAI-compatible base URL (OpenRouter, LiteLLM, Portkey, Vercel and Cloudflare gateways, your own), the Anthropic and Google APIs, and model servers you run yourself, such as vLLM and Ollama.
Send what your app sends
A model upgrade can reject a parameter your app sends, and a plain ping won’t see that. It happened to us: our gateway moved our assistant to a newer Claude model, which answers 400 “`temperature` is deprecated for this model” to any request that sets temperature. The gateway was up, the models list was fine and a bare prompt worked. Every real request failed.
So a gateway monitor can carry your app’s real request. params are merged into each model’s tiny request, in the provider’s own spelling: temperature, top_p, max_tokens, stop, system, tools, tool_choice, thinking, response_format. headers go on every request: a gateway’s own sign-in or routing header, anthropic-beta, a newer anthropic-version. path sets the chat path when your gateway uses its own route.
monitors:
- name: Assistant gateway
type: gateway
url: https://gateway.yourapp.com
api_key: secret.GATEWAY_KEY
models: [claude-sonnet-5-5]
every: 15m
config:
api: anthropic
request:
params:
temperature: 0
max_tokens: 16
tool_choice: { type: tool, name: verdict }
tools:
- name: verdict
description: Pass or fail
input_schema: { type: object, properties: { pass: { type: boolean } } }
headers:
cf-aig-authorization: secret.CF_GATEWAY_TOKEN- The check keeps the model, the question and streaming.
model,messagesandstreamare refused inparams. - The reply stays at 5 tokens unless you copy your app’s
max_tokens(at most 4,096). Withthinkingon,max_tokensmust be above its budget, as the API itself requires. - A header that carries a credential (
authorization, an API key, a token) must be asecret.NAME. Its value is never shown in a result. - A reply that starts with a tool call or a thought counts as an answer.
Why it failed, in plain words
Each model that fails gets a cause, and the alert leads with it:
- Setting rejected: “The model refuses the setting
temperature. Your app sends it, so your app’s own requests fail the same way.” - Unknown model and model retired.
- Key refused, out of credit (even when the provider sends it as a 429), overloaded, rate limited (with the wait it asked for).
- Timed out, stream broke (it started answering and stopped) and empty answer.
When a provider’s replies announce that a model is being retired (a Sunset, Deprecation or warning header), you get one heads-up per model and date, well before the day: “claude-sonnet-4-5 is being retired on 30 November 2026.” It never changes the monitor’s state.
Down and degraded
- Down: the key is refused, the models list fails, or no model answers, confirmed from another region.
- Degraded: some models fail while others answer (“2 of 3 models failing”, each with its cause), the first token is slower than usual, or the gateway answered from a fallback model: “You asked for claude-sonnet-5-5 and claude-haiku-4-5 answered.” A dated name for the same model (
gpt-4o-mini-2024-07-18forgpt-4o-mini) is not a fallback. - A rate limit on its own never pages anyone.
An agent behind the gateway
An agent test can reach an agent that sits behind your gateway: give it the gateway’s headers, and for an OpenAI-compatible agent a body with your app’s own settings and {{ask}} where the question goes. A refused setting gets the same cause as here.
- name: Support bot answers
type: agent
endpoint: https://gateway.yourapp.com/v1
protocol: openai
headers: { x-portkey-config: pc-support }
body: '{"model":"support-bot","temperature":0,"messages":[{"role":"user","content":"{{ask}}"}]}'
cases:
- ask: "What is the refund window?"
expect:
- answered: { within: 20s }
- contains: "30 days"What it costs
Every monitor has a monthly token budget (200,000 by default). When it runs out the monitor pauses itself and tells you. Replayed settings such as a long system prompt or many tools add input tokens to every run, so keep them to what your app really sends and check every 5 to 15 minutes.