◀ Knowledge hub

03, AI & intelligence

Provider fallback chains and a cost ceiling

codeAmani Labs Engineering
Cinematic still for Provider fallback chains and a cost ceiling

A single model provider is a single point of failure

If every AI feature in your product calls one vendor, then that vendor's bad afternoon is your bad afternoon. Fallback chains turn a hard dependency into a soft one: a primary provider for normal operation, one or more backups for when the primary is down, slow, or rate limiting you.

Order the chain by what each tier is for

A good chain is not three copies of the same model. It is a deliberate ordering by capability and cost. Put the model you actually want first, a comparable alternative second for resilience, and a cheaper tier last so that "degraded" still means "answered".

const CHAIN = [
  { model: "openai/gpt", timeoutMs: 8000 },
  { model: "anthropic/claude", timeoutMs: 8000 },
  { model: "deepseek/chat", timeoutMs: 12000 },
];

async function withFallback(prompt: string) {
  let lastError: unknown;
  for (const step of CHAIN) {
    try {
      return await callWithTimeout(step.model, prompt, step.timeoutMs);
    } catch (err) {
      lastError = err;
    }
  }
  throw lastError;
}

Time out aggressively

A provider that is slow is, from the user's point of view, a provider that is down. A timeout on each step is what converts "hanging for thirty seconds" into "fell through to the backup in eight". Without it, the fallback never triggers because the request never fails, it just waits.

Put a ceiling on cost

Resilience and budget pull in the same direction here. Measuring every call through one gateway lets you set a spend ceiling and route accordingly: send the bulk, low stakes work to the cheap tier by default and reserve the premium model for the requests that justify it. The chain handles failure; the routing handles cost; both rely on the same uniform interface.

What to log

Record which step answered. If the primary is quietly failing half the time and the backup is silently carrying the load, you want to know before the backup also fails. The chain hides failure from the user on purpose, so the logs are the only place the failure is visible. Watch them.

Qualified conversation

Have a build to de-risk? Let's talk.

Tell us what you are building. We respond within two business days.