Middleware used to mean an Express function that checked a JWT and logged the request. A request that ends in an LLM call needs more than that. Something has to pick the model, check the tenant’s token budget, redact the email address in the response and decide what to do after the third 429 in a minute. If a dedicated layer doesn’t do it, forty route handlers do, each slightly differently.

ai aware backend middleware

Quick answer

An AI gateway (also called an LLM gateway or AI middleware) sits between your app and model providers. It owns five jobs: model routing, caching, token-aware rate limiting, observability and failover. You need one once you call more than one model, serve more than one tenant, or get a bill you can’t explain.

AI gateway vs API gateway

API gateway AI gateway
Unit of traffic Requests Tokens and requests
Reads the payload No Yes: prompts, completions, tool calls
Rate limiting Requests per second Tokens per minute + requests per minute
Caching Exact key match Semantic similarity
Routing Path, header Task difficulty, cost, provider health

Keep the API gateway at the edge for auth and validation (the usual Express middleware patterns). The AI layer sits behind it, with a thin provider adapter underneath:

API gateway (auth, validation)  →  AI middleware (decisions)  →  provider adapter (SDK calls)

Route handlers should never call a model ID directly. Once they do, you can’t change the model through config, attribute its cost or cache its responses.

1. Model routing

Most production traffic doesn’t need a frontier model. LMSYS’s RouteLLM cut costs by over 85% on MT Bench compared with sending everything to GPT-4, while keeping 95% of its quality score. Rules are enough for a first version:

typescript
type Tier = 'small' | 'frontier';

interface AIRequest {
  tenant: string;
  prompt: string;
  maxTokens?: number;
  requireCapability?: 'vision' | 'code' | 'reasoning';
}

function pickTier(req: AIRequest, inputTokens: number): Tier {
  if (req.requireCapability === 'reasoning') return 'frontier';
  if (inputTokens > 8_000) return 'frontier';
  return 'small';
}

// Ordered lists: primary model first, then fallbacks on other providers.
const MODELS: Record<Tier, string[]> = {
  small: process.env.MODELS_SMALL!.split(','),
  frontier: process.env.MODELS_FRONTIER!.split(','),
};

Log the tier on every call. A week of that data tells you whether a trained router would pay for itself.

2. Caching

Prompt caching is the cheap win. Anthropic charges up to 90% less for cached input reads, and OpenAI caches prompts over 1,024 tokens automatically. Keep the prefix stable: system prompt and tool definitions first, user content last, and no timestamps near the top.

Semantic caching skips the provider call when a new prompt means the same as one already answered. The similarity threshold decides everything. Prem AI’s measurements show 0.85 giving 45-70% hits but 15-30% wrong answers, while 0.93-0.95 gives around 30% hits with 3-7% wrong. Tune the threshold on your own labelled query pairs.

typescript
interface SemanticCache {
  // Keyed by tenant: a shared index will eventually serve one customer's answer to another.
  lookup(tenant: string, prompt: string): Promise<AIResponse | null>;
  store(tenant: string, prompt: string, res: AIResponse, ttlSec: number): Promise<void>;
}

function isCacheable(ctx: { usesTools: boolean; personalised: boolean }) {
  return !ctx.usesTools && !ctx.personalised; // "what's my balance?" must never hit the cache
}

A 20% hit rate on a $5,000 monthly bill saves $1,000 a month.

3. Token-aware rate limiting

Providers enforce tokens per minute (TPM) and requests per minute (RPM) at the same time. One 80,000-token job can be under the RPM limit and still block dozens of short chats. Use two buckets: reserve the worst case before the call, then correct it with real usage afterwards.

typescript
class TokenBucket {
  private tokens: number;
  private last = Date.now();
  constructor(private ratePerSec: number, private capacity: number) { this.tokens = capacity; }

  private refill() {
    const now = Date.now();
    this.tokens = Math.min(this.capacity, this.tokens + ((now - this.last) / 1000) * this.ratePerSec);
    this.last = now;
  }
  tryConsume(n: number) { this.refill(); if (this.tokens < n) return false; this.tokens -= n; return true; }
  adjust(n: number) { this.refill(); this.tokens = Math.min(this.capacity, this.tokens - n); }
}

export class DualBucketLimiter {
  private tpm: TokenBucket;
  private rpm: TokenBucket;

  constructor(tpmLimit: number, rpmLimit: number) {
    this.tpm = new TokenBucket(tpmLimit / 60, tpmLimit);
    this.rpm = new TokenBucket(rpmLimit / 60, rpmLimit);
  }

  acquire(estimatedInput: number, maxOutput: number): number | null {
    const reserved = estimatedInput + maxOutput;
    if (!this.tpm.tryConsume(reserved)) return null;
    if (!this.rpm.tryConsume(1)) { this.tpm.adjust(-reserved); return null; }
    return reserved; // returned, not stored, so concurrent requests can't overwrite it
  }

  reconcile(reserved: number, actualInput: number, actualOutput: number) {
    this.tpm.adjust(actualInput + actualOutput - reserved); // charge overruns, refund the rest
  }
}

Estimate input tokens with gpt-tokenizer or js-tiktoken, or use characters ÷ 4 as a fallback. Keep one limiter per tenant so a noisy tenant can’t use up a shared key. With more than one process, move the buckets to Redis: in-memory counters under cluster multiply the real limit.

4. Observability with OpenTelemetry

Use the OpenTelemetry GenAI semantic conventions so calls to OpenAI, Anthropic and local models share one schema. Note that v1.37.0 replaced gen_ai.system with gen_ai.provider.name.

typescript
import { trace, SpanStatusCode } from '@opentelemetry/api';

const tracer = trace.getTracer('ai-middleware');

function tracedCall(provider: ModelProvider, model: string, req: AIRequest, signal: AbortSignal) {
  return tracer.startActiveSpan(`chat ${model}`, async (span) => {
    span.setAttributes({
      'gen_ai.operation.name': 'chat',
      'gen_ai.provider.name': provider.name,
      'gen_ai.request.model': model,
      'app.tenant': req.tenant,
    });
    try {
      const res = await provider.complete(model, req, signal);
      span.setAttributes({
        'gen_ai.usage.input_tokens': res.usage.input,
        'gen_ai.usage.output_tokens': res.usage.output,
        'app.cost_usd': res.costUsd,
      });
      return res;
    } catch (err) {
      span.recordException(err as Error);
      span.setStatus({ code: SpanStatusCode.ERROR });
      throw err;
    } finally {
      span.end();
    }
  });
}

Redact PII before spans and logs leave the process (Pino redaction paths handle the log side). The agent and tool-call conventions are still marked Development, so pin your semconv version.

5. Resilience

Retry 408, 429, 5xx and timeouts. Never retry a 400, because it fails the same way every time. Give each attempt its own AbortSignal deadline, and stop calling a provider after repeated failures.

typescript
const RETRYABLE = new Set([408, 429, 500, 502, 503, 504]);
const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));

async function withRetry<T>(fn: (signal: AbortSignal) => Promise<T>, attempts = 3) {
  for (let i = 0; ; i++) {
    try {
      return await fn(AbortSignal.timeout(15_000));
    } catch (err: any) {
      if (err?.status !== undefined && !RETRYABLE.has(err.status)) throw err;
      if (i === attempts - 1) throw err;
      const base = 250 * 2 ** i;
      await sleep(base + Math.random() * base); // backoff with jitter
    }
  }
}

class CircuitBreaker {
  private failures = 0;
  private openedAt = 0;
  constructor(private threshold = 5, private coolDownMs = 30_000) {}

  canCall() { return this.failures < this.threshold || Date.now() - this.openedAt > this.coolDownMs; }
  success() { this.failures = 0; }
  failure() { if (++this.failures >= this.threshold) this.openedAt = Date.now(); }
}

Fall back across providers within the same tier. Don’t quietly downgrade a reasoning task to a small model; fail and say so.

Putting it together

typescript
class PolicyViolation extends Error {}
class RateLimited extends Error {}

export class AIMiddleware {
  private limiters = new Map<string, DualBucketLimiter>();
  private breakers = new Map<string, CircuitBreaker>();

  constructor(private providers: Map<string, ModelProvider>, private cache: SemanticCache) {}

  async complete(req: AIRequest, ctx = { usesTools: false, personalised: false }) {
    const violation = checkInputPolicy(req.prompt);
    if (violation) throw new PolicyViolation(violation);

    const cacheable = isCacheable(ctx);
    const hit = cacheable && (await this.cache.lookup(req.tenant, req.prompt));
    if (hit) return hit;

    const inputTokens = estimateTokens(req.prompt);
    const limiter = this.get(this.limiters, req.tenant, () => new DualBucketLimiter(200_000, 500));
    const reserved = limiter.acquire(inputTokens, req.maxTokens ?? 1024);
    if (reserved === null) throw new RateLimited();

    for (const model of MODELS[pickTier(req, inputTokens)]) {
      const breaker = this.get(this.breakers, model, () => new CircuitBreaker());
      if (!breaker.canCall()) continue;
      try {
        const res = await withRetry((signal) => tracedCall(this.providers.get(model)!, model, req, signal));
        breaker.success();
        limiter.reconcile(reserved, res.usage.input, res.usage.output);
        res.text = redactPII(res.text);
        if (cacheable) await this.cache.store(req.tenant, req.prompt, res, 3600);
        return res;
      } catch {
        breaker.failure();
      }
    }
    limiter.reconcile(reserved, 0, 0); // nothing consumed, return the budget
    throw new Error('No healthy provider');
  }

  private get<T>(map: Map<string, T>, key: string, make: () => T): T {
    if (!map.has(key)) map.set(key, make());
    return map.get(key)!;
  }
}

The Express route doesn’t know which model answered. The tenant comes from the verified JWT:

typescript
app.post('/api/chat', async (req, res, next) => {
  try {
    const answer = await ai.complete({ tenant: req.user.orgId, prompt: req.body.message, maxTokens: 800 });
    res.json({ text: answer.text });
  } catch (err) {
    if (err instanceof RateLimited) return res.set('Retry-After', '10').status(429).json({ error: 'Token budget exceeded' });
    if (err instanceof PolicyViolation) return res.status(422).json({ error: err.message });
    next(err);
  }
});

Agent middleware and MCP

Agent frameworks expose the same interception points inside the agent loop. LangChain 1.0 has before_model, after_model and wrap_model_call hooks; before_* hooks run in list order and after_* hooks run in reverse, as in Express. The Vercel AI SDK uses wrapLanguageModel:

typescript
import { wrapLanguageModel } from 'ai';
import type { LanguageModelV4Middleware } from '@ai-sdk/provider'; // V2 in AI SDK 5

const redactOutput: LanguageModelV4Middleware = {
  specificationVersion: 'v4',
  wrapGenerate: async ({ doGenerate }) => {
    const result = await doGenerate();
    return {
      ...result,
      content: result.content.map((p) => (p.type === 'text' ? { ...p, text: redactPII(p.text) } : p)),
    };
  },
};

const model = wrapLanguageModel({ model: baseModel, middleware: [cacheMiddleware, redactOutput] });

MCP doesn’t replace any of this. It is the wire protocol agents use to discover and call tools. Middleware decides whether a tool call should happen at all. Treat tool descriptions and results as untrusted input, for the reasons covered in indirect prompt injection.

Which AI gateway to use

Gateway Best at Status in late 2026
LiteLLM 100+ providers, self-hosting Active; PyPI releases 1.82.7 and 1.82.8 were hijacked on 2026-03-24, so pin versions and verify hashes
Bifrost Throughput (~11µs overhead at 5,000 RPS, vendor-reported) Active, Go
Portkey Guardrails, governance, MCP gateway Acquired by Palo Alto Networks
Vercel AI Gateway Next.js and AI SDK apps Active, managed
Helicone Observability Maintenance mode since March 2026; no new signups

Start with one of these rather than the class above. Write your own layer only once you’ve outgrown them; by then you’ll know which of the five jobs you need to own.

Five mistakes to avoid

  1. Calling LLMs from route handlers. No cache, no limit, no cost tracking, no failover.
  2. Hardcoding one model ID. Model choice belongs in config.
  3. Sharing a semantic cache across tenants. Similarity search doesn’t check ownership.
  4. Retrying every error. A retried 400 costs money and still fails.
  5. Logging raw prompts. Redact before anything leaves the process.

Frequently Asked Questions

What is an AI gateway? A proxy between your app and LLM providers that routes, caches, rate-limits by tokens, applies guardrails and records cost per call. “LLM gateway” means the same thing.

Do I need one with a single provider? Not on day one. A thin wrapper like the one above covers budgets, retries and tracing. Add a gateway when you add a second provider or tenant.

How do I rate-limit LLM calls by tokens? Use two buckets per tenant (TPM and RPM). Reserve estimated input plus max_tokens before the call and reconcile with real usage after it.

Is semantic caching safe? It is, with a threshold tuned on your own data, entries scoped per tenant, and no caching of personalised or tool-based answers.