Below you will find pages that utilize the taxonomy term “AI Gateway”
LLM Response Caching Pays Off in CI, Evals and Agent Retries, Not in Chat
A repository has 300 integration tests that each send a prompt to a model. The suite runs on every push to every branch, then again in the merge queue. Most of those runs don’t touch a prompt (the diff was a stylesheet or a migration), so the model receives the same 300 requests it saw an hour earlier, writes roughly the same 300 answers, and the provider bills every one. Nobody chose that. It’s what happens when a test calls a live API.
Routing LLM Requests by Token Count, Privacy Tag and Cost Ceiling From One Config File
LLM routing usually starts as if statements. One service picks the cheap model for ticket summaries. Another hardcodes the strong model because it needs tool calling. A third has a comment reading “never send this to the cloud” directly above a fallback that sends it to the cloud when the local server times out. Changing a model, a price or a provider means editing all three.
The router worth building is a small one: ordered rules over properties you can compute before the request leaves, written like an nginx config in a single file, with no database and no UI. Input size, whether tools or images are present, a data classification tag, a cost ceiling for the single request, which upstreams are healthy. One binary on a cheap VPS reads the file, accepts the OpenAI request shape and picks where each request goes. The rule that justifies the exercise is that privacy routing fails closed.
AI Platforms for Designing APIs in 2026: Spec Editors, SDK Generators, MCP Builders and AI Gateways Reviewed
Ask two developers in 2026 what platform they use to design an API, and you will get two answers that have almost nothing in common. One of them means the tool where they write the OpenAPI document, lint it, mock it and publish the reference docs. The other means the layer that sits between their application and a dozen model providers, routing requests to whichever LLM is cheapest or still up. Both groups call it “API design.” Both groups are right, because the two stacks have quietly grown into each other.