Make for APIs: Rerun Only the Steps Downstream of an Endpoint That Changed
A nightly job pulls a users endpoint and an orders endpoint, normalizes both, joins them on customer ID and renders a report. On most nights the upstream data hasn’t changed since the last run. The job doesn’t know that, so it downloads and parses everything, reruns every transform and writes the same report again. If the API bills per call, or one step takes twenty minutes, you pay for the same answer every night.
Mapping an Unfamiliar API by Following the IDs Between Its Endpoints
You inherit an integration with a partner API that has 140 endpoints, documented in alphabetical order. GET /accounts sits next to GET /adjustments, and POST /refunds is a long scroll away from the GET /charges/{id} it depends on. Nobody wrote down that a refund hangs off a charge, which hangs off an order, which may or may not have a customer. You find out the slow way: call an endpoint, copy an ID out of the response, paste it into the next call, and keep a diagram in a notebook that’s out of date by Thursday.
MCP Tool Calls Are Plain HTTP Now: Route Them With a Small Proxy Before Buying a Gateway
A platform team runs four MCP servers for its coding agents: GitHub, internal docs search, a read-only Postgres server and the ticketing system. Someone wants per-tool rate limits and a log of which agent called what, so a gateway evaluation goes on the calendar. Nobody has asked the cheaper question yet, which is how much of that the reverse proxy they already run can handle.
Since July, quite a lot. The 2026-07-28 revision made MCP stateless at the protocol layer: the initialize handshake is gone, the Mcp-Session-Id header is gone, and every request carries its own protocol version, client info and capabilities. Two new headers, Mcp-Method and Mcp-Name, copy the JSON-RPC method and (for a tool call) the tool name out of the body and into headers, where proxies already look. Without sessions, an MCP server behaves much like any other HTTP API, which is why the difference between an API and MCP now comes down to who reads the docs. Before the revision, a proxy needed a body-parsing module to learn which tool was being called, and usually session affinity to keep each client on the backend holding its session. Now stock nginx covers the basics:
Most MCP Token Waste Is in Tool Results: Put a Deterministic Reducer Between Server and Model
An agent calls a code-search tool and gets back a hundred hits. Each hit carries dozens of fields: node IDs, a URL for every related resource, avatar links, permission flags. The model needed three of them, a repo, a path and a snippet. The rest now sits in the context window for the remainder of the session, and the model reads past it on every later turn.
Tool definitions get most of the attention in agent token costs, and they’ve earned it. A server that exposes 150 tools puts 150 schemas in front of the model on every turn. Results are the other half of the bill, and they have fewer standard answers. An ordinary API client ignores the fields it doesn’t use, and ignoring is free. A model pays to read every token it’s handed.
Package a Failed API Request Into One File Anyone Can Replay Locally
A customer’s checkout returns a 500. Support pastes the request ID into the ticket, and the engineer on call finds the log line: KeyError: 'tax_region' in the pricing module. They send the same request locally and get a 200. Of course they do. Their database has no customer with a null tax region, the feature flag that routes to the new tax engine is off in development, the rates service answers differently today, and the clock is a day later. The bug is a function of all of that, and the ticket contains none of it.
Record an Agent's MCP Traffic Once, Then Replay It in CI Without Servers or Credentials
Your agent test passes on your laptop and fails in CI. The laptop has a token for the issue tracker’s MCP server; CI doesn’t, and shouldn’t. Even with a token, the tracker listed three open bugs yesterday and lists four today, so the agent’s summary changes and the assertion on it breaks.
Web developers dealt with this years ago. VCR (Ruby), Polly.js, Betamax and go-vcr record real HTTP responses into a file the first time a test runs, then serve them from that file on every run after. The file is called a cassette. MCP needs the same thing: a proxy that sits between an agent and its MCP servers, writes every exchange into one cassette, and plays it back later with no server, no credentials and no network. mcprec below is an illustrative name for a tool you’d have to build.
Routing LLM Requests by Token Count, Privacy Tag and Cost Ceiling From One Config File
LLM routing usually starts as if statements. One service picks the cheap model for ticket summaries. Another hardcodes the strong model because it needs tool calling. A third has a comment reading “never send this to the cloud” directly above a fallback that sends it to the cloud when the local server times out. Changing a model, a price or a provider means editing all three.
The router worth building is a small one: ordered rules over properties you can compute before the request leaves, written like an nginx config in a single file, with no database and no UI. Input size, whether tools or images are present, a data classification tag, a cost ceiling for the single request, which upstreams are healthy. One binary on a cheap VPS reads the file, accepts the OpenAI request shape and picks where each request goes. The rule that justifies the exercise is that privacy routing fails closed.
Stop Reparsing the Same Big JSON Documents: Persist Them as Indexed Binary Instead
A service loads a 40 MB product catalog from disk every time a worker starts. It parses the JSON into objects, builds a map from SKU to product, and answers price lookups until the next deploy. The parse costs seconds of startup. The object tree often costs several times the file’s size in memory, and every worker holds its own copy. A typical request touches two fields of one product. Multiply that by every deploy, every autoscale event and every cold start. Nothing here is broken. The format was built for exchange and it’s being used as a database.
The Best Small Infrastructure Tools Reduce Data Near the Source Instead of Storing More
A gigabyte of disk costs almost nothing. The same gigabyte sent to a log vendor that prices by ingestion costs real money, and pasted into a model’s context window it costs money and answer quality at the same time. Storage is cheap per byte. Everything that touches the bytes afterward is not: ingestion fees, query time, bandwidth on a thin link, a context window that holds only so much, and the attention of whoever has to read the result. Observability is the oldest place this shows up, and agents have just made it louder.
The New MCP Spec Caches Tool Lists but Not Tool Calls, and a Caching Proxy Fills the Gap
An agent works through a ticket and asks a docs-search tool the same question on turn 3, again on turn 14, and again the next morning in a fresh session. Each repeat costs a charge against the upstream’s rate limit and a wait with the model idle, for an answer that never changed. Agents retry after errors and start every session by looking up what the last one already knew.