Below you will find pages that utilize the taxonomy term “API Tooling”
A Tiny ETL Binary Competes With curl, jq and SQLite in a Cron Job, So Build It That Small
The real competitor is a shell script in a crontab, and it usually looks like this:
curl -s "https://api.example.com/v1/launches?limit=100" \
| jq -r '.data[] | [.id, .name, .net] | @csv' \
| sqlite3 -csv launches.db ".import /dev/stdin launches"
It works on the day you write it. Then the API answers 429 and curl pipes an error page into jq. Or the API has a second page, and the script never asks for it. When the job dies halfway, the rerun inserts the same rows again or trips over the primary key, depending on how the table was made. A field that starts arriving as a string goes unnoticed until a chart looks wrong. Retries, backoff, pagination, incremental state, idempotent writes and schema drift: that’s the list, and shell scripts get every item on it wrong in predictable ways. (To be fair, curl --retry covers the first two.)
An AI Agent Flight Recorder Belongs in One Portable File, the Way HAR Did It for HTTP
An agent edits the wrong config file, then spends forty minutes trying to repair its own repair. The user files a bug with a screenshot of the last message. The maintainer asks for logs, and the logs live in four places: the provider’s usage page, the tool server’s stdout, the framework’s debug output, and a terminal scrollback that closed with the window. Nobody can say what the agent saw on turn six.
Benchmarking MCP Servers and Stateful API Workflows: Latency, Throughput and Tokens per Call
Point wrk at an MCP server and you get a clean report: every response a 200. Some of those 200s carry isError: true results, and nothing in the output says what the tool definitions cost the model on every turn. The tool is measuring requests per second. An agent runs a workflow, and what the server costs it is latency per call times calls per task, plus the tokens each response and each definition puts into its context.
Diagnose a Slow API From One Request: DNS, Connect, TLS, Server Wait and Transfer
“The API is slow” starts a hunt. Somebody opens the APM dashboard, somebody else greps the load balancer logs, a third person checks whether the database is on fire. An hour later the team has three theories and no measurement. One request, timed phase by phase, answers the first question in under a second: whether the time went to name lookup, the TCP connect, the TLS handshake, the server or the transfer of the body.
Finding API Endpoints, Tables and Flags Nobody Uses Takes Code and Traffic Together
A team wants to delete GET /v1/invoices/{id}/legacy-pdf. A search across the main repo finds no caller, and the gateway logs show zero hits in the last 30 days. The route goes out in a cleanup PR. On the first business day of the next quarter, a partner’s reconciliation job starts failing: it calls that endpoint four times a year, and nobody on the team knew the partner was still there.
Fuzzing MCP Servers: Generate Bad Arguments From the Tool Schema and Watch What Breaks
You test an MCP server by chatting with it. Ask for the open bugs, get the open bugs, ship it. The model never sends limit: "ten", an empty path or a 200 KB query while you’re watching, so those paths stay dark until a real session hits one. Say it’s the limit. The handler throws, the framework wraps the exception in a 40 KB stack trace, and the model reads all of it, adjusts, and retries with the same bug in a new shape.
Give Any API a History: Poll It, Hash It and Query Old Versions With SQL
Ask a launch schedule API when a rocket flies and you get one date. Ask what the date was last Tuesday, or how many times it has moved, and there’s no endpoint for that. Most APIs describe the present. Prices, timetables, government datasets and status pages all change in place, and the old value is gone the moment the new one is written. That’s a pity, because launch dates slip so often that the list of changes in a launch schedule is more interesting than the current date.
Infer Your API's Real Contract From Traffic, Then Diff It Against the Docs
Your OpenAPI file marks shipping_address as required on GET /orders/{id}. After Tuesday’s deploy, about one response in 300 leaves it out, because a new code path for orders created by the import job skips the field. The docs still say required. The tests still pass, since nobody wrote one for import-job orders. A client with a strict deserializer starts throwing on 0.3% of order pages, and the first report you get is “sometimes the order screen is blank”.
Keeping Third-Party API Responses in SQLite Gets You a Cache, an Offline Mode and a History
The first version of an API cache is a dictionary with a timeout. The second is Redis holding a JSON string under a key built from the URL. The third gets written after an incident, when somebody needs to know what the weather provider returned on Tuesday and the cache has already replaced it with Wednesday’s answer. Every app that depends on an outside API walks the same path: a cache, then a retry layer, then a debugging log, then a wish that it had kept the old responses.
LLM Response Caching Pays Off in CI, Evals and Agent Retries, Not in Chat
A repository has 300 integration tests that each send a prompt to a model. The suite runs on every push to every branch, then again in the merge queue. Most of those runs don’t touch a prompt (the diff was a stylesheet or a migration), so the model receives the same 300 requests it saw an hour earlier, writes roughly the same 300 answers, and the provider bills every one. Nobody chose that. It’s what happens when a test calls a live API.
Make for APIs: Rerun Only the Steps Downstream of an Endpoint That Changed
A nightly job pulls a users endpoint and an orders endpoint, normalizes both, joins them on customer ID and renders a report. On most nights the upstream data hasn’t changed since the last run. The job doesn’t know that, so it downloads and parses everything, reruns every transform and writes the same report again. If the API bills per call, or one step takes twenty minutes, you pay for the same answer every night.
Mapping an Unfamiliar API by Following the IDs Between Its Endpoints
You inherit an integration with a partner API that has 140 endpoints, documented in alphabetical order. GET /accounts sits next to GET /adjustments, and POST /refunds is a long scroll away from the GET /charges/{id} it depends on. Nobody wrote down that a refund hangs off a charge, which hangs off an order, which may or may not have a customer. You find out the slow way: call an endpoint, copy an ID out of the response, paste it into the next call, and keep a diagram in a notebook that’s out of date by Thursday.
MCP Tool Calls Are Plain HTTP Now: Route Them With a Small Proxy Before Buying a Gateway
A platform team runs four MCP servers for its coding agents: GitHub, internal docs search, a read-only Postgres server and the ticketing system. Someone wants per-tool rate limits and a log of which agent called what, so a gateway evaluation goes on the calendar. Nobody has asked the cheaper question yet, which is how much of that the reverse proxy they already run can handle.
Since July, quite a lot. The 2026-07-28 revision made MCP stateless at the protocol layer: the initialize handshake is gone, the Mcp-Session-Id header is gone, and every request carries its own protocol version, client info and capabilities. Two new headers, Mcp-Method and Mcp-Name, copy the JSON-RPC method and (for a tool call) the tool name out of the body and into headers, where proxies already look. Without sessions, an MCP server behaves much like any other HTTP API, which is why the difference between an API and MCP now comes down to who reads the docs. Before the revision, a proxy needed a body-parsing module to learn which tool was being called, and usually session affinity to keep each client on the backend holding its session. Now stock nginx covers the basics:
Most MCP Token Waste Is in Tool Results: Put a Deterministic Reducer Between Server and Model
An agent calls a code-search tool and gets back a hundred hits. Each hit carries dozens of fields: node IDs, a URL for every related resource, avatar links, permission flags. The model needed three of them, a repo, a path and a snippet. The rest now sits in the context window for the remainder of the session, and the model reads past it on every later turn.
Tool definitions get most of the attention in agent token costs, and they’ve earned it. A server that exposes 150 tools puts 150 schemas in front of the model on every turn. Results are the other half of the bill, and they have fewer standard answers. An ordinary API client ignores the fields it doesn’t use, and ignoring is free. A model pays to read every token it’s handed.
Package a Failed API Request Into One File Anyone Can Replay Locally
A customer’s checkout returns a 500. Support pastes the request ID into the ticket, and the engineer on call finds the log line: KeyError: 'tax_region' in the pricing module. They send the same request locally and get a 200. Of course they do. Their database has no customer with a null tax region, the feature flag that routes to the new tax engine is off in development, the rates service answers differently today, and the clock is a day later. The bug is a function of all of that, and the ticket contains none of it.
Record an Agent's MCP Traffic Once, Then Replay It in CI Without Servers or Credentials
Your agent test passes on your laptop and fails in CI. The laptop has a token for the issue tracker’s MCP server; CI doesn’t, and shouldn’t. Even with a token, the tracker listed three open bugs yesterday and lists four today, so the agent’s summary changes and the assertion on it breaks.
Web developers dealt with this years ago. VCR (Ruby), Polly.js, Betamax and go-vcr record real HTTP responses into a file the first time a test runs, then serve them from that file on every run after. The file is called a cassette. MCP needs the same thing: a proxy that sits between an agent and its MCP servers, writes every exchange into one cassette, and plays it back later with no server, no credentials and no network. mcprec below is an illustrative name for a tool you’d have to build.
Routing LLM Requests by Token Count, Privacy Tag and Cost Ceiling From One Config File
LLM routing usually starts as if statements. One service picks the cheap model for ticket summaries. Another hardcodes the strong model because it needs tool calling. A third has a comment reading “never send this to the cloud” directly above a fallback that sends it to the cloud when the local server times out. Changing a model, a price or a provider means editing all three.
The router worth building is a small one: ordered rules over properties you can compute before the request leaves, written like an nginx config in a single file, with no database and no UI. Input size, whether tools or images are present, a data classification tag, a cost ceiling for the single request, which upstreams are healthy. One binary on a cheap VPS reads the file, accepts the OpenAI request shape and picks where each request goes. The rule that justifies the exercise is that privacy routing fails closed.
Stop Reparsing the Same Big JSON Documents: Persist Them as Indexed Binary Instead
A service loads a 40 MB product catalog from disk every time a worker starts. It parses the JSON into objects, builds a map from SKU to product, and answers price lookups until the next deploy. The parse costs seconds of startup. The object tree often costs several times the file’s size in memory, and every worker holds its own copy. A typical request touches two fields of one product. Multiply that by every deploy, every autoscale event and every cold start. Nothing here is broken. The format was built for exchange and it’s being used as a database.
The Best Small Infrastructure Tools Reduce Data Near the Source Instead of Storing More
A gigabyte of disk costs almost nothing. The same gigabyte sent to a log vendor that prices by ingestion costs real money, and pasted into a model’s context window it costs money and answer quality at the same time. Storage is cheap per byte. Everything that touches the bytes afterward is not: ingestion fees, query time, bandwidth on a thin link, a context window that holds only so much, and the attention of whoever has to read the result. Observability is the oldest place this shows up, and agents have just made it louder.
The New MCP Spec Caches Tool Lists but Not Tool Calls, and a Caching Proxy Fills the Gap
An agent works through a ticket and asks a docs-search tool the same question on turn 3, again on turn 14, and again the next morning in a fresh session. Each repeat costs a charge against the upstream’s rate limit and a wait with the model idle, for an answer that never changed. Agents retry after errors and start every session by looking up what the last one already knew.
Trimming 100 KB API Responses to the Five Fields Your Client Uses, at the Proxy
Take a mobile screen that lists orders: a customer name, a status, a total, a date. The endpoint behind it returns 100 KB for one page, because the vendor’s schema carries about sixty fields per order, plus nested addresses, hypermedia links and an audit trail. The phone downloads all of it over whatever connection it has, parses all of it, and keeps five fields. You can’t change the vendor’s API. You can change what reaches the phone.
Turning Ten Minutes of Production Traffic Into an API Regression Suite
You’re about to refactor the billing endpoints. The service has a few dozen tests, mostly happy paths, and nobody trusts them to catch a changed rounding rule or a renamed field. Writing better ones by hand means reading every handler and inventing inputs. Meanwhile production receives thousands of real inputs a minute, and each one comes labelled with the response your current code gives.
Capturing that traffic is the easy part; a proxy, a packet tap or a log line with the body in it will do. The product is everything after. Ten minutes of traffic holds thousands of near-duplicate requests and a handful that exercise something different, and a tool is only useful if it can tell them apart. Cluster by endpoint, request shape and response shape. Pick representatives that cover the status codes and branches. Assert only on fields that are stable. Mock the downstream calls. What comes out is a few dozen characterization tests (Michael Feathers’ name for tests that pin down what code does today).
Five SDK Generators Compared: Speakeasy, Stainless, Fern, APIMatic, and OpenAPI Generator
Speakeasy has published a head-to-head comparison of the five most widely deployed SDK generators for OpenAPI specifications, covering Speakeasy itself alongside Stainless, Fern, APIMatic, and the open-source OpenAPI Generator. The evaluation spans language coverage, runtime type safety, dependency footprint, OpenAPI fidelity, enterprise feature sets, and deployment flexibility — the criteria that now determine whether an SDK generator clears enterprise procurement rather than just developer preference.
SDK generators have moved from convenience tooling to infrastructure. API-first companies, including those building for AI agent integrations, now treat generated client libraries as a production artifact, and the differences between generators surface in SOC 2 audits and supply chain reviews rather than only in developer experience surveys.