Below you will find pages that utilize the taxonomy term “Performance”
Benchmarking MCP Servers and Stateful API Workflows: Latency, Throughput and Tokens per Call
Point wrk at an MCP server and you get a clean report: every response a 200. Some of those 200s carry isError: true results, and nothing in the output says what the tool definitions cost the model on every turn. The tool is measuring requests per second. An agent runs a workflow, and what the server costs it is latency per call times calls per task, plus the tokens each response and each definition puts into its context.
Diagnose a Slow API From One Request: DNS, Connect, TLS, Server Wait and Transfer
“The API is slow” starts a hunt. Somebody opens the APM dashboard, somebody else greps the load balancer logs, a third person checks whether the database is on fire. An hour later the team has three theories and no measurement. One request, timed phase by phase, answers the first question in under a second: whether the time went to name lookup, the TCP connect, the TLS handshake, the server or the transfer of the body.
Stop Reparsing the Same Big JSON Documents: Persist Them as Indexed Binary Instead
A service loads a 40 MB product catalog from disk every time a worker starts. It parses the JSON into objects, builds a map from SKU to product, and answers price lookups until the next deploy. The parse costs seconds of startup. The object tree often costs several times the file’s size in memory, and every worker holds its own copy. A typical request touches two fields of one product. Multiply that by every deploy, every autoscale event and every cold start. Nothing here is broken. The format was built for exchange and it’s being used as a database.
Trimming 100 KB API Responses to the Five Fields Your Client Uses, at the Proxy
Take a mobile screen that lists orders: a customer name, a status, a total, a date. The endpoint behind it returns 100 KB for one page, because the vendor’s schema carries about sixty fields per order, plus nested addresses, hypermedia links and an audit trail. The phone downloads all of it over whatever connection it has, parses all of it, and keeps five fields. You can’t change the vendor’s API. You can change what reaches the phone.
gRPC in Production: What the Documentation Doesn't Tell You
gRPC documentation is thorough on the protocol’s features and sparse on the operational realities of running it at production scale. The gap between the getting-started experience and the production experience is wide enough to have surprised most of the teams that have made the journey. The surprises are not fatal. They are the kind that would have been useful to know before the architecture decision was made.
The gRPC pitch is compelling: Protocol Buffers serialization that is faster and smaller than JSON, HTTP/2 multiplexing that reduces connection overhead, generated client and server stubs that eliminate serialization bugs, strong typing that catches integration errors at compile time rather than runtime. All of these benefits are real. All of them come with operational requirements that the pitch does not emphasize.
The GraphQL N+1 Problem and How to Actually Fix It
The N+1 query problem is the GraphQL performance issue that every team encounters and few solve completely before it causes a production incident. Its mechanics are straightforward. Its solutions are well-documented. Its persistence in production systems reflects the gap between understanding a problem and implementing a solution that holds under all the query patterns a flexible API allows.
A GraphQL query that requests a list of posts and the author of each post produces, in a naive resolver implementation, one database query to fetch the posts and one database query per post to fetch each author. A query requesting 100 posts with their authors produces 101 database queries. A query requesting 1,000 posts produces 1,001. The number of database queries grows linearly with the number of items in the list — hence N+1, where N is the list length and 1 is the initial list query.
Pagination Strategies for Large Datasets and Why Offset Pagination Fails
Offset pagination — the pattern where a consumer requests a page by specifying how many records to skip — is the default choice for most APIs because it maps naturally to SQL’s LIMIT and OFFSET clauses and allows consumers to request any page directly by number. It is also the pagination strategy that fails most visibly at scale and produces the most confusing behavior when underlying data changes between page requests.