By Malay Mehta, Software Architect and Mentor · 22 min read
From 14 Seconds to 149ms: What Agentic Performance Engineering Actually Looks Like
A field report from two days of load testing a multi-tenant platform with Claude Code, including the parts where the agent was wrong and the moments a human unblocked it in one sentence. Service names, tenant identifiers and hostnames are genericised. All numbers, measurements, code patterns and quoted prompts are verbatim from the actual engagement.
The Setup: Eight Node.js Services, Kubernetes, MySQL Per Tenant
The system is a multi-tenant point-of-sale platform. Eight consolidated Node.js services on Kubernetes, MySQL per tenant behind Aurora, RabbitMQ for async work, Redis for caching. The question was the simple one every platform team eventually has to answer.
How much load can we actually take, and what breaks first?
We had a k6 harness. What we did not have was a production shaped load profile, or any real understanding of where CPU went. We started with a profile built from New Relic production data and a 12 rps baseline.
The first run came back with a p95 of 14,250 ms.
Two days later the same load ran at 149 mswith a 0.12% error rate, and we had a measured capacity ceiling of 10 invoices per second. That is roughly 14 times production's busiest minute, with a precise account of what stopped it there.
This is how that happened, and what the division of labour between human and agent actually looked like. If you have read my post on why vibe coded apps break in production, this is the other side of the same coin. Not AI writing code nobody reviewed, but AI doing investigation work under supervision.
The numbers up front
- p95 latency at 12 rps: 14,250 ms to 149 ms
- Load test cycles: 12, each around 50 minutes
- Distinct root causes found and fixed: 10
- Repositories touched: 8
- Documented corrections where the agent was wrong: 14
- Proven sustained capacity at the end: 10 invoices per second, 63 rps
The Loop We Ran Twelve Times
Each cycle looked like this:
- Human decides what to run and why.
- Agent dispatches, monitors, and waits out a 50 minute run.
- Agent pulls k6 output, New Relic NRQL,
kubectl, CloudWatch and CPU flame graphs, forms a hypothesis, and identifies the constraint. - Human reviews, corrects, decides scope.
- Agent implements the fix across however many repos it touches, updates tests, validates.
- Repeat.
Wall clock time was dominated by step two. The runs themselves do not get faster because an agent is driving them. What collapsed was the thinking, from hours to minutes. That changes the economics of investigation in a way that is easy to miss. When analysis is cheap you stop rationing hypotheses. More on that later, because it cuts both ways.
What Twelve Load Test Runs Actually Found
Twelve runs, twelve different binding constraints. Each one had to be removed before the next became visible. This is the part that makes performance work slow and it is why you cannot parallelise it. You do not know what is behind the current wall until you take it down.
1. The 42% of CPU That Should Not Have Existed
The first CPU flame graph, taken from a live pod via SIGUSR1 and the Chrome DevTools Protocol, showed something absurd.
performAttachment 16.4s of 39.3s non-idle CPU (42%)
garbage collector 29%
application code 0.7%performAttachment ran on every single request and rebuilt the entire Sequelize model layer. Re-defining every model, re-wiring every association, producing a byte for byte identical result each time. The business logic was 0.7% of CPU. The framework was doing 42%.
The fix was a WeakMap keyed on the Sequelize instance rather than the tenant id, because the connection manager expires and recreates instances. A tenant keyed cache would happily hand back models bound to a dead connection.
Result at 12 rps: p95 14,250 ms to 301 ms. DB time 2,604 ms to 50 ms. Thirteen pod restarts became zero.
It also dissolved two invariant failures. Stock conservation and missing stock movements had both looked like data integrity bugs. They were downstream of CPU starvation. The confirmation worker simply never got scheduled. Two correctness bugs that were really one performance bug.
2. The Observability Agent Eating 38.5% of CPU
With the application no longer pathological, the next flame graph showed New Relic accounting for 38.5% of non-idle CPU.
This is the finding that best illustrates why naive profiling misleads you. A self time by leaf view put New Relic at 1.5%, because the agent's work lands in frames attributed to node:internal, (program) and other. Only walking each sample's ancestry revealed the real share, and it reproduced across two independent captures, 36.4% at 5 rps and 38.5% at 12 rps. So it was not sampling noise.
The culprit was Code Level Metrics, on by default since agent v11. A full require() including module resolution, internalModuleStat and readPackageJSON on every Express handler invocation, purely to attach a source file and line number to a span.
Two config lines removed roughly 10% of CPU fleet wide.
Except on two services, where nothing happened at all, because their Docker images never contained the config file. The setting had been silently ignored for months. Nobody had checked, because nobody had a reason to.
3. Five Autoscalers That Had Never Scaled
Every HorizontalPodAutoscaler used AverageValue: 400m as its CPU target. Measured per pod averages were 16 to 84 mCPU.
checkout-service 84 mCPU/pod 400m was 5x above anything measurable
inventory-service 43 mCPU/pod 9x
promotions-service 26 mCPU/pod 15x
auth-service 16 mCPU/pod 25xFive HPAs, none of which had ever scaled on CPU in their lives. They were decoration. Recalibrated to roughly 3 times the measured average, four of five scaled to two replicas on the very next run.
If you have autoscalers in your cluster right now, go and compare the target against what your pods actually use. This takes five minutes and I would bet money on what you find.
4. Resource Requests Backwards by 4x
group used mCPU requested
app services 1701 600 284% of its reservation
newrelic-bundle 66 900 7%
redis 44 200 22%Nodes sat at 95% of allocatable CPU requests while only 53% of real CPU was in use. Observability reserved 23% of the cluster and used 7% of that. Meanwhile 15 of 27 production pods had no CPU request at all, which means QoS BestEffort, first to be evicted, and invisible to the scheduler's arithmetic.
The scheduler admits pods on requests. An over-reservation is a hard wall regardless of how idle the machine actually is. This is the single most common Kubernetes capacity mistake I see in mentoring sessions, and it is almost always invisible until something needs to scale.
5. The Five Minute Error Storm on Every Scale Up
One run failed its error SLO at 0.0254 against a 0.025 target. Roughly half the errors came from one pod.
pod started 12:24:23
readinessProbe /health, initialDelay 10s -> Ready at ~12:24:34
FIRST DB FAILURE 12:24:39 5s after joining the Service
last failure 12:29:42 5 minutes of 500s
startupProbe NONE/health returned a static {status:"ok"} and served as both liveness and readiness, so it could not express "running but not yet able to serve". The process began listening before AWS secrets loaded. Kubernetes added the pod to the Service and 409 requests failed with Could not initialize database instance while its sibling served normally. So the service looked degraded, not broken, which is the worst kind of failure to diagnose.
The underlying reason was worse than a missing await:
static async getEnvironmentValues() {
let interval = setInterval(async () => { ...fetch... }, 10000);
}That async function resolves immediately and the first fetch attempt is ten seconds away. Four services did await it and were simply winning a race they did not know they were in.
The fix was a separate /ready gated on credentials actually being present, wired through a new readiness_path input on the shared deployment action. Liveness deliberately stayed shallow, because a database dependent liveness probe restarts every replica at once on one database blip. Verified live during an HPA scale up: pod held at 0/1, Ready 12 seconds after start instead of 10, and zero database attachment failures fleet wide.
6. The Async Ceiling No HTTP Dashboard Could Show
Then we switched from request rate anchored profiles to invoice rate anchored ones, and found a completely different wall.
At 5 invoices per second every HTTP metric looked superb. p95 103 ms, error rate 0.18%. But the post load drain took 345 seconds where it had been instant.
The drain progress logging, added earlier that day because a run had looked stuck for ten silent minutes, turned an opaque wait into a measurement:
t(s) remaining rate/s
0 574
60 466 1.93
150 322 1.80
240 174 1.70
330 16 2.07
confirmed during run 5994 / 1500s = 4.00/s
invoices created 7716 / 1500s = 5.14/s
net accumulation = 1.15/s
predicted backlog 1722 vs observed 1722 exact matchThe confirmation worker sustained 4.00 per second while 5.14 per second arrived. It fell behind for the entire 25 minute plateau and ended 7.2 minutes in arrears, which means on-hand stock was overstated by around 1,700 invoices at peak. Invisible to every HTTP dashboard.
The cause was one line. { prefetch: 1 } on the confirmation consumer, with no comment explaining it, while sibling consumers in the same codebase used 2, 5 and 10. Each message was around 179 ms of which roughly half was I/O wait, so prefetch 1 left the worker idle half the time by construction, at 9% of one CPU core.
Before raising it we traced every mutation on that path for concurrency safety. Message level dedup, per step markers, SELECT ... FOR UPDATE on customer rows, an atomic Redis counter. All already hardened. And a test comment named the exact scenario:
"two concurrent INVOICE_CONFIRMATION messages for the same invoice (split tenders, prefetch=5) could interleave their read compute write"
Someone had already prepared the ground and never raised the consumer. That happens more often than you would think.
prefetch 1 to 4 plus a pool raise: drain 345 s to 0 s, throughput 4.28 to 6.89 messages per second with per message cost unchanged.
Pure overlap. Nothing got faster, work just stopped queueing single file.
7. The Failure That Needed a Human to Solve
At 15 invoices per second the system collapsed. createP95 30,002 ms, 12.85% errors, 10,930 create failures, drain timeouts on all three tenants.
Everything the agent could see said nothing was saturated. Application CPU at 34% of limit. Zero throttling. Query counts unchanged. And every statement uniformly slow, including a trivial indexed lookup on a tiny table going from 5.7 ms to 109 ms, and Redis at 79 ms, which has no locks and is not the database at all.
The first hypothesis was row lock contention. It fit the pattern, since the two worst affected services both touch customer rows, and it was wrong. Unlocked reference tables were equally slow.
The agent tried to check connection counts. New Relic returned null. It said so, and stated plainly that the case was "strong but circumstantial".
Then:
"rds metrics can be in cloud watch if you want to see"
One sentence. CloudWatch gave the answer immediately, by being the opposite of what anyone expected.
connections DB CPU
sales-2x (passed) 175/199 36.7%
sales-3x (FAILED) 227/240 16.1% DB CPU FELLThe database was less busy during the failure. It was starved of work. And 240 connections was almost exactly the application's own pool ceiling: 15 per tenant times 3 tenants times 5 pods equals 225, plus the worker's 45, against an instance allowing around 2,000.
The app had capped itself at 240 connections while the database idled at 16% CPU.
DB_POOL_MAX 15 to 40 produced a 25x latency improvement. 7,133 ms to 283 ms average, p95 29,500 ms to 922 ms, errors 12.85% to 0.15%.
Sequelize measures a query from acquire() to result, so pool queue wait gets billed to the statement. The queries had never slowed down. They were all standing in the same line. That is why everything, Redis included, looked uniformly broken.
Where the Human Was Indispensable
This is the part most write-ups skip, and it is the most useful part.
Domain Knowledge the Agent Could Not Have
"rds metrics can be in cloud watch if you want to see." The single highest leverage sentence of the session. The agent had exhausted its instrumentation, said so, and labelled its conclusion as inference. One pointer to a data source it had not considered turned a hypothesis into a measurement, and revealed the counter-intuitive truth that no amount of reasoning from application telemetry would have produced.
"DB_POOL_MAX is set to 15 for checkout-service check and confirm." The agent had read the code default of 5 and reported it as fact. Wrong. The deployment overrode it. Trusting source code over runtime configuration is a specific, repeatable failure mode, and it took a human who knew the environment to catch it.
Catching the Agent Contradicting Itself
"dont you think you needed to change inventory-service hpa?"
The agent had just refused to copy a 400m HPA target to two new autoscalers, correctly arguing it was unreachable. Then it left that same unreachable target in place on three existing ones. Same defect, same evidence, inconsistent action.
The prompt forced a re-examination which found something worse. The agent had sized its own two new targets from peak per pod CPU rather than average, 7.3x and 12.4x too high, reproducing the exact bug it was fixing. One question corrected five autoscalers instead of three.
Pushing Back on a Wrong Conclusion
"actually topology should spread to another node where resource is available i dont think we would be constrained with pending state."
The agent had warned about Pending pods using cluster total arithmetic. The pushback was partly right and forced precision. The scheduler tests per node capacity, so cluster totals are the wrong unit. Re-measured, one node had 165 mCPU free and the pod needed 150, so it would schedule.
But the agent also got to correct the mental model on the other side. topologySpreadConstraints with ScheduleAnyway expresses a preference among feasible nodes. It cannot create capacity, and by balancing pods it actually leaves symmetric small gaps that make fragmentation more likely.
Both parties learned something. That exchange produced a better answer than either would have alone, and that is genuinely what good pairing looks like.
Setting the Direction
"i am genuinely concerned that low tps is not sustaining with less resources. let's use new relic, logs, etc to understand if it is realistic."
This reframed the whole engagement. Up to that point the work was "make the test pass". That sentence turned it into "find out whether the resource cost is legitimate", which is what produced the CPU profiling tooling, the 42% discovery, and everything downstream.
An agent will optimise the goal you give it. Give it a narrow goal and you get a narrow answer, fast. That is not a limitation of the agent. It is the whole job of the person directing it.
Scope Control
"make only pool changes to begin with." The agent had proposed two changes together. The human split them to isolate the variable. Textbook discipline the agent had skipped in its eagerness. Then one exchange later: "let's even make the acquire change as its good to have" , having judged it was instrumentation rather than a competing variable.
"i would say that we should for now consider it is max capacity with this resources and that should be fine. I just wanted to see how far system can sustain." Knowing when to stop. The agent would have kept optimising forever.
Correcting the Basic Mistakes
- "man firstly i dont see your changes in my repo git". Work built in the wrong repository entirely.
- "i cant selectall the options you said... i only can select profile from github actions". The agent had designed a workflow the human could not actually drive.
- "unit tests failed". A config change with a stale test assertion.
Where the Agent Was Wrong: Fourteen Documented Corrections
A credible field report has to include this. Here are the ones that matter.
| What the agent claimed | What was actually true |
|---|---|
slow_sql: false will save 3.8% of CPU | It already defaults to false on every deployed agent version. Measured: zero change. The cost was transaction_tracer.record_sql, gated before the slow_sql check |
| A fix was not deployed on two services | Wrong image layout assumption, /app versus /usr/src/app |
| tenant-service needs more resources | It uses 9 mCPU. Its latency breach was per request DNS and TCP connects with no keep-alive agent |
| sales-2x will fail the stock movement invariant | It passed. The projection assumed worker throughput stayed flat |
| Every service has the readiness bug | Only 2 of 6 |
| 3x needs around 4 more nodes | Built on contaminated node CPU figures. 3x was comfortable |
| HPA targets sized from peak CPU | The HPA compares a rolling average, so the targets were 7 to 12 times too high |
There are two systemic lessons in that table.
The agent over-generalised from single observations. "This service has bug X" became "the fleet has bug X" more than once. Each time, checking all six showed a narrower truth.
It misread its own measurements. The most instructive one: dividing a single pod's CPU by the whole service's request rate, producing "187 ms per request, 7.5x worse than checkout-service". Internally consistent, directionally right, numerically wrong by 3x. It took a later cross-check against a different tool to catch.
What made all of this survivable is that every claim came with the command that produced it. A wrong conclusion was auditable rather than load bearing. If you take one thing away from this post about working with agents, take that.
Manual Versus Agentic: An Honest Effort Comparison
Wall clock time was dominated by the runs. Twelve runs at roughly 50 minutes is about 10 hours of pure waiting that no amount of automation removes. What changed is everything around them.
| What the agent produced | Count |
|---|---|
| Load test runs dispatched, monitored, analysed | 12 |
| Repositories modified | 8 |
| New tooling | 3 scripts: capacity report, CPU profiler, CI capture harness |
| Load profiles authored | 9 |
| Documentation | 10 files, 2,112 lines |
| Harness tests kept green | 149 |
| Distinct root causes found and fixed | 10 |
Where the Compression Actually Was
Cross-source correlation. A single analysis routinely joined k6 output, New Relic NRQL across five apps, kubectl pod and HPA state, CloudWatch RDS metrics, V8 CPU profiles, and source code across six repos. Manually that is a lot of context switching and a lot of half remembered query syntax. Each analysis took the agent a handful of minutes.
Flame graph interpretation. Recognising that a 1.5% leaf time share was really 38.5% once you walk sample ancestry required writing a custom analyser. A human would plausibly have accepted the 1.5% and moved on, because the misleading number is the one the standard tooling shows you.
Fan-out fixes. "Disable Code Level Metrics" touched six repositories, two of which needed a Dockerfile change first because their images never contained the config file. Mechanical, tedious, easy to do inconsistently by hand.
Documentation as a by-product. 2,112 lines of analysis written while investigating rather than reconstructed afterwards, including the mistakes, which is exactly the material that never survives into a manual write-up.
An Honest Estimate
A senior performance engineer doing this manually: the twelve runs are fixed at around 10 hours. Analysis at perhaps 1 to 3 hours each is 12 to 36 hours. Implementation across 8 repos with tests, maybe 8 to 16 hours. Building the profiler and capacity tooling, another 8 to 16 hours. Documentation to this depth, realistically, does not get written.
Call it 40 to 80 hours of skilled work, spread over two to three weeks because nobody does this continuously.
We did it in two days of elapsed time, with the human engaged for a fraction of it.
But the comparison is not 40 hours versus 2 days, and I do not want to sell it that way. A human would have found fewer things, because rationing hypotheses is rational when each one costs an hour. And a human would have made different mistakes. Probably fewer arithmetic slips, certainly fewer over-generalisations, and they would have known about CloudWatch on day one.
The right framing is not replacement. It is that the cost of testing a hypothesis fell far enough that you stop choosing between them, and the human job shifts from doing the analysis to directing it and catching the errors. That is the same argument I make about AI-augmented engineering generally. The judgement does not move.
What Made This Work
The agent showed its commands. Every number traceable to the query that produced it. Wrong conclusions were auditable.
It said "I don't know" in the right places. "This is inference, not measurement." "I have not isolated the full 174 ms." "Strong but circumstantial." That honesty is what made the CloudWatch pointer possible. A confident agent would have reported a conclusion instead of a gap, and we would still be chasing row lock contention.
Documentation captured errors, not just findings. The record includes fourteen corrections, and those are the most reusable content in it.
The human corrected rather than accepted. Every one of the five most valuable moments was a human either supplying knowledge the agent lacked or refusing a plausible sounding answer.
Measurement before action, consistently. The best example is a non-change. The agent proposed giving tenant-service more resources, measured first, found 9 mCPU, and retracted. The fix would have been harmless, useless, and would have looked like progress. Most performance work I get called into is full of that kind of change.
Where It Ended
Proven sustained capacity: 10 invoices/sec (63 rps)
Production's busiest minute: 0.71 invoices/sec
Headroom: ~14xThe thing that stops it is pleasingly boring. With everything fixed, the limit is HPA replica ceilings pinned by a CPU request budget on a two node cluster, roughly a third of which is consumed by workloads outside the system under test. Application CPU is 37%. The cluster is oversubscribed on reservations while genuinely half idle.
That is a capacity planning decision, not an engineering problem. It is the best place a performance investigation can end.
One defect remains open and is deliberately called out above the performance work. HTTP idempotency writes its dedup record after the response is flushed, so a retry landing on a sibling pod creates a second invoice. Three failures in eight runs, and the probability rises with replica count. It costs money. Everything else costs milliseconds.
What I Would Tell You Before You Try This
If you are planning to point an agent at a performance problem, here is the short version of what actually mattered.
- Make it show the command behind every number. A claim you cannot re-run is not evidence, it is a guess with good grammar.
- Give it the data sources it cannot discover. CloudWatch, the deployment overrides, the reason a config value is what it is. That is your unfair advantage in the loop.
- Watch for the fleet generalisation. One service having a bug is not six services having a bug, and the agent will make that jump.
- Change one thing at a time even when the agent wants to bundle. You are protecting the attribution, not the schedule.
- Measure before you fix. The non-change is often the most valuable output, and it is the one an eager agent will never propose.
- Decide when to stop. An agent has no sense of good enough. That is yours to supply.
None of the ten root causes above needed a clever engineer to fix. They needed someone to look. The agent made looking cheap enough that we looked twelve times instead of twice, and that is the whole story.
Stuck on a performance problem you cannot explain?
I help engineers find the real constraint instead of guessing at it. Profiling, capacity work, Kubernetes resourcing, and the database problems that only show up under load. 726+ sessions, 5.0 rating.
Get notified when I publish new posts
Architecture, system design, and production engineering.
Related reading
- Get a second pair of eyes on your performance problem
Free intro call to scope where the time is actually going.
- System design for engineers who ship, not interview
- Where AI-generated database schemas fall apart
- How I use AI in engineering work