Why Your API Is Slow — Understanding P95 and P99 Latency
Your API has an average response time of 120ms.
Looks good.
Then you check P99.
"2.8 seconds."
Suddenly, the system doesn't look quite as fast.
"The Problem With Averages"
Average response time can hide slow requests.
P50 tells you where 50% of requests fall.
P95 covers 95%.
P99 covers 99%.
If P50 is 100ms but P99 is 2.8 seconds, most users get a fast response — but some are waiting much longer.
At one million requests, even 1% means 10,000 slow requests.
"Why Is P99 High?"
High P99 often means something only becomes slow under certain conditions.
It could be:
A slow database query.
A cache miss.
An exhausted connection pool.
Database lock contention.
A slow external API.
Traffic or resource spikes.
That's why optimizing random code isn't the answer.
"Trace the Slow Requests"
A high P99 tells you there is a problem.
Tracing tells you where it is.
Request: 2,800ms
Application: 80ms
Database: 120ms
External API: 2,600ms
Now you know where to investigate.
The goal isn't simply to make the average faster.
It's to understand why your slowest requests are slow.
So when your dashboard says:
"Average response time: 120ms"
Don't stop there.
Check P95.
Check P99.
Your average tells you how the system usually performs.
Your tail latency tells you where the problems are hiding.




Comments (0)
Log in to join the discussion.
No comments yet. Be the first.