Key takeaways
- Response time is the full round trip: from the moment a request is sent until the complete response arrives back.
- Averages hide the slow tail. Report the median alongside p95 or p99, because those are what your slowest users actually feel.
- A response time only means something next to the load, location and environment that produced it, so test it under realistic traffic.
Is Your Infrastructure Ready for Global Traffic Spikes?
Unexpected load surges can disrupt your services. With LoadFocus’s cutting-edge Load Testing solutions, simulate real-world traffic from multiple global locations in a single test. Our advanced engine dynamically upscales and downscales virtual users in real time, delivering comprehensive reports that empower you to identify and resolve performance bottlenecks before they affect your users.
What is response time?
Response time is the time from the moment a client sends a request until it has received the complete response. In performance testing it is measured per transaction: one page load, one API call, one database query, or one step in a user journey such as “add to cart”.
It is the number users care about most, because it is the wait they actually experience. A server that processes a request in 40 ms is no comfort to a user who waits two seconds because the request crossed an ocean, sat in a queue and then downloaded a 3 MB payload.
Broken down, a single response time is roughly:
Think your website can handle a traffic spike?
Fair enough, but why leave it to chance? Uncover your website’s true limits with LoadFocus’s cloud-based Load Testing for Web Apps, Websites, and APIs. Avoid the risk of costly downtimes and missed opportunities—find out before your users do!
- Network time: DNS lookup, TCP and TLS handshakes, and the travel time of the request and response over the network.
- Waiting time: time the request spends queued behind others when the server, a thread pool or a database connection pool is busy.
- Processing time: the work your application does, including calls to databases, caches and third party services.
- Transfer time: how long it takes to send the response body back, which grows with payload size.
Under light load, waiting time is close to zero. Under heavy load it often becomes the largest part, which is why response time has to be tested with realistic traffic and not just measured once from your laptop.
Response time vs latency, TTFB and throughput
These terms get mixed up constantly, and mixing them up leads to the wrong conclusions from a test.
| Metric | What it measures | Typical use |
|---|---|---|
| Response time | Request sent until the full response is received | What the user waits for; the main pass or fail metric in load tests |
| Latency | The delay before data starts moving, mostly network travel time | Explaining why the same server is slower from some regions |
| Time to first byte (TTFB) | Request sent until the first byte of the response arrives | Separating server and network slowness from payload size |
| Throughput | How many requests (or bytes) the system completes per second | Capacity: how much work the system does, not how fast each request is |
One caution when reading reports: in JMeter and LoadFocus results, “Latency” means the time to the first byte of the response, so it includes server processing and not only network time.
Response time and throughput are best read together. See what throughput means in performance testing for the other half of the picture, and what latency is for the network side.
LoadFocus is an all-in-one Cloud Testing Platform for Websites and APIs for Load Testing, Apache JMeter Load Testing, Page Speed Monitoring and API Monitoring!
Why the average response time misleads you
Say a test sends 100 requests. 90 of them come back in 200 ms and 10 of them take 2 seconds.
- The average is 380 ms, a value that not a single request actually took.
- The median (p50) is 200 ms, which is what a typical request experiences.
- The 95th percentile (p95) is 2 seconds, ten times the median.
Note that p90 is still 200 ms here: the slow 10% sits just above it. Which percentile you watch matters.
The average makes this system look fine. The percentiles show a real problem that one request in ten will hit, and on a page that makes seven or more such requests, most page views will hit at least one. That is why performance teams report the median together with p90, p95 or p99, and set their targets on the percentiles. We go deeper on this in why percentiles are more useful than averages.
Response time metrics to report
- Average: total response time divided by the number of requests. Easy to compute, easy to be misled by.
- Median (p50): half of requests were faster, half slower.
- Percentiles (p90, p95, p99): the time within which 90, 95 or 99 percent of requests finished. These describe the slow tail.
- Minimum and maximum: the extremes. Useful for spotting outliers, not for setting targets.
- Standard deviation: how spread out the times are. A high value means inconsistent performance, even when the average looks good.
What is a good response time?
There is no single number, but there are well established reference points.
For people waiting on an interface, Jakob Nielsen’s three response time limits still hold up:
- 0.1 seconds: feels instant. The user feels the system reacted directly to what they did.
- 1 second: the user notices the delay, but their flow of thought stays uninterrupted.
- 10 seconds: about the limit for keeping attention on the task. Beyond this, people switch away.
For web pages, Google’s web.dev guidance treats a time to first byte of 0.8 seconds or less as good. The server response is only the start of a page load, so a slow TTFB leaves little room for everything that follows.
For APIs, set targets per endpoint rather than one global number. A search endpoint and a report export do very different amounts of work. A common starting point is a p95 target for each critical endpoint, based on what the calling page or service can tolerate, then tightened as you learn what the system can actually deliver.
And if a tool reports something like 0.03 ms, be suspicious rather than pleased. Numbers that small usually mean you measured a cache hit, a local call or an error that returned immediately, not a real round trip over the network.
How to test response time
Response time testing measures how long a system takes to answer requests under a given load, and checks those times against targets such as a p95 limit. It is usually part of load testing.
Measuring one request tells you very little. The useful question is how response time behaves as traffic grows. A practical sequence:
- Pick the transactions that matter. Login, search, checkout, the API calls your mobile app makes on launch. Test user journeys, not just the homepage.
- Get a baseline with light load. Run a few virtual users first. This is the best case, and everything later is compared against it.
- Ramp up gradually. Increase users over several minutes instead of all at once, so you can see where response times start to climb. The ramp-up time guide explains how to choose it.
- Watch response time against the number of users. Healthy systems stay flat and then bend upward at some point. That bend, where throughput stops growing and response times start rising, is your practical capacity.
- Test from where your users are. Network time from another continent can add hundreds of milliseconds. Running load from several regions shows what users there actually experience; see how network latency affects load test accuracy.
- Check errors next to timings. A fast response time with a rising error rate is not good news. Failed requests often return quickly and pull the numbers down. Timeouts do the opposite and inflate the tail.
- Repeat under the same conditions. Same script, same load profile, same regions. Only then can you compare one release with the next.
Pushing past the bend on purpose, to find where the system breaks, is stress testing. The difference is covered in load testing vs stress testing.
Common causes of slow response times
- Database queries: missing indexes, N+1 query patterns and lock contention are behind a large share of slow endpoints.
- Exhausted pools: too few threads, workers or database connections, so requests queue even though CPU looks fine.
- Slow dependencies: a third party API or internal service that is slow under load, and passes that delay on to every caller.
- Large payloads: uncompressed responses, oversized images or APIs that return far more data than the client uses.
- No caching: repeated work for data that rarely changes.
- Distance: users far from your servers with no CDN in front of static content.
- Resource limits: CPU or memory saturation, garbage collection pauses and memory leaks that only show up after a test has run for a while.
Measuring response time with LoadFocus
LoadFocus runs load tests from cloud locations around the world, with no infrastructure to set up, and it can also run your existing Apache JMeter and k6 scripts. For response time, test reports give you:
- Response time percentiles (median, 90th, 95th and 99th) next to latency, hits per second, throughput and error rate.
- A response time versus users chart, so you can see exactly where the curve bends as load increases.
- Optional thresholds per test, such as a maximum p95 and p99 response time and a maximum error rate, so every run is judged pass or fail automatically.
- Comparison with previous runs, to catch a release that made things slower.
To try it on your own site, run a free load test in your browser, or use the free API load test for an endpoint. To keep an eye on response times between tests, API monitoring checks your endpoints on a schedule and tracks average and p95 response time over time.
Frequently Asked Questions
What does response time measure?
The time from a client sending a request until it receives the complete response, for one transaction such as a page load or an API call. It includes network time, time spent waiting in queues, server processing and the transfer of the response.
What is a good response time?
For interactive use, under 0.1 seconds feels instant and under 1 second keeps the user’s flow. Google’s web.dev guidance treats a time to first byte of 0.8 seconds or less as good. For APIs, set a p95 target per endpoint based on what its callers can tolerate.
What are the 3 important limits for response times?
Jakob Nielsen’s limits: 0.1 seconds feels instant, 1 second keeps the user’s flow of thought uninterrupted, and 10 seconds is about the limit for keeping their attention on the task.
How is response time different from latency?
Latency is the delay before data starts moving, mostly network travel time. Response time is the whole round trip, which includes latency plus queueing, server processing and transferring the response.
Is 0.03 ms response time good?
Over a real network, 0.03 ms (30 microseconds) is almost certainly a cache hit, a local call or an instant error rather than a real round trip. If you mean 0.03 seconds (30 ms), that is very fast for an API or a server response.
Is an average response time enough?
No. Averages hide the slow tail, so report the median with a percentile such as p95 or p99. Users experience individual requests, not your mean.
How do you test response time?
Pick the transactions that matter, measure a baseline with a few users, then ramp up load gradually while watching response time percentiles and error rate against the number of users. Repeat with the same script and conditions to compare releases.