Nginx Response Buffering - Decouple Slow Clients Without Delaying Streams
An application can finish generating a response quickly while the person downloading it is on a slow connection. Who should wait in that situation: the application, or the reverse proxy in front of it? The answer changes when the response is not an ordinary HTML page or JSON document but a stream whose value depends on each message arriving promptly.
Nginx response buffering sits inside that timing problem. For an HTTP upstream reached with proxy_pass, buffering is enabled by default. That default is useful for many conventional responses, but it is not neutral for Server-Sent Events, incremental progress, or other timing-sensitive output. The goal is therefore not to turn buffering off everywhere. It is to decide what each endpoint promises, then preserve that promise deliberately.
There are three clocks, not one
A reverse-proxied response involves at least three participants: the client, Nginx, and the upstream application. They can move at different speeds. The application may produce bytes quickly, Nginx may read them quickly, and the client may consume them slowly. Alternatively, the application may produce a small message every few seconds while the client is ready to receive each one immediately.
According to the official Nginx proxy module reference, when response buffering is enabled, Nginx reads from the upstream as soon as possible into configured memory buffers. If the response does not fit, part of it can be written to a temporary file. Nginx can then continue sending the response at the client's pace.
This separates two relationships. The upstream talks to Nginx; the client also talks to Nginx. A slow client does not necessarily have to keep the upstream occupied for the whole download. The official reverse proxy guide identifies that isolation from slow clients as a reason buffering can help.
With buffering disabled, the relationship becomes tighter. Nginx passes response data to the client synchronously as it receives that data from upstream. That is valuable when delivery timing is part of the endpoint's behavior, but it also means the upstream side is less insulated from a client that reads slowly.
Buffering is not caching
The word “buffer” is easy to confuse with several neighboring features. Response buffering is temporary handling for one response in flight. It does not, by itself, make that response reusable for a later request. Reuse across requests belongs to response caching and its own directives, keys, freshness rules, and safety decisions.
Request buffering is different again. proxy_request_buffering controls whether Nginx reads a client request body before sending it upstream. That affects uploads and request retry behavior. proxy_buffering controls the response traveling in the opposite direction. Changing one does not implicitly change the other.
This distinction matters during diagnosis. “The upload reaches my application late” and “the browser receives streamed events in a burst” may both be described casually as buffering, but they point to different data paths and different directives.
What the response buffers actually control
For proxied HTTP responses, proxy_buffer_size sets the buffer used for the first part of the upstream response, which usually contains the response headers. proxy_buffers sets the number and size of buffers available for one connection. Their documented defaults depend on the platform's memory page size, so a copied number is not automatically an improvement.
If a buffered response exceeds the available memory buffers, Nginx may use a temporary file. proxy_max_temp_file_size limits that temporary file for a response, while a value of zero disables response buffering to temporary files. That option does not make the response disappear, nor does it prove memory use will be harmless under concurrency. It changes where excess buffered data may go.
There is no credible universal answer such as “use sixteen buffers” without the response-size distribution, concurrency, memory budget, storage behavior, and client population of the actual service. More memory buffers can reduce temporary-file use while increasing the potential memory held across simultaneous requests. Smaller buffers can move pressure toward storage or expose oversized response-header problems. The useful question is not which number looks fast, but which constraint is currently visible.
Streaming makes timing part of correctness
An ordinary page remains useful if it arrives in a few larger pieces instead of many tiny ones. A progress stream is different. If an application emits “10%”, “20%”, and “30%” over time but an intermediary accumulates those messages, the bytes can still be correct while the behavior is wrong for the user.
Server-Sent Events provide a concrete example. The WHATWG HTML Living Standard defines event streams with the text/event-stream media type and parses them line by line. Its processing notes warn that block buffering can delay event dispatch. That does not mean every endpoint called “stream” needs identical Nginx settings, but it does establish that buffering can affect observable SSE behavior.
A focused location can disable proxy response buffering without changing ordinary routes:
location /events/ {
proxy_pass http://app_backend;
proxy_buffering off;
}
This example addresses one mechanism only. The upstream still needs to emit complete event records and flush them through its own runtime. Compression, another proxy, a CDN, the client library, and network conditions can also affect delivery. Turning off Nginx buffering cannot compensate for an application that keeps its output in a language-level buffer.
Let the application identify exceptional responses
Sometimes one route returns both conventional responses and a stream. The proxy module also recognizes the upstream response header X-Accel-Buffering. An application can return X-Accel-Buffering: no for the response that needs pass-through behavior while leaving buffering enabled for ordinary responses. Nginx processes this header unless its handling has been disabled with proxy_ignore_headers.
This header is an Nginx control mechanism, not a general HTTP promise that every intermediary understands. If another reverse proxy or CDN sits in front of Nginx, its buffering behavior must be checked separately. The application-controlled approach is useful when the application knows the response semantics better than a broad URL pattern, but it also creates a cross-layer contract that deserves documentation and a test.
What is lost when buffering is disabled
Lower delivery delay is not a free setting. Without response buffering, upstream production and client consumption are more directly coupled. A slow or paused client can keep the response path active longer. Capacity consequences depend on the upstream server, its concurrency model, connection limits, response rate, and the number of simultaneous streams; they cannot be inferred from one directive alone.
Error recovery also has a hard boundary. The proxy_next_upstream documentation notes that Nginx can pass a request to another upstream only if nothing has yet been sent to the client. Once a partial response is visible, a second server cannot quietly replace its beginning. This limitation matters for streaming, where output intentionally starts before the operation is complete.
Disabling buffering also does not define liveness. proxy_read_timeout measures the gap between successive reads from upstream, not the duration of the entire response. A long-lived stream therefore needs an application-level decision about idle periods, heartbeats, reconnection, and what a missing message means. Raising a timeout without that model can merely postpone failure detection.
Observe both sides of Nginx
A browser stopwatch alone cannot explain where time was spent. Nginx exposes client-facing and upstream-facing timing variables that can make the boundary more visible. The log module reference defines $request_time as the time from reading the first client bytes until logging after the last response bytes are sent. The upstream module separately provides $upstream_header_time and $upstream_response_time.
A temporary or carefully retained access-log format might include those boundaries:
log_format upstream_timing
'$request status=$status bytes=$body_bytes_sent '
'request_time=$request_time '
'upstream_header_time=$upstream_header_time '
'upstream_response_time=$upstream_response_time';
These values need interpretation. A large $request_time beside a smaller upstream time can be consistent with slow client delivery, but it is not proof that one buffer setting is wrong. Retries can produce multiple upstream values, streaming responses make “completion” intentionally late, and aggregate traffic matters more than one request. Compare like endpoints under representative conditions and inspect Nginx error logs for temporary-file warnings rather than tuning from an isolated line.
Change one location, then verify the promise
A cautious process begins with behavior rather than directives:
- Classify the endpoint as a bounded response, a large transfer, or a timing-sensitive stream.
- Record the symptom: delayed first event, upstream held by slow clients, temporary-file warnings, memory pressure, or something else.
- Measure client-visible timing and upstream timing before changing configuration.
- Apply the smallest location-specific change that tests the hypothesis.
- Validate both normal clients and deliberately slow clients, then watch concurrency and resource use.
Before reloading, use the documented configuration test:
sudo nginx -t
The Nginx command-line reference says -t checks configuration syntax and attempts to open referenced files. Only after it succeeds should the tested configuration be reloaded through the server's normal service-management procedure. A successful syntax test does not prove that events arrive on time, that memory use is acceptable, or that an upstream flushes correctly. Those are runtime properties and need runtime checks.
Keep the default ordinary, make exceptions explicit
Response buffering is not an old optimization that should automatically be removed, and streaming is not a reason to disable it for an entire site. Buffering can let an upstream finish while Nginx handles a slower client. Pass-through behavior can preserve timing when incremental delivery is part of the endpoint's contract. Each solves a different problem.
A defensible starting point is modest: keep buffering for ordinary bounded responses, make streaming exceptions narrow and visible, and avoid buffer-size folklore. The final decision belongs to evidence from the actual response path. If the application, proxy, and client each keep a different clock, configuration should begin by asking which clock the user is waiting for.
