Benchmarks Caching proxies
HTTP caching proxies for HLS delivery
Varnish, Vinyl, and NGINX across four HLS workloads: hit-path TTFB, segment serve, miss-storm coalescing, and origin-flap grace.
Hit-path TTFB at 5000 rps
p99 TTFB at 5000 rps on a 2 KB manifest across three caching proxies and three TLS topologies.
Test 1 p50 TTFB
50th-percentile time to first byte at 5000 rps (LOWER ms = FASTER)
Test 2 p99 TTFB
99th-percentile time to first byte at 5000 rps (LOWER ms = FASTER)
p99 TTFB ranged from 0.774 ms (Vinyl c67fd5f57 (plaintext)) to 6.002 ms (Varnish 9.0.3 (PROXYv2))
4 MB segment serve
4 MB segment delivery latency, full GET and 64 KB range, at 500 rps.
Test 1 p50 full GET
50th-percentile TTFB for 4 MB full GET at 500 rps (LOWER ms = FASTER)
Test 2 p50 range GET
50th-percentile TTFB for 64 KB range GET at 500 rps (LOWER ms = FASTER)
Test 3 p99 full GET
99th-percentile TTFB for 4 MB full GET at 500 rps (LOWER ms = FASTER)
Test 4 p99 range GET
99th-percentile TTFB for 64 KB range GET at 500 rps (LOWER ms = FASTER)
Full GET p99 TTFB ranged from 2.783 ms (Varnish 9.0.3 (plaintext)) to 64.856 ms (NGINX stable (TLS)); Range GET p99 0.96–3.75 ms
Miss-storm coalescing
200-client burst on a cold segment, 20 repetitions, plaintext only.
Test 1 p50 TTFB
50th-percentile TTFB per 200-client burst on a cold segment (LOWER ms = FASTER)
Test 2 p99 TTFB
99th-percentile TTFB per 200-client burst on a cold segment (LOWER ms = FASTER)
Test 3 Coalescing efficiency
Ratio of concurrent clients to origin requests (200x = perfect coalescing) (HIGHER x = BETTER)
All three engines coalesced 200 concurrent requests to 1 origin fetch
Origin-flap grace
Grace/stale-if-error behavior across five origin failure cycles.
Test 1 p50 TTFB
50th-percentile TTFB during origin failure cycles (LOWER ms = FASTER)
Test 2 Grace hit ratio
Percentage of requests served from stale cache during origin failure (HIGHER % = BETTER)
All three engines maintained 100% grace hit ratio across five origin failure cycles
Test conditions: rig, corpus, protocol
- Machine
- MacBook Pro
- Chip
- Apple M2 Max
- Cores
- 12 CPU cores
- Memory
- 96 GB
- OS
- macOS 26.0 (arm64)
- Runtime
- Docker (OrbStack)
- Warmups
- 3 unmeasured passes
- Measured
- 20 passes per tool
- Process
- Container per engine, shared origin
- Cache
- Warm cache after three curl requests (W1/W2) or cold by construction (W3)
- Output
- TTFB via oha 1.16.0 with --latency-correction
What did we learn?
No single engine dominated all four workloads. Vinyl c67fd5f57 (plaintext) led on hit-path, Varnish 9.0.3 (plaintext) led on segment serve, all three coalesced 200 concurrent requests to one origin fetch, and all three maintained 100% grace under origin failure.
What this does not prove
- Varnish Cache and Vinyl Cache share a maintainer team and a common ancestor commit (63806461); this benchmark measures shipped binaries and configuration only.
- Docker/OrbStack on an M2 Max is a virtualized arm64 rig, not bare-metal x86_64; numbers are relative only.
- Vinyl c67fd5f57 has no in-process TLS listener in this configuration; its TLS-topology coverage is PROXYv2-only.
- NGINX hit/miss accounting is log-derived ($upstream_cache_status), coarser than varnishstat/vinylstat's live counters.
- Purge/ban stampede is out of scope this pass (queued).
- Client, terminator, engine, and origin share one host's CPU/network stack — results reflect relative efficiency under shared contention, not isolated network performance.
- Traffic is synthetic HLS-shaped load, not a captured production trace.
Rerun it yourself: the runner and configuration are in the repository, and the methodology page covers what every run holds constant. Think a result is wrong? Open an issue with your rig and your samples.
Logos identify the products under test and imply no endorsement. Varnish is a registered trademark of Varnish Software AB. NGINX is a trademark of F5, Inc. Vinyl Cache logo: CC BY 4.0 Rhubarbe.design.