The bug that halved a gateway's throughput without a single error

Pfizer runs an AI gateway — the single service every internal tool talks to when it wants a model. During a routine upgrade check, its throughput fell by half. No errors. No failed requests. Nothing wrong on any dashboard.
They published what they found together with the LiteLLM team, and the cause is small enough to fit in a paragraph: someone wrote a setting that said don't encrypt this connection, and the system encrypted it anyway.
What actually went wrong
The gateway keeps a cache — a fast side-store called Redis that it checks before doing expensive work. The connection to that cache can be encrypted or not, and that's a setting you write in a config file.
The operator wrote the setting to off. Plain connection, no encryption.
The code that read that setting checked whether the setting existed, rather than
what it was set to. Writing off creates the setting. The setting now exists. So
the check passed, and the gateway opened an encrypted connection — to a cache that
wasn't expecting one.
Here's the part that makes this dangerous rather than merely wrong. When a program tries to start an encrypted conversation with something that doesn't speak encryption, it doesn't get turned away. It says hello and waits for a reply that never comes. Not refused — ignored.
So the cache wasn't down. The cache was slow. And slow is much harder to see.
The four lines
For anyone who wants to look at it directly, this is
get_redis_connection_pool
as it stood in the affected version:
connection_class = async_redis.Connection
if "ssl" in redis_kwargs:
connection_class = async_redis.SSLConnection
redis_kwargs.pop("ssl", None)
redis_kwargs["connection_class"] = connection_class
"ssl" in redis_kwargs asks does this key exist. It does — you created it when you
set it to false. The branch fires, and you get the encrypted connection you
explicitly turned off.
Leave the setting out entirely and everything works. Write it down and say no, and it breaks. The careful configuration is the broken one.
Why nobody got an error
Every layer above the cache is built to tolerate a cache that isn't answering, and that's correct design — if the cache is unavailable, you do the work the slow way and carry on. Which is exactly what turns this into silence:
Redis timeout errors appeared in the proxy's internal logs, but the gateway still returned HTTP 200 to every caller
HTTP 200 means success. Every single request reported success, for the entire duration. The post is precise about where the damage landed:
Redis operations were stalling on TLS handshake timeouts in the async hot path, degrading throughput without surfacing HTTP-level errors
The reported numbers: throughput down from about 300 requests per second to about 156 — roughly half — with median response time under load at 4,200 ms, measured against 750 simulated concurrent users over a sustained 60-second run.
Those measurements come from Pfizer's internal testing and are reported in the post rather than independently reproducible. The code is a different matter — you can open the file at that version and read it yourself.
Their own one-line summary is the best sentence in the writeup:
A throughput drop with a 0% HTTP error rate is the worst kind of regression to catch after the fact
The bug was already there
I wanted to know whether this arrived with the upgrade that exposed it. It didn't. I compared the relevant file between the version that performed fine and the version that didn't — they're byte-for-byte identical. The post says the same thing:
The underlying Redis bug existed in older versions too, but surfaced during Pfizer's v1.89.2 upgrade validation under this workload
That's the detail worth sitting with. This wasn't a bad release you could roll back. The flaw had been sitting in the connection code across versions that benchmarked perfectly well. It needed a particular configuration and a particular load before it did any visible damage — and then it cost half a production gateway's capacity without raising anything.
How they actually found it
Not through monitoring, and not by hunting through recent code changes. By turning features off and on until the number moved:
Enabled Redis caching. Median latency jumped to 4,200ms. That isolated it to the Redis connection path.
A team with real production traffic, real observability and a direct line to the vendor found this by hand. There was no signal in the error rate, because there were no errors. Nothing in the status codes, because they all said success. The only place it showed up was response time under sustained load — which you see only if you go looking on purpose.
That matches what I ran into at a far smaller scale when I load tested a gateway to find the stall the median hides: in this kind of infrastructure, the failures that cost you most are the ones that never raise an error. A blocking callback and a connection stuck waiting on a handshake are completely different bugs with the same fingerprint — throughput collapses, the slowest requests get much slower, and the error rate never moves.
I should be straight about the limits here: I don't run anything at Pfizer's scale, so I can't reproduce their measurements. What I can do is read the code, and the code is public.
The fix
One line, in the current version:
if redis_kwargs.pop("ssl", None):
redis_kwargs["connection_class"] = async_redis.SSLConnection
This version reads the setting's value instead of merely noting that it exists. Off now means off.
It's tracked as "Redis ssl handling — Presence check → value check" under LIT-4307 and PR #32590, shipped in v1.93.0. The pull request title puts it plainly: honor ssl value instead of key presence when building async connection pool.
What to take from it
Two things, if you run anything with a cache behind it.
Check what your system actually did, not what you told it to do. Writing a setting and having that setting take effect are different events, and this is a clean example of the gap. The connection Pfizer got was the opposite of the one they asked for, and nothing anywhere said so.
If your testing measures averages, it cannot see this. Neither can alerting built on error rates. Sustained load, slowest-request timings, and throughput compared against a known baseline are what turn a silent stall into something visible.
The dashboards stayed green through the whole thing. That's not a gap you close by adding another alert on errors — there weren't any to alert on.
