A slow Redis call can tie up the PHP worker serving a visitor. I set a budget for connection time, reads and retries together, then test what the application does when the cache does not respond. A short connection timeout alone leaves other waits unbounded.
Count the whole wait
There are several separate questions: how long may connecting take, how long may a read take, and how many times should a failed operation be retried? A short limit multiplied by several retries can still use most of a page’s response budget. Add repeated cache calls inside one request and the cost grows again.
I start with the time the application can reasonably spend on optional work, then divide that budget between the operations. I measure on the actual network. Copying a tiny timeout from a local socket setup into a remote managed service is a good way to manufacture failures.
Failing open needs a working fallback
For a disposable object cache, a failed read can become a cache miss and the application can obtain the value from its original source. That is what I mean by failing open here. It does not apply to authorisation, payment records, sessions or a datastore that holds the only copy of something.
The fallback has a cost. If every request suddenly returns to the database, the database needs enough headroom to cope. I also avoid trying the same broken connection repeatedly within one request. Remembering that it failed prevents an otherwise sensible fallback turning into dozens of identical waits.
Make the failure visible
A site that still responds can hide a cache outage for days. I want a concise operational signal that the cache is unavailable, without logging credentials or a line for every visitor. I also made failed cache drop-in installation visible in the admin area: a green plugin status is no use when the file that actually provides the cache was never installed.
The check I keep
- Confirm normal reads and writes work from both web requests and scheduled jobs.
- On a test environment, make the cache unavailable and measure how long a page takes to recover through its fallback.
- Check worker usage and database load while several requests do this together.
- Restore the cache and confirm the application reconnects without a manual site restart.
I treat this as resilience work, not a claim that every site needs the same timeout. The useful result is a known upper bound, a tested fallback and a signal when the faster path is missing.