Serving an eligible cached page while a worker refreshes it can keep visitors waiting less. The refresh still consumes PHP workers, database time and remote API capacity. I give it a clear freshness limit, one pending job per page and a budget which leaves room for normal requests.
Decide what may be stale
A public information page can sometimes tolerate a short delay after an edit. A basket, account page or a response containing a private token cannot be treated the same way. I decide the eligibility and maximum stale window before adding a background worker.
There are also changes which deserve a hard purge, such as withdrawing information that must no longer be shown. Stale serving needs an explicit policy for those cases; it is not permission to keep every old page available indefinitely.
One page, one pending refresh
When hundreds of visitors request the same stale page, I want one refresh job rather than hundreds. A short-lived lock or deduplicated queue entry makes that possible. The worker removes or expires the coordination state so a failed job cannot leave the page permanently stuck.
It also matters where the queue is drained. If the early cache responds before the normal application starts, assuming that an ordinary application hook will always process the queue can leave it untouched. A real scheduler gives the work an independent route to run.
Keep capacity available for visitors
I cap concurrent renders, limit how long one run can work and slow it down when the server is busy. Load average can be one signal, but it does not describe database pressure or a remote API limit on its own. The limits need to reflect the expensive parts of the actual application.
A small pause between jobs can be more valuable than a higher worker count. Visitors still need PHP workers while the cache catches up, and the database does not care whether its sudden workload came from users or a well-intentioned warmer.
Watch the age of the queue
A queue length of ten could mean a healthy worker is nearly finished, or that ten pages have been waiting all morning. I look at the oldest pending item, failures and the age of the stored content as well as throughput.
The test I like is an edit followed by a burst of requests: readers should receive an allowed response quickly, one refresh should be scheduled, and the new version should appear within the agreed window. Then I repeat it with the worker stopped. That exposes whether the stale limit and failure reporting are doing their job.