no healthy upstream is the kind of error that makes you expect wreckage.
Then you open the dashboards and find… almost nothing. CPU is low. No pods have crashed. The last deployment was hours ago. By the time you refresh the page, the service has recovered by itself.
That was the scene a few weeks ago when I started chasing an intermittent failure in a search ba...