Nobody logs a ticket, because search was fine last month and is only slightly worse this month. A year later it is slow enough that people search in other systems first.
The estate is much bigger than it was when the search tier was designed, and the design assumptions of year one were never revisited.
The full-text indexes had fragmented as the estate grew. Index maintenance had been assumed, not scheduled, and it fell between the application team and the infrastructure team.
Object counts were well past the design assumptions of year one. Perceived speed, the thing users actually complain about, belonged to nobody.
Schedule defragment and rebuild cycles sized to the change rate, and verify them with the same search timings that proved the problem.
Give the maintenance an owner and a place in the runbook. Then it stops being a surprise.
Month by monththe rate at which search degrades: too slowly for a ticket, too surely to ignore
What the logs said
search timings p50 rising month on month system report full-text index fragmentation high; no defrag or rebuild schedule present verdict failure mode THREE: maintenance assumed, not scheduled; no owner fix scheduled defrag sized to the change rate, owner named, retest with the same timings
