The Elasticsearch deadlock that parked every index: ILM warm phases on a single node
, 1 min read
This cluster is a single Elasticsearch 9.1.5 node, and it holds compliance audit evidence for a US fintech. A stuck lifecycle there is not cosmetic. Twice it stopped moving indices through their lifecycle, and I treated those as two separate incidents. They weren't.
What broke
First, the cluster climbed to its shard ceiling — 999 of 1,000 — and indices stopped moving through their lifecycle. I spotted it from the host's disk space.
A month later it came back in a different form: a warm-phase deadlock, with indices waiting in check-migration indefinitely.
What I tried
I treated the shard-ceiling incident and the warm-phase deadlock as two separate problems. They were the same root cause in two different forms.
What worked
On a single node, any ILM policy with a warm or cold phase must set replicas to zero in that phase. Otherwise check-migration waits for replica copies that can never be allocated.
In the policy, that's one action in the phase:
"warm": {
"actions": {
"allocate": { "number_of_replicas": 0 }
}
}Fixing the policy doesn't move indices that are already stuck. I moved those to their next phase manually.
After the fix: shards 810 → 433, ILM errors at zero, disk steady at 71%. I verified the warm-phase deadlock closed on 12 August 2026.
Worth knowing
- YELLOW with around 40 unassigned shards is now the expected steady state. Fleet creates new indices with one replica, and the warm phase strips that replica at its first transition.
- The first incident was caught from disk space, not from Elasticsearch itself. Afterwards I set up an alert rule so it can't go unnoticed again.