Provider
AstronomerIncident detail
Tasks remain in the queued state on some worker queues.
Timeline window
to
Get alerted the next time Astronomer breaks
Free email alerts for the handful of vendors you cannot afford to miss. No card, live in about a minute. Paid plans add Slack, Teams, Discord, and webhook delivery across your whole stack, plus higher API quotas.
Timeline
Incident updates
Every update Astronomer posted, oldest to newest, exactly as it appeared on their official status page.
Identified
We have identified the cause.
On Astro, each worker queue is backed by a Kubernetes autoscaling resource whose name is built from a fixed platform prefix, the deployment release name, and the worker queue name.
For a very small number of customers, where that combined name exceeds the Kubernetes 63-character limit, the autoscaling resource is rejected at creation. As a result, workers do not scale for the affected queue and tasks routed there remain in the queued state.
This is why the limit can be reached even when the worker queue name itself looks short: the release name and platform prefix already consume a large portion of the 63 characters before the queue name is appended.
Workaround: If affected, until the permanent fix is deployed, please keep worker queue name at or below 10 characters.
We are preparing a permanent fix. Further updates to follow.
Monitoring
A fix has been implemented, and we are monitoring the results.
Resolved
The incident has been resolved.
Keep exploring
More from Astronomer
Neighboring incidents on Astronomer's timeline and the rest of their record on OutageDeck.