Skip to content

Incident detail

Tasks remain in the queued state on some worker queues.

Resolved incidentMajor

Timeline window

to

Get alerted the next time Astronomer breaks

Free email alerts for the handful of vendors you cannot afford to miss. No card, live in about a minute. Paid plans add Slack, Teams, Discord, and webhook delivery across your whole stack, plus higher API quotas.

Timeline

Incident updates

Every update Astronomer posted, oldest to newest, exactly as it appeared on their official status page.

  1. Identified

    We have identified the cause.

    On Astro, each worker queue is backed by a Kubernetes autoscaling resource whose name is built from a fixed platform prefix, the deployment release name, and the worker queue name.

    For a very small number of customers, where that combined name exceeds the Kubernetes 63-character limit, the autoscaling resource is rejected at creation. As a result, workers do not scale for the affected queue and tasks routed there remain in the queued state.

    This is why the limit can be reached even when the worker queue name itself looks short: the release name and platform prefix already consume a large portion of the 63 characters before the queue name is appended.

    Workaround: If affected, until the permanent fix is deployed, please keep worker queue name at or below 10 characters.

    We are preparing a permanent fix. Further updates to follow.

  2. Monitoring

    A fix has been implemented, and we are monitoring the results.

  3. Resolved

    The incident has been resolved.

Keep exploring

More from Astronomer

Neighboring incidents on Astronomer's timeline and the rest of their record on OutageDeck.