Skip to content

Incident detail

Serverless Inference - High error rates for open source models ( Qwen 3 32B)

Resolved incidentMinor

Timeline window

to

Outage alerts

Get alerted the next time DigitalOcean breaks

Free email alerts for up to 5 providers. No card, live in about a minute. Paid plans add Slack, Teams, Discord, and webhook delivery across your whole stack, plus higher API quotas.

Timeline

Incident updates

Every update DigitalOcean posted, oldest to newest, exactly as it appeared on their official status page.

  1. Investigating

    Serverless inference for alibaba-qwen3-32b (Qwen 3 32B) in tor1 is experiencing high error rates starting at 10:46 UTC.

  2. Identified

    We are currently investigating reports of elevated latency affecting requests to this model when using Serverless Inference and Agents.

    Earlier observations indicated increased error rates for the open-source Qwen 3 32B model. The Ray dashboard also showed multiple workers in a pending state, suggesting capacity constraints.

    Our analysis determined that the model was experiencing higher-than-expected request volume without sufficient resources to scale accordingly. To address this, the node pool size has been increased to improve available capacity. However, there are still insufficient nodes to fully support the desired number of model replicas.

    Following the node pool expansion, a new pod-related error has been identified. Our Engineering team is actively working to resolve this issue and restore full service performance.

  3. Resolved

    Service has been fully restored, and the model is now operating normally. We have implemented improvements to enhance stability and reduce the likelihood of similar issues in the future.

Keep exploring

More from DigitalOcean

Neighboring incidents on DigitalOcean's timeline and the rest of their record on OutageDeck.