Skip to content

Incident detail

Delayed scaling due to node failure

Resolved incidentMinor

Timeline window

to

Outage alerts

Get alerted the next time Replicate breaks

Free email alerts for up to 5 providers. No card, live in about a minute. Paid plans add Slack, Teams, Discord, and webhook delivery across your whole stack, plus higher API quotas.

Timeline

Incident updates

Every update Replicate posted, oldest to newest, exactly as it appeared on their official status page.

  1. Monitoring

    Scaling decisions were delayed by nearly 1 hour after the controller responsible for emitting queue metrics failed to schedule on a soft-failed node. We have since cordoned and drained the node, and the controller is emitting queue metrics again.

  2. Resolved

    All affected models have been scaling correctly for more than 30 minutes at this point, and we see no residual prediction queues. Thank you for your patience!

Keep exploring

More from Replicate

Neighboring incidents on Replicate's timeline and the rest of their record on OutageDeck.