Provider
ReplicateIncident detail
Hitting GPU Capacity for H100s creating large queue times for some models
Active incidentMinor
Timeline window
Started
Outage alerts
Get alerted the next time Replicate breaks
Free email alerts for up to 5 providers — no card, live in about a minute. Paid plans add Slack, Discord, and webhook delivery across your whole stack, plus higher API quotas.
Timeline
Incident updates
Updates are normalized from the official source chronology so timeline changes remain easy to scan.
Identified
We are over provisioned on H100s currently which is causing long queue times for some models running on H100s
Monitoring
GPU usage is back below capacity. We're continuing to monitor.