Skip to content

Incident detail

Hitting GPU Capacity for H100s creating large queue times for some models

Active incidentMinor

Timeline window

Started

Outage alerts

Get alerted the next time Replicate breaks

Free email alerts for up to 5 providers — no card, live in about a minute. Paid plans add Slack, Discord, and webhook delivery across your whole stack, plus higher API quotas.

Timeline

Incident updates

Updates are normalized from the official source chronology so timeline changes remain easy to scan.

  1. Identified

    We are over provisioned on H100s currently which is causing long queue times for some models running on H100s

  2. Monitoring

    GPU usage is back below capacity. We're continuing to monitor.