Skip to content

Incident detail

Fly.io Upstash Redis Service Distruption

Resolved incidentCritical

Timeline window

to

Outage alerts

Get alerted the next time Upstash breaks

Free email alerts for up to 5 providers. No card, live in about a minute. Paid plans add Slack, Teams, Discord, and webhook delivery across your whole stack, plus higher API quotas.

Timeline

Incident updates

Every update Upstash posted, oldest to newest, exactly as it appeared on their official status page.

  1. Investigating

    Some regions are experiencing connectivity issues due to an ongoing network problem. We are currently investigating

  2. Identified

    The issue has been identified and the fix is being implemented.

  3. Monitoring

    A fix has been implemented and we are monitoring the results.

  4. Resolved

    This incident has been resolved.

  5. Resolved

    On May 12th and 13th at various times, a subset of Upstash Redis instances on Fly.io experienced intermittent hangs and elevated error rates. The Redis process would stall inside a logging syscall — alive but not making progress — which made the issue hard to spot from our usual telemetry. After investigating with Fly's team, we identified the root cause as a bad interaction between a recent guest kernel update on Fly's newer machines and an upstream Cloud Hypervisor bug (cloud-hypervisor#7672) affecting log writes from inside the VM. We mitigated by disabling the affected logging paths, and Fly has since rolled out a hypervisor-side patch, fully resolving the issue. No data was lost. Sorry for the disruption.

Keep exploring

More from Upstash

Neighboring incidents on Upstash's timeline and the rest of their record on OutageDeck.