Skip to content

Incident detail

Elevated error rate in AWS

Resolved incidentMajor

Timeline window

to

Outage alerts

Get alerted the next time ClickHouse Cloud breaks

Free email alerts for up to 5 providers — no card, live in about a minute. Paid plans add Slack, Discord, and webhook delivery across your whole stack, plus higher API quotas.

Timeline

Incident updates

Updates are normalized from the official source chronology so timeline changes remain easy to scan.

  1. Investigating

    We are investigating a partial outage affecting services in the AWS region us-east-1.

  2. Identified

    Our engineering team has identified the root cause of the service disruption affecting ClickHouse some instances in US-EAST-1. We are currently deploying a fix across the region. ClickHouse instances should begin resuming normal operations within the next 10-30 minutes. We will continue to monitor closely and provide updates as services are restored.

  3. Resolved

    This issue is now fixed.

  4. Investigating

    We are investigating an elevated error rate in AWS us-east-1 and eu-west-2.

  5. Resolved

    All components are operational after the restart of an internal component (Cilium) in the affected regions. The team is working on an RCA.

  6. Investigating

    We are investigating an elevated error rate in AWS eu-central-1 and ap-southeast-1.

  7. Monitoring

    The team mitigated the issue in affected regions and continues to investigate the root cause.

  8. Monitoring

    Additional configuration changes have been applied to prevent the recurrence of this issue. The team continues to monitor the systems.

  9. Monitoring

    We believe this is related to a known bug in the version of our networking stack currently in use and are mitigating the issue by resetting this layer across all AWS regions. The team continues to monitor all systems closely.

  10. Monitoring

    We have completed the reset of the networking stack across all regions and continue to monitor all systems closely.

  11. Resolved

    We have confirmed that all systems are stable and operating normally. The networking configuration has been fully restored across all regions, and we have implemented additional monitoring and safeguards to prevent recurrence.

    We will continue to monitor systems closely over the coming days as part of our standard post-incident procedures, and share an RCA as soon as possible.