Provider
ClickHouse CloudIncident detail
We are investigating an issue with CH servers in Azure germanywestcentral region
Timeline window
to
Get alerted the next time ClickHouse Cloud breaks
Free email alerts for the handful of vendors you cannot afford to miss. No card, live in about a minute. Paid plans add Slack, Teams, Discord, and webhook delivery across your whole stack, plus higher API quotas.
Timeline
Incident updates
Every update on the official status source, oldest to newest, exactly as it appeared there.
Investigating
We detected an issue with CHC clusters in some Azure regions. ClickHouse cloud team is investigating. Some operations might be degraded.
Investigating
Only germanywestcentral Azure region is affected.
Investigating
Only one Availability zone is affected in germanywestcentral Azure region. The problem affects internal component (ClickHouse Keepers). Some ClickHouse Services in this region may see increased latency on inserts or DDL queries.
Investigating
The problem is caused by Azure outage in the germanywestcentral region. ClickHouse Cloud team is working with Azure support.
Impact:
• New ClickHouse services can't be provisioned (appear as stuck in UI).
• Some services may experience higher latency due to reduced capacity.
Identified
The Azure team is working on resolution of the issue. The ClickHouse Cloud team will relay any updates related to ClickHouse clusters operation.
An excerpt from Azure incident status:
> As per our investigation, we believe this stemmed from multiple underlying storage systems becoming unavailable at the same time, which led to attached virtual machines becoming unresponsive, restarting unexpectedly, or being temporarily unavailable.
Monitoring
The issue is still present on Azure infrastructure side and ClickHouse Cloud team is actively monitoring for any impact for existing ClickHouse Clusters.
Correction on earlier impact description: Provisioning of new instances is expected to work. Operations like autoscaling of the ClickHouse Keeper replicas can be stuck or slow. Latency for insert and DDL queries are not affected, thanks to cross-AZ redundancy.
Monitoring
Azure support shared they are working on fixing the underlying issue, we started seeing some recovery but the incident is not completely resolved yet. We are working with Azure to get this fixed as soon as possible.
Resolved
The underlying AZ has recovered and we are fully back online. Since ClickHouse cloud architecture uses 3 AZs we were operational during this outage but there might have been some degradations.
Azure support shared the details of the issue as following- Between 12:04 UTC and 17:46 UTC on 22 September 2022, a platform issue resulted in an impact to the Virtual Machines (VMs) and Virtual Machine Scale Sets (VMSS) in Germany West Central. Customers may have experienced their virtual machines becoming unresponsive, restarting unexpectedly, or being temporarily unavailable during the impact window.
Keep exploring
More from ClickHouse Cloud
Neighboring incidents on ClickHouse Cloud's timeline and the rest of their record on OutageDeck.