Skip to content

Incident detail

All modules are running slow in Prod1/2/3/4 due to cloud provider incident

Resolved incidentMajor1 affected service

Timeline window

to

Get alerted the next time Harness breaks

Free email alerts for the handful of vendors you cannot afford to miss. No card, live in about a minute. Paid plans add Slack, Teams, Discord, and webhook delivery across your whole stack, plus higher API quotas.

Timeline

Incident updates

Every update Harness posted, oldest to newest, exactly as it appeared on their official status page.

  1. Investigating

    We are currently investigating this issue.

  2. Investigating

    The slowness could cause either of below symptoms:

    • Pipelines not starting
    • Delays in execution
    • Pipelines being cancelled due to timeouts
  3. Investigating

    Our cloud provider is facing an active incident and we are following up.

  4. Investigating

    We are continuing to investigate this issue.

  5. Identified

    Our cloud provider has confirmed an ongoing incident impacting multiple regions. Harness pipelines have not experienced failures as a result, though some users may continue to experience slowness. We are monitoring the situation closely and will provide updates as more information becomes available.

  6. Monitoring

    We are observing improved latencies across the board following the fix implemented by our cloud provider. We are continuing to monitor the situation closely and will provide further updates as warranted.

    We noted some stuck executions for CI for couple of customers, which we are investigating

  7. Resolved

    This incident has been resolved.

  8. Resolved

    # Summary

    On 20 August 2026, beginning at approximately 15:00 UTC, the Harness platform experienced widespread performance degradation across all production environments. Pipeline executions that normally complete in around two minutes took seven to ten minutes. Continuous Delivery, Continuous Integration, pipeline orchestration, and Feature Management & Experimentation were all affected.

    Google Cloud Platform experienced a multi-product incident in the us-west1 region affecting Bigtable, Compute Engine, Google Kubernetes Engine, and persistent-disk I/O. Harness production infrastructure runs on persistent disks in that region. The degradation raised database operation latency from approximately 2 ms to over 10 ms at the 95th percentile, which in turn caused message-queue processing lag and propagated to every service that depends on timely database access.

    # Impact

    This was a degradation, not an outage. Pipelines continued to execute and complete successfully throughout; they were slow rather than failing. No data was lost, and no customer work was dropped as a result of this incident.

    # Root cause

    Harness production infrastructure in the affected environments runs on Google Cloud Platform persistent disks in the us-west1 region. When that storage layer degraded, the effect propagated through the platform in a predictable chain:

    Persistent-disk I/O degradation in us-west1. Google Cloud Platform experienced a multi-product incident affecting Bigtable, Compute Engine, Google Kubernetes Engine, and persistent-disk performance. This was an infrastructure failure in the provider’s environment, outside Harness’s control.

    # Preventive actions

    Although Harness cannot prevent a cloud provider infrastructure failure. The actions below are aimed at detecting one faster and being better positioned to act on it.

    | Action |

    | --- |

    | Continue routine pre-testing of targeted cross-region database failovers, as performed during this incident, to keep failover readiness verified rather than assumed |

    | Assess full-stack multi-region failover readiness for future scenarios in which cross-region latency would be unacceptable |

Keep exploring

More from Harness

Neighboring incidents on Harness's timeline and the rest of their record on OutageDeck.