Skip to content

Incident detail

SEI 2.0 dashboards are not loading

Resolved incidentMajor

Timeline window

to

Get alerted the next time Harness breaks

Free email alerts for the handful of vendors you cannot afford to miss. No card, live in about a minute. Paid plans add Slack, Teams, Discord, and webhook delivery across your whole stack, plus higher API quotas.

Timeline

Incident updates

Every update on the official status source, oldest to newest, exactly as it appeared there.

  1. Investigating

    We are currently investigating a Harness component that is experiencing issues. We are working to identify the cause and restore normal operations as soon as possible.

  2. Investigating

    We are continuing to investigate this issue.

  3. Identified

    The issue has been identified and a fix is being implemented.

  4. Resolved

    This incident has been resolved.

  5. Resolved

    ## Summary

    Customers on Prod1, Prod2, and Prod3 (US) clusters experienced failures when loading SEI 2.0 dashboards on August 6, 2026, from 7:22 AM PDT to 9:03 AM PDT. Customers calling the SEI 2.0 API also experienced similar failures.

    No customer data was lost, and ingestion of all integration data continued to work uninterrupted. SEI customers using 1.0 were not impacted.

    ## Root Cause

    The incident was caused by resource exhaustion on the nodes serving queries. This resource degradation developed in a pattern that did not cross our existing alerting thresholds early enough to provide sufficient warning or allow mitigation before customer impact occurred.

    ## Impact

    Customers on Prod1, Prod2, and Prod3 (US) clusters were unable to load SEI 2.0 dashboards during the incident window.

    Duration: August 6, 2026, 7:22 AM PDT – 9:03 AM PDT (~1 hour 41 minutes)

    ### What was not impacted?

    • Data ingestion and processing
    • SEI 1.0 customers
    • Integrations and metadata flows

    No customer data was lost.

    ## Remediation

    Upon identifying the root cause, our team took immediate corrective action by adding capacity to restore the affected systems. Services were fully recovered, and all dashboards resumed normal operation at 9:03 AM PDT.

    ## Action Items

    To prevent from such issues happening again, Harness is/has

    Proactively added capacity updates have been applied to prevent this issue from recurring

    #### Enhanced Monitoring and Alerting

    Additional monitoring and alerting have been put in place to detect anomalies early, focused on a leading indicator, which in this case was thread pool exhaustion, before they can impact dashboard availability and data rendering.

    #### System Patch in Progress

    We are working with our vendor to apply a patch to remediate this and similar issues completely.

Keep exploring

More from Harness

Neighboring incidents on Harness's timeline and the rest of their record on OutageDeck.