Skip to content

Incident detail

Platform access issues in Prod1/Prod2/Prod3

Resolved incidentMajor1 affected service

Timeline window

to

Outage alerts

Get alerted the next time Harness breaks

Free email alerts for up to 5 providers — no card, live in about a minute. Paid plans add Slack, Discord, and webhook delivery across your whole stack, plus higher API quotas.

Timeline

Incident updates

Updates are normalized from the official source chronology so timeline changes remain easy to scan.

  1. Investigating

    We are currently investigating this issue.

  2. Investigating

    We are continuing to investigate this issue.

  3. Identified

    The issue has been identified and a fix is being implemented.

  4. Monitoring

    A fix has been implemented and we are monitoring the results.

  5. Resolved

    This incident has been resolved.

  6. Resolved

    ## Summary

    Between 10:07 AM–10:39 AM PST on Tuesday, May 12, 2026, customers using the prod1, prod2, and prod3 Production clusters experienced elevated latency and intermittent service degradation. During this timeframe, customers observed delegate timeouts, login failures, and pipeline execution failures

    ## Root Cause

    A recently introduced configuration change to a common infrastructure component caused unexpected resource pressure across nodes in the prod1, prod2, and prod3 production clusters. The peak traffic exacerbated the resource utilization and introduced elevated latency across several critical platform services.

    ## Impact

    1. Customers in prod1, prod2, and prod3 experienced login and access failures for approximately 20 minutes.

    2. Delegate connectivity was intermittently impacted during the incident window.

    3. Pipeline executions and API requests experienced elevated failure rates and latency during the incident window.

    ## Remediation

    • Immediately Rolled back to the previous stable release, restoring customer pipeline functionality and alleviating node pressure.

    ## Action Items

    1. Enhance perf testing to include such workloads so that we can catch issues before we hit production.

    2. Increase capacity across clusters to make sure we have enough headroom to absorb the traffic surges.