Skip to content

Incident detail

Login to Traceable Clusters Impacted

Resolved incidentMinor

Timeline window

to

Outage alerts

Get alerted the next time Harness breaks

Free email alerts for up to 5 providers — no card, live in about a minute. Paid plans add Slack, Discord, and webhook delivery across your whole stack, plus higher API quotas.

Timeline

Incident updates

Updates are normalized from the official source chronology so timeline changes remain easy to scan.

  1. Investigating

    We are currently experiencing intermittent login issues with the Traceable cluster due to an issue with our login provider. We are actively working with the provider to resolve the issue.

    This incident does not impact data ingestion, and all ingestion pipelines continue to function normally.

    We will provide updates as we have more information.

  2. Identified

    The issue has been identified and a fix is being implemented.

  3. Monitoring

    Issue has been fixed and we are closely monitoring

  4. Monitoring

    We are continuing to monitor for any further issues.

  5. Resolved

    This incident has been resolved.

  6. Resolved

    ## Summary

    _Between 06:50 UTC and 07:20 UTC on June 26, 2026, customers experienced intermittent login failures while accessing Traceable environments across multiple US clusters. The incident was traced to an issue with the external authentication provider (Auth0), where elevated socket timeouts caused increased login latency and authentication request failures. Login functionality gradually recovered as the upstream issue stabilized, and the incident was resolved after successful login validation across multiple impacted clusters._

    ## Root Cause

    _The root cause was an issue with the external authentication provider (Auth0), which experienced elevated socket timeouts while processing authentication requests. These upstream timeouts increased login latency and caused intermittent authentication failures across multiple US-region clusters. Since authentication requests depended on the external provider, affected login attempts failed despite Traceable platform services remaining healthy. The incident was resolved once the upstream authentication service recovered and login requests consistently completed successfully._

    ## Impact

    _Starting at approximately 06:50 UTC, customers experienced intermittent login failures when accessing Traceable environments across multiple US clusters. The issue affected user authentication, preventing some users from accessing the platform while underlying application services remained operational. Login functionality progressively recovered during the incident, and normal authentication was restored by 07:20 UTC._

    ## Remediation

    _The engineering team worked with the external authentication provider while continuously monitoring authentication health across affected clusters. Login functionality was validated through platform metrics and manual verification across representative environments. After confirming consistent authentication success across impacted clusters, the incident was declared resolved._

    ## Action Items

    _To prevent such issues going forward, Harness will,_

    _Increase authentication resilience: Evaluate improvements to authentication request handling, including timeout tuning, retry strategies where appropriate, and graceful degradation for transient upstream failures._