Provider
HarnessIncident detail
Intermittent 503s on IDP services
Timeline window
to
Outage alerts
Get alerted the next time Harness breaks
Free email alerts for up to 5 providers — no card, live in about a minute. Paid plans add Slack, Discord, and webhook delivery across your whole stack, plus higher API quotas.
Timeline
Incident updates
Updates are normalized from the official source chronology so timeline changes remain easy to scan.
Investigating
We are currently investigating this issue.
Monitoring
A fix has been implemented and we are monitoring the results.
Resolved
This incident has been resolved.
Resolved
## Summary
_Between 13:30 UTC on June 5, 2026, and 00:27:00 UTC on June 6, 2026, customers using Harness IDP experienced intermittent 503 errors when accessing IDP web pages and APIs. The incident was mitigated through a configuration update and normal operation was restored by 00:27 UTC._
## Root Cause
_The Root Cause was traced back to a performance optimization update deployed around 10:35 UTC on June 5, 2026. This network transport configuration change altered idle TCP connection behavior, causing long-lived connections to be terminated prematurely under certain conditions. Consequently, requests failed to reach backend services, leading to intermittent 503 errors in the Harness IDP service. Stability was restored by recalibrating connection timeouts and lifecycle settings to synchronize connection management across the network path and eliminate stale connection reuse._
## Impact
_Starting at approximately ~13:30 UTC, customers experienced intermittent 503 service errors while using the Harness Internal Developer Portal (IDP) until we resolved the issue at 00:27:00 UTC. The issue impacted both UI and API traffic, resulting in failed requests and a degradation of service availability until mitigation measures were implemented._
## Remediation
_The team implemented a configuration change to connection management settings, reducing idle connection timeouts and limiting connection lifetime._
_This prevented the reuse of stale connections and restored service stability._
## Action Items
1. _Enhance monitoring & alerting – Enhance synthetic alerting and coverage for upstream connection failures, connection resets, and elevated 503 error rates to enable earlier detection of network transport issues._
2. _Pre-production validation of networking changes – Expand pre-production checks for configuration alignment and testing for network timeout and connection management scenarios to reduce the risk of similar issues in the future._