Provider
HarnessIncident detail
Platform is experiencing degraded performance for some organizations.
Timeline window
to
Outage alerts
Get alerted the next time Harness breaks
Free email alerts for up to 5 providers — no card, live in about a minute. Paid plans add Slack, Discord, and webhook delivery across your whole stack, plus higher API quotas.
Timeline
Incident updates
Updates are normalized from the official source chronology so timeline changes remain easy to scan.
Investigating
We are currently investigating this issue.
Identified
Issue has been identified and mitigated
Resolved
This incident has been resolved.
Resolved
# Summary
On April 30, 2026, between approximately 15:29 UTC and 17:00 UTC, customers in Prod3 experienced degradation impacting delegate connectivity, instance synchronization, pipeline executions, and connector operations due to spike in load on one of our services. .
Service stability was restored through service scaling, infrastructure capacity increases, and database resource expansion.
# Impact
Customer Impact:
- Delegates disconnected intermittently during the incident window
- Instance synchronization operations were delayed
- Some pipeline executions and connector operations experienced failures or delays
Duration:
- Delegate connectivity impact: ~15 minutes
- Elevated service degradation: ~90 minutes
# Root Cause
The incident was caused by spike causing thread exhaustion and elevated request contention between internal services during a period of increased synchronization and delegate activity.
# Mitigation and Recovery
The following actions were taken to restore service stability:
- Scaled management service replicas horizontally
- Increased autoscaling thresholds and maximum replica counts
- Expanded Database compute capacity
- Upgraded MongoDB infrastructure components
- Stabilized delegate reassignment and reconnection processing
Services recovered progressively beginning at approximately 15:47 UTC, with full stability restored by ~17:00 UTC.
# Preventive Actions
To prevent such issues from happening again, We are implementing the following improvements
- Improving our circuit breakers and fail-fast protections between dependent services
- Enhancing monitoring and alerting for thread pool saturation and queue buildup
- Increasing baseline service headroom and resiliency protections