Provider
HarnessIncident detail
Deployment Degradation – Failures in Prod1,2,3
Timeline window
to
Get alerted the next time Harness breaks
Free email alerts for the handful of vendors you cannot afford to miss. No card, live in about a minute. Paid plans add Slack, Teams, Discord, and webhook delivery across your whole stack, plus higher API quotas.
Timeline
Incident updates
Every update Harness posted, oldest to newest, exactly as it appeared on their official status page.
Investigating
Deployment Degradation – Failures in Prod1,2,3
Investigating
We are continuing to investigate this issue.
Investigating
We are continuing to investigate this issue.
Monitoring
A fix has been implemented and we are monitoring the results.
Resolved
This incident has been resolved.
Resolved
### Incident Summary
On May 6 at 11:50 PM PST, we deployed a configuration change to one of our core pipeline
services. This change introduced an unintended interaction with our database layer, causing a
significant increase in write load. The resulting pressure degraded query and command
throughput across the platform
### Root Cause
The configuration change introduced a blocking condition on expression evaluation in the
pipeline service. When executions encountered blocked expressions, they failed and retried
repeatedly, generating a write storm against the database and the throughput went 4x
### Remediation
● Rolled back the configuration change.
● Applied database-level tuning to reduce write pressure and accelerate backlog drainage
● Performed a controlled failover to a healthy database node to restore throughput
● Scaled up database nodes to provide sufficient capacity for full recovery
### Preventive Actions
To prevent from such Issues happening again, we are focussing on:
1. Increased database resilience: We are implementing automated load-shedding thresholds
that trigger on leading indicators (replication lag, session depth, op latency) before the database
reaches saturation, preventing retry storms from compounding into full degradation events.
2. We are optimizing databases so that we can increase write throughput by an order of magnitude and enable independent scaling of customer data workloads. This would have allowed us to drain the message backlog nearly instantaneously during this incident.
Keep exploring
More from Harness
Neighboring incidents on Harness's timeline and the rest of their record on OutageDeck.