Provider
HarnessIncident detail
Prod2: CV is failing with errors for all customers
Timeline window
to
Get alerted the next time Harness breaks
Free email alerts for the handful of vendors you cannot afford to miss. No card, live in about a minute. Paid plans add Slack, Teams, Discord, and webhook delivery across your whole stack, plus higher API quotas.
Timeline
Incident updates
Every update on the official status source, oldest to newest, exactly as it appeared there.
Investigating
We are currently investigating this issue.
Monitoring
A fix has been implemented and we are monitoring the results.
Monitoring
We are continuing to monitor for any further issues.
Resolved
This incident has been resolved.
Resolved
Summary
Customers on Prod-2 experienced elevated failure rates in Continuous Verification pipelines. Verification jobs failed to complete or reported errors because CV data-collection workers could not communicate with the platform service that coordinates their work. This was caused by an unexpected traffic spike to one of our backend systems.
Service was restored by scaling out the affected platform component. No data loss or corruption occurred.
Customer Impact
Customers using Continuous Verification on Prod-2 saw verification pipeline jobs fail or time out during the incident window. A significant fraction of verification job instances that ran during this period ended in an operational failure. Pipelines not using CV continued to operate normally.
Root Cause
A large automated provisioning run created a significant number of persistent CV monitoring workers in a short window, which overwhelmed a shared platform component.
The sudden increase in concurrent outbound calls exhausted the available network connections on each service instance.
Mitigation
The maximum scaling limit for the CV service was raised, allowing it to expand capacity to meet demand. Failures stopped shortly after the scale-out was completed.
Next Steps
To prevent recurrence, Harness will:
- Increase capacity headroom and add provisioning guardrails: Raise the CV service scaling limit to maintain headroom under burst loads, and introduce rate limiting for bulk provisioning of monitored services.
- Improve Resiliency: Add exponential backoff and jitter for worker retries, and use a reusable connection pool to reduce overhead and prevent transient failures from compounding under high concurrency.
- Improve monitoring and alerting: Enhance alerting for abnormal per-account worker growth rates, so similar conditions are detected and acted on earlier.
Keep exploring
More from Harness
Neighboring incidents on Harness's timeline and the rest of their record on OutageDeck.