Skip to content

Incident detail

Prod2: CV is failing with errors for all customers

Resolved incidentMinor1 affected service

Timeline window

to

Get alerted the next time Harness breaks

Free email alerts for the handful of vendors you cannot afford to miss. No card, live in about a minute. Paid plans add Slack, Teams, Discord, and webhook delivery across your whole stack, plus higher API quotas.

Timeline

Incident updates

Every update on the official status source, oldest to newest, exactly as it appeared there.

  1. Investigating

    We are currently investigating this issue.

  2. Monitoring

    A fix has been implemented and we are monitoring the results.

  3. Monitoring

    We are continuing to monitor for any further issues.

  4. Resolved

    This incident has been resolved.

  5. Resolved

    Summary

    Customers on Prod-2 experienced elevated failure rates in Continuous Verification pipelines. Verification jobs failed to complete or reported errors because CV data-collection workers could not communicate with the platform service that coordinates their work. This was caused by an unexpected traffic spike to one of our backend systems.

    Service was restored by scaling out the affected platform component. No data loss or corruption occurred.

    ‌

    Customer Impact

    Customers using Continuous Verification on Prod-2 saw verification pipeline jobs fail or time out during the incident window. A significant fraction of verification job instances that ran during this period ended in an operational failure. Pipelines not using CV continued to operate normally.

    ‌

    Root Cause

    A large automated provisioning run created a significant number of persistent CV monitoring workers in a short window, which overwhelmed a shared platform component.

    The sudden increase in concurrent outbound calls exhausted the available network connections on each service instance.

    ‌

    Mitigation

    The maximum scaling limit for the CV service was raised, allowing it to expand capacity to meet demand. Failures stopped shortly after the scale-out was completed.

    ‌

    Next Steps

    To prevent recurrence, Harness will:

    • Increase capacity headroom and add provisioning guardrails: Raise the CV service scaling limit to maintain headroom under burst loads, and introduce rate limiting for bulk provisioning of monitored services.
    • Improve Resiliency: Add exponential backoff and jitter for worker retries, and use a reusable connection pool to reduce overhead and prevent transient failures from compounding under high concurrency.
    • Improve monitoring and alerting: Enhance alerting for abnormal per-account worker growth rates, so similar conditions are detected and acted on earlier.

Keep exploring

More from Harness

Neighboring incidents on Harness's timeline and the rest of their record on OutageDeck.