Skip to content

Incident detail

Platform is experiencing degraded performance for some organizations.

Resolved incidentMinor3 affected services

Timeline window

to

Outage alerts

Get alerted the next time Harness breaks

Free email alerts for up to 5 providers — no card, live in about a minute. Paid plans add Slack, Discord, and webhook delivery across your whole stack, plus higher API quotas.

Timeline

Incident updates

Updates are normalized from the official source chronology so timeline changes remain easy to scan.

  1. Investigating

    We are currently investigating this issue.

  2. Identified

    Issue has been identified and mitigated

  3. Resolved

    This incident has been resolved.

  4. Resolved

    # Summary

    On April 30, 2026, between approximately 15:29 UTC and 17:00 UTC, customers in Prod3 experienced degradation impacting delegate connectivity, instance synchronization, pipeline executions, and connector operations due to spike in load on one of our services. .

    Service stability was restored through service scaling, infrastructure capacity increases, and database resource expansion.

    # Impact

    Customer Impact:

    • Delegates disconnected intermittently during the incident window
    • Instance synchronization operations were delayed
    • Some pipeline executions and connector operations experienced failures or delays

    Duration:

    • Delegate connectivity impact: ~15 minutes
    • Elevated service degradation: ~90 minutes

    # Root Cause

    The incident was caused by spike causing thread exhaustion and elevated request contention between internal services during a period of increased synchronization and delegate activity.

    # Mitigation and Recovery

    The following actions were taken to restore service stability:

    • Scaled management service replicas horizontally
    • Increased autoscaling thresholds and maximum replica counts
    • Expanded Database compute capacity
    • Upgraded MongoDB infrastructure components
    • Stabilized delegate reassignment and reconnection processing

    Services recovered progressively beginning at approximately 15:47 UTC, with full stability restored by ~17:00 UTC.

    # Preventive Actions

    To prevent such issues from happening again, We are implementing the following improvements

    • Improving our circuit breakers and fail-fast protections between dependent services
    • Enhancing monitoring and alerting for thread pool saturation and queue buildup
    • Increasing baseline service headroom and resiliency protections