Skip to content

Incident detail

Intermittent slowness during pipeline executions (Prod1, Prod2)

Resolved incidentMinor1 affected service

Timeline window

to

Outage alerts

Get alerted the next time Harness breaks

Free email alerts for up to 5 providers — no card, live in about a minute. Paid plans add Slack, Discord, and webhook delivery across your whole stack, plus higher API quotas.

Timeline

Incident updates

Updates are normalized from the official source chronology so timeline changes remain easy to scan.

  1. Investigating

    We are currently investigating this issue.

  2. Monitoring

    A fix has been implemented and we are monitoring the results.

  3. Monitoring

    We are largely mitigated and most pipelines are running normally. We are monitoring all parameters to make sure there are no issues before closing it.

  4. Resolved

    This incident has been resolved.

  5. Resolved

    ### Summary

    A rollout involving OpenTelemetry instrumentation changes introduced a memory leak in the OTEL eBPF collector running in production clusters. Under sustained production traffic, the leak caused increasing JVM heap utilization, elevated garbage collection pressure, and eventual out-of-memory (OOM) conditions across several core platform services.

    ### Impact

    • Elevated latency and intermittent instability in Prod1, Prod2, and Prod3
    • Some customers experienced slow pipeline execution and degraded responsiveness

    No customer data loss occurred.

    ### Root Cause

    The root cause was an upstream defect in the OpenTelemetry eBPF instrumentation library that introduced a memory leak under production-scale workloads. The leak continuously increased telemetry-related memory consumption, leading to sustained JVM garbage collection pressure and eventual heap exhaustion.

    ### Mitigation and Recovery

    Immediate Actions

    • scaled up clusters to stabilize impacted clusters
    • Disabled OTEL instrumentation components and restarted affected services

    ### Next Steps

    To prevent such issues from happening again, we are:

    • Enhance our load testing process to test in higher workloads to identify such issues prior to going production.
    • Add additional granular instrumentation to catch such issues sooner.