{
  "meta": {
    "version": "v1",
    "pricing": {
      "public": {
        "label": "Public",
        "description": "Read-only API access for lightweight status checks and public integrations."
      },
      "premium": {
        "label": "Premium",
        "description": "API keys with higher hourly quotas, plus Slack, Discord, webhook, and email outage alerts across your vendor stack."
      }
    },
    "generatedAt": "2026-07-23T09:51:50.388Z"
  },
  "data": {
    "id": "incident_statuspage_harness_yx38bdw9gryq",
    "slug": "harness-prod3-experiences-slowness-in-pipelines-2026-05-01",
    "title": "Prod3 experiences slowness in pipelines",
    "summary": "Prod3 experiences slowness in pipelines",
    "status": "resolved",
    "severity": "minor",
    "startedAt": "2026-05-01T16:28:47.917+00:00",
    "updatedAt": "2026-05-11T21:09:01.77+00:00",
    "resolvedAt": "2026-05-01T20:10:46.509+00:00",
    "provider": {
      "slug": "harness",
      "name": "Harness"
    },
    "affectedServices": [
      {
        "slug": "harness-cicd",
        "name": "CI/CD"
      },
      {
        "slug": "harness-feature-flags",
        "name": "Feature flags & FME"
      }
    ],
    "links": {
      "html": "/incidents/harness-prod3-experiences-slowness-in-pipelines-2026-05-01",
      "api": "/api/v1/incidents/harness-prod3-experiences-slowness-in-pipelines-2026-05-01",
      "providerHtml": "/providers/harness"
    },
    "impactSummary": "Harness reported a minor event for the affected tracked services.",
    "source": {
      "id": "source_harness_status",
      "kind": "official_api",
      "name": "Harness Status",
      "checkedAt": "2026-07-23T09:45:03.435+00:00",
      "officialUrl": "https://status.harness.io",
      "statusPageUrl": "https://status.harness.io"
    },
    "updates": [
      {
        "id": "update_statuspage_harness_yx38bdw9gryq_2c9ld4r2bzmr",
        "status": "investigating",
        "body": "We are currently investigating this issue.",
        "createdAt": "2026-05-01T16:28:48.025+00:00"
      },
      {
        "id": "update_statuspage_harness_yx38bdw9gryq_ndzlqnxmq3by",
        "status": "monitoring",
        "body": "A fix has been implemented and we are monitoring the results.",
        "createdAt": "2026-05-01T16:56:40.48+00:00"
      },
      {
        "id": "update_statuspage_harness_yx38bdw9gryq_99dkr3v3n2v1",
        "status": "monitoring",
        "body": "We are continuing to monitor for any further issues.",
        "createdAt": "2026-05-01T17:14:47.927+00:00"
      },
      {
        "id": "update_statuspage_harness_yx38bdw9gryq_kw5t63xswxmx",
        "status": "monitoring",
        "body": "We are continuing to monitor for any further issues.",
        "createdAt": "2026-05-01T17:15:39.259+00:00"
      },
      {
        "id": "update_statuspage_harness_yx38bdw9gryq_0y85987lfhl6",
        "status": "monitoring",
        "body": "We are largely mitigated and most pipelines are running normally. We are monitoring all parameters to make sure there are no issues before closing it.",
        "createdAt": "2026-05-01T19:58:18.724+00:00"
      },
      {
        "id": "update_statuspage_harness_yx38bdw9gryq_jbqxyphys0qv",
        "status": "resolved",
        "body": "This incident has been resolved.",
        "createdAt": "2026-05-01T20:10:46.509+00:00"
      },
      {
        "id": "update_statuspage_harness_yx38bdw9gryq_2c2hz3bfsdy3",
        "status": "resolved",
        "body": "### **Summary**\n\nA rollout involving OpenTelemetry instrumentation changes introduced a memory leak in the OTEL eBPF collector running in production clusters. Under sustained production traffic, the leak caused increasing JVM heap utilization, elevated garbage collection pressure, and eventual out-of-memory \\(OOM\\) conditions across several core platform services.\n\n### **Impact**\n\n* Elevated latency and intermittent instability in Prod1, Prod2, and Prod3\n* Some customers experienced slow pipeline execution and degraded responsiveness\n\nNo customer data loss occurred.\n\n### **Root Cause**\n\nThe root cause was an upstream defect in the OpenTelemetry eBPF instrumentation library that introduced a memory leak under production-scale workloads. The leak continuously increased telemetry-related memory consumption, leading to sustained JVM garbage collection pressure and eventual heap exhaustion.\n\n### **Mitigation and Recovery**\n\n**Immediate Actions**\n\n* scaled up clusters to stabilize impacted clusters\n* Disabled OTEL instrumentation components and restarted affected services\n\n### **Next Steps**\n\nTo prevent such issues from happening again, we are:\n\n* Enhance our load testing process to test in higher workloads to identify such issues prior to going production.\n* Add additional granular instrumentation to catch such issues sooner.",
        "createdAt": "2026-05-11T21:08:59.833+00:00"
      }
    ],
    "access": {
      "plan": "public",
      "keyed": false
    }
  }
}