{
  "meta": {
    "version": "v1",
    "pricing": {
      "public": {
        "label": "Public",
        "description": "Read-only API access for lightweight status checks and public integrations."
      },
      "premium": {
        "label": "Premium",
        "description": "API keys with higher hourly quotas, plus Slack, Discord, webhook, and email outage alerts across your vendor stack."
      }
    },
    "generatedAt": "2026-07-24T10:19:59.057Z"
  },
  "data": {
    "id": "incident_statuspage_harness_x0yrvcxzqj12",
    "slug": "harness-intermittent-slowness-during-pipeline-executions-prod1-prod2-2026-05-01",
    "title": "Intermittent slowness during pipeline executions (Prod1, Prod2)",
    "summary": "Intermittent slowness during pipeline executions (Prod1, Prod2)",
    "status": "resolved",
    "severity": "minor",
    "startedAt": "2026-05-01T15:02:27.269+00:00",
    "updatedAt": "2026-05-11T21:08:23.025+00:00",
    "resolvedAt": "2026-05-01T20:10:33.382+00:00",
    "provider": {
      "slug": "harness",
      "name": "Harness"
    },
    "affectedServices": [
      {
        "slug": "harness-cicd",
        "name": "CI/CD"
      }
    ],
    "links": {
      "html": "/incidents/harness-intermittent-slowness-during-pipeline-executions-prod1-prod2-2026-05-01",
      "api": "/api/v1/incidents/harness-intermittent-slowness-during-pipeline-executions-prod1-prod2-2026-05-01",
      "providerHtml": "/providers/harness"
    },
    "impactSummary": "Harness reported a minor event for the affected tracked services.",
    "source": {
      "id": "source_harness_status",
      "kind": "official_status_page",
      "name": "Harness Status",
      "checkedAt": "2026-07-23T12:00:00Z",
      "officialUrl": "https://status.harness.io",
      "statusPageUrl": "https://status.harness.io"
    },
    "updates": [
      {
        "id": "update_statuspage_harness_x0yrvcxzqj12_z7q2yz3kqlg4",
        "status": "investigating",
        "body": "We are currently investigating this issue.",
        "createdAt": "2026-05-01T15:02:27.465+00:00"
      },
      {
        "id": "update_statuspage_harness_x0yrvcxzqj12_pmwwmlg1m9kb",
        "status": "monitoring",
        "body": "A fix has been implemented and we are monitoring the results.",
        "createdAt": "2026-05-01T15:37:32.779+00:00"
      },
      {
        "id": "update_statuspage_harness_x0yrvcxzqj12_qk72zvmrwpc3",
        "status": "monitoring",
        "body": "We are largely mitigated and most pipelines are running normally. We are monitoring all parameters to make sure there are no issues before closing it.",
        "createdAt": "2026-05-01T19:58:35.564+00:00"
      },
      {
        "id": "update_statuspage_harness_x0yrvcxzqj12_m55ky1nwqj1d",
        "status": "resolved",
        "body": "This incident has been resolved.",
        "createdAt": "2026-05-01T20:10:33.382+00:00"
      },
      {
        "id": "update_statuspage_harness_x0yrvcxzqj12_gkpc1kch1qfd",
        "status": "resolved",
        "body": "### **Summary**\n\nA rollout involving OpenTelemetry instrumentation changes introduced a memory leak in the OTEL eBPF collector running in production clusters. Under sustained production traffic, the leak caused increasing JVM heap utilization, elevated garbage collection pressure, and eventual out-of-memory \\(OOM\\) conditions across several core platform services.  \n\n### **Impact**\n\n* Elevated latency and intermittent instability in Prod1, Prod2, and Prod3\n* Some customers experienced slow pipeline execution and degraded responsiveness\n\nNo customer data loss occurred.\n\n### **Root Cause**\n\nThe root cause was an upstream defect in the OpenTelemetry eBPF instrumentation library that introduced a memory leak under production-scale workloads. The leak continuously increased telemetry-related memory consumption, leading to sustained JVM garbage collection pressure and eventual heap exhaustion.    \n\n### **Mitigation and Recovery**\n\n**Immediate Actions**\n\n* scaled up clusters  to stabilize impacted clusters\n* Disabled OTEL instrumentation components and restarted affected services\n\n### **Next Steps**\n\nTo prevent such issues from happening again, we are: \n\n* Enhance our load testing process to test in higher workloads to identify such issues prior to going production.\n* Add additional granular instrumentation to catch such issues sooner.",
        "createdAt": "2026-05-11T20:58:50.278+00:00"
      }
    ],
    "access": {
      "plan": "public",
      "keyed": false
    }
  }
}