{
  "meta": {
    "version": "v1",
    "pricing": {
      "public": {
        "label": "Public",
        "description": "Read-only API access for lightweight status checks and public integrations."
      },
      "premium": {
        "label": "Premium",
        "description": "API keys with higher hourly quotas, plus Slack, Discord, webhook, and email outage alerts across your vendor stack."
      }
    },
    "generatedAt": "2026-07-23T08:04:58.180Z"
  },
  "data": {
    "id": "incident_statuspage_buildkite_3vngjcnbkycl",
    "slug": "buildkite-delayed-notifications-2026-05-20",
    "title": "Delayed notifications",
    "summary": "Delayed notifications",
    "status": "resolved",
    "severity": "major",
    "startedAt": "2026-05-20T16:40:24.52+00:00",
    "updatedAt": "2026-05-25T07:04:02.67+00:00",
    "resolvedAt": "2026-05-20T17:39:34.444+00:00",
    "provider": {
      "slug": "buildkite",
      "name": "Buildkite"
    },
    "affectedServices": [],
    "links": {
      "html": "/incidents/buildkite-delayed-notifications-2026-05-20",
      "api": "/api/v1/incidents/buildkite-delayed-notifications-2026-05-20",
      "providerHtml": "/providers/buildkite"
    },
    "impactSummary": "Buildkite reported a major event for the affected tracked services.",
    "source": {
      "id": "source_buildkite_status",
      "kind": "official_api",
      "name": "Buildkite Status",
      "checkedAt": "2026-07-23T07:55:03.022+00:00",
      "officialUrl": "https://www.buildkitestatus.com",
      "statusPageUrl": "https://www.buildkitestatus.com"
    },
    "updates": [
      {
        "id": "update_statuspage_buildkite_3vngjcnbkycl_gf2x7pqwlswl",
        "status": "investigating",
        "body": "We are investigating delays to notifications across all customers",
        "createdAt": "2026-05-20T16:40:24.671+00:00"
      },
      {
        "id": "update_statuspage_buildkite_3vngjcnbkycl_2f84259nx8dk",
        "status": "identified",
        "body": "We have identified the issue and applied mitigations and are monitoring recovery\n\nWe have determined that only a subset of customers are affected by the notification latency.",
        "createdAt": "2026-05-20T17:06:51.515+00:00"
      },
      {
        "id": "update_statuspage_buildkite_3vngjcnbkycl_lv4cnzhxs1cn",
        "status": "monitoring",
        "body": "We are seeing recovery across affected customers and continue to monitor",
        "createdAt": "2026-05-20T17:26:30.42+00:00"
      },
      {
        "id": "update_statuspage_buildkite_3vngjcnbkycl_tstpvbsy2v0t",
        "status": "resolved",
        "body": "The incident is resolved",
        "createdAt": "2026-05-20T17:39:34.444+00:00"
      },
      {
        "id": "update_statuspage_buildkite_3vngjcnbkycl_1zn25t44nvgw",
        "status": "resolved",
        "body": "## Service Impact\n\nA subset of our customers experienced elevated latency in our notification delivery, build dispatch and metrics services.\n\n## Incident Summary\n\nWe are in the process of migrating our underlying compute platform from AWS Fargate to AWS EKS for our production workloads. We are migrating our services in small batches so we can verify stability as we go.\n\nBetween 15:42 and 17:33 our EKS Prometheus server began to need more memory than was available on the host where it was running. This was caused by autoscaling operations that increased the number of pods tracked by Prometheus, which in turn increased the Prometheus server's memory requirement. The host killed the Prometheus server process, which was restarted shortly after by the Kubernetes control plane. In the interim, the metrics used for application autoscaling were unavailable. The unavailable metrics meant that the affected services were not being triggered to scale up, resulting in the observed delays. Prometheus exceeded the host's available memory again soon after restarting, which caused the cycle to repeat.\n\nThe on call team followed a prepared documentation to shift load on the affected services back to Fargate. The majority of customers saw complete recovery from 16:49. A handful of customers had developed such a large backlog during the period of higher latency, that they had to be manually scaled up further. All customers saw full recovery by 17:33.\n\n## Changes we're making\n\nWe have already made the following changes to our rollout of EKS for production workloads:\n\n* Upsized the underlying system nodes.\n* Set higher requests and limits for the Prometheus server so it can handle more product load.\n* Reviewed and set any missing requests and limits for all new EKS resources, ensuring that EKS has all the required information to prevent accidental resource contention.\n* Added more observability and monitors for EKS pod and node health to help us identify root causes quickly during future incidents.\n\nWe have since migrated all these services back to EKS and observed successful scaling well beyond the limits we encountered during this incident.",
        "createdAt": "2026-05-25T07:02:51.303+00:00"
      }
    ],
    "access": {
      "plan": "public",
      "keyed": false
    }
  }
}