{
  "meta": {
    "version": "v1",
    "pricing": {
      "public": {
        "label": "Public",
        "description": "Read-only API access for lightweight status checks and public integrations."
      },
      "premium": {
        "label": "Premium",
        "description": "API keys with higher hourly quotas, plus Slack, Discord, webhook, and email outage alerts across your vendor stack."
      }
    },
    "generatedAt": "2026-07-22T20:05:54.610Z"
  },
  "data": {
    "id": "incident_statuspage_scaleway_bsp2y5fysy9w",
    "slug": "scaleway-fr-par-1-issue-with-compute-instances-2026-07-21",
    "title": "[fr-par-1] - Issue with Compute Instances",
    "summary": "[fr-par-1] - Issue with Compute Instances",
    "status": "resolved",
    "severity": "major",
    "startedAt": "2026-07-21T21:03:41.752+00:00",
    "updatedAt": "2026-07-22T16:02:26.107+00:00",
    "resolvedAt": "2026-07-22T04:55:46.186+00:00",
    "provider": {
      "slug": "scaleway",
      "name": "Scaleway"
    },
    "affectedServices": [],
    "links": {
      "html": "/incidents/scaleway-fr-par-1-issue-with-compute-instances-2026-07-21",
      "api": "/api/v1/incidents/scaleway-fr-par-1-issue-with-compute-instances-2026-07-21",
      "providerHtml": "/providers/scaleway"
    },
    "impactSummary": "Scaleway reported a major event for the affected tracked services.",
    "source": {
      "id": "source_scaleway_status",
      "kind": "official_api",
      "name": "Scaleway Status",
      "checkedAt": "2026-07-22T20:00:22.253+00:00",
      "officialUrl": "https://status.scaleway.com",
      "statusPageUrl": "https://status.scaleway.com"
    },
    "updates": [
      {
        "id": "update_statuspage_scaleway_bsp2y5fysy9w_3l24jyzlrhwp",
        "status": "investigating",
        "body": "Following a crash of one node on the block storage cluster at 20h15 UTC, some VMs are stuck.\nThe oncall team is investigating it.",
        "createdAt": "2026-07-21T21:03:41.846+00:00"
      },
      {
        "id": "update_statuspage_scaleway_bsp2y5fysy9w_vwrtfbzstjzr",
        "status": "investigating",
        "body": "We are continuing to investigate this issue.",
        "createdAt": "2026-07-21T21:13:42.471+00:00"
      },
      {
        "id": "update_statuspage_scaleway_bsp2y5fysy9w_g7wmtkw00qyr",
        "status": "investigating",
        "body": "We are still investigating this critical issue with the utmost priority.",
        "createdAt": "2026-07-21T21:39:56.712+00:00"
      },
      {
        "id": "update_statuspage_scaleway_bsp2y5fysy9w_phtxy0x3cp8j",
        "status": "investigating",
        "body": "Our team is currently working on solving the issue.",
        "createdAt": "2026-07-21T22:50:21.139+00:00"
      },
      {
        "id": "update_statuspage_scaleway_bsp2y5fysy9w_0wzh11jpc2m3",
        "status": "investigating",
        "body": "Disk I/O operations are starting to recover, which should begin resolving the issue for the instances.",
        "createdAt": "2026-07-21T23:35:49.443+00:00"
      },
      {
        "id": "update_statuspage_scaleway_bsp2y5fysy9w_68vqdfx55dbr",
        "status": "investigating",
        "body": "The situation is stable.\n\nAll products are now operational except for the Kapsule public API, which is still unavailable.",
        "createdAt": "2026-07-22T00:27:17.818+00:00"
      },
      {
        "id": "update_statuspage_scaleway_bsp2y5fysy9w_tfn4wfbk1x6w",
        "status": "monitoring",
        "body": "The Kapsule public API is available again.",
        "createdAt": "2026-07-22T01:41:42.989+00:00"
      },
      {
        "id": "update_statuspage_scaleway_bsp2y5fysy9w_srxkvd3q7mjh",
        "status": "investigating",
        "body": "We are currently investigating this issue.",
        "createdAt": "2026-07-22T03:30:23.393+00:00"
      },
      {
        "id": "update_statuspage_scaleway_bsp2y5fysy9w_p6mtnr347stm",
        "status": "monitoring",
        "body": "The situation has returned to normal",
        "createdAt": "2026-07-22T03:53:16.45+00:00"
      },
      {
        "id": "update_statuspage_scaleway_bsp2y5fysy9w_2bnvbs40sbcz",
        "status": "resolved",
        "body": "This incident has been resolved.",
        "createdAt": "2026-07-22T04:55:46.186+00:00"
      },
      {
        "id": "update_statuspage_scaleway_bsp2y5fysy9w_1dq9h46kzh7k",
        "status": "resolved",
        "body": "**Incident Overview**\n\nOn July 21, a block storage cluster incident affected the Instance product on FR-PAR-1 between 20:14 and 03:45 UTC. This incident caused a regional service degradation \\(FR-PAR\\), resulting in the following impacts:\n\n* Kapsule operations were unavailable on a regional level between 21:30 and 01:30 UTC.\n* Instances on FR-PAR-1 were non-functional.\n* All managed products relying on Instance and Kapsule experienced disruptions.\n\n#### **Root Cause Analysis** \n\n**5 whys**\n\n* **Why were Instance and its dependent products unavailable?**  \n  Because disk I/O was impossible.\n* **Why was disk I/O impossible?**  \n  Because cluster doesn't accept I/O operations.\n* **Why cluster doesn’t accept I/O operations?**  \n  Due to a fault in a cluster component \\(crash\\).\n* **Why was a cluster component faulty?**  \n  It suffered an OOM \\(Out of Memory\\) kill.\n* **Why did an OOM kill occur?**  \n  The memory limit was incorrectly configured, causing the component to consume excessive memory and triggering an OS-level OOM kill.\n* **Why was the memory limit incorrectly configured?**  \n  The cluster hardware is heterogeneous; nodes have varying memory capacities. A global setting was applied without accounting for nodes with lower memory, leading to node-level OOM kills.\n* **Why was the configuration error not detected?**  \n  We lack sufficient safeguards for this specific configuration.\n\n### Impact on Kapsule Product\n\n‌\n\n* **Why did an incident in a single Availability Zone \\(AZ\\) impact the entire FR-PAR region for Kapsule?**The Kapsule API was proactively disabled to prevent a cascading failure \\(\"snowball effect\"\\) caused by auto-healing mechanisms reacting to widespread instance unavailability.\n* **Why is a multi-AZ product affected by a single API endpoint failure?**The Kapsule API architecture is currently regional and lacks the granularity to isolate or disable specific Availability Zones. Consequently, disabling the regional API was the necessary precautionary measure to protect block storage convergence.\n\n#### **Summary of Events**\n\n#### **Incident Timeline \\(UTC\\)**\n\n| **Time \\(UTC\\)** | **Event Description** |\n| --- | --- |\n| 20:14 UTC | OOM kill on one OSD |\n| 20:28 UTC | First alert on block storage team |\n| 20:43 UTC | Escalation to larger teams, incident open at company level |\n| 20:48 UTC | sbs-api stops processing river jobs \\(no more volume update\\) |\n| 21:30 UTC | API Kapsule is unavailable |\n| 21:43 UTC | First restart of blk-api \\(internal api, not customer facing\\) |\n| 22:27 UTC | Second restart of blk-api, helped unstick api calls |\n| 23:48 UTC | Faulty OSD removed from production, throughput restored on cluster |\n| 23:50 UTC | Teams begin relaunching operations on disk |\n| 00:00 UTC | System was unavailable |\n| 01:30 UTC | API Kapsule is available |\n| 02:08 UTC | Beginning of second block cluster global failure |\n| 03:45 UTC | Block Cluster state restored successfully, end of impact |\n\n#### **Resolution and Improvements**\n\n**Short-term Actions**\n\n* During the incident timelapse, we identified a configuration issue and corrected the memory limits across all cluster nodes.\n\n**Mid-term Actions**\n\n* Enhance cluster configuration management to prevent environment-specific mismatches \\(e.g., node-level hardware heterogeneity\\).\n* Implement isolation mechanisms to ensure that individual component failures do not impact the overall cluster stability.\n* We also notice network saturation during the recovery leading to latencies increase, we may need to rework this part.\n* Optimize network performance during recovery phases to prevent latency spikes caused by saturation.\n* Improve the incident response to avoid logical bias\n\n#### **Contact**\n\nIf you have any further questions or need assistance, please contact our support team.",
        "createdAt": "2026-07-22T16:01:12.366+00:00"
      }
    ],
    "access": {
      "plan": "public",
      "keyed": false
    }
  }
}