Provider
ScalewayIncident detail
[COCKPIT] - [fr-par/pl-waw] - Scaleway Metrics/Logs ingestion down
Timeline window
to
Outage alerts
Get alerted the next time Scaleway breaks
Free email alerts for up to 5 providers. No card, live in about a minute. Paid plans add Slack, Teams, Discord, and webhook delivery across your whole stack, plus higher API quotas.
Timeline
Incident updates
Every update Scaleway posted, oldest to newest, exactly as it appeared on their official status page.
Monitoring
Cockpit impacted this night betwen 1am30 and 8am30 with data loss for all Products metrics and logs on FR-PAR & PL-WAW Regions
Root cause : based on last week Cockpit ingestion solution, a high RAM consumption saturated nodes and corrupted replay solution.
Status : Cockpit fixed at 8am30 and fix in development to avoid replay corruption on OOM events
Monitoring
Root cause was finally identified yesterday evening, we make an evolution on staging environments and monitored its behavior today. We plan to deploy this fix next monday morning.
Metrics & Logs acquisition are fine since yesterday with a scaling workaround, we keep this solution for the week-end to not impact production on a friday afternoon.
Monitoring
We applied and monitored today the last fix for high RAM consumption. All seems fine now on Cockpit performances for all regions.
Resolved
This incident has been resolved.
Keep exploring
More from Scaleway
Neighboring incidents on Scaleway's timeline and the rest of their record on OutageDeck.