Provider
Neo4j AuraIncident detail
Aura metrics issue.
Timeline window
to
Get alerted the next time Neo4j Aura breaks
Free email alerts for the handful of vendors you cannot afford to miss. No card, live in about a minute. Paid plans add Slack, Teams, Discord, and webhook delivery across your whole stack, plus higher API quotas.
Timeline
Incident updates
Every update on the official status source, oldest to newest, exactly as it appeared there.
Investigating
Our team has identified an issue with collecting the metrics for Aura instances. This will impact the metric visibility and forwarding.
We are currently investigating the issue.
Investigating
We continue to investigate the Aura metrics issue.
Investigating
We continue to investigate the Aura metrics issue.
Resolved
Since the last update, the investigation has moved on. The issue was identified, and we monitored the service as it recovered from degradation. We confirm the issue is now resolved.
Resolved
# What happened
On Thursday, 20 Aug 2026 14:58:00 UTC, our metrics ingestion pipeline began experiencing elevated export failures when sending data to Google Managed Prometheus (GMP). The underlying cause was a Google Cloud Platform service incident in the `us-west1` region, which disrupted the export path used by our metrics collection infrastructure.
As a result, metric export queues began building up across a subset of production environments. The failures manifested primarily as network timeouts when attempting to deliver metric data to GMP, preventing normal ingestion for affected clusters. Our engineering team identified the geographic concentration of affected systems — spanning AWS `us-west-2`, Google Cloud `us-west1`, and Azure `westus2` — all of which route through infrastructure in the Oregon area.
Once the GCP regional incident was mitigated by Google, export errors subsided and queues began draining.
# How the service was affected
The impact was limited to metrics visibility for a subset of Aura environments. Approximately 45 customer environments were affected, meaning that metric data from databases was not being ingested into our monitoring platform during the incident window.
The most affected region was AWS `us-west-2`, with additional impact in Google Cloud `us-west1` and Azure `westus2`. Customers with databases in these regions may have experienced gaps in their metrics dashboards and monitoring data for the duration of the incident. Core database operations — including reads, writes, and availability — were not affected; only the collection and visibility of performance metrics was disrupted.
Once Google mitigated the underlying regional issue, metric ingestion resumed and the backlog of queued data was processed.
# What we are doing now
We have conducted a thorough review of this incident and are taking the following steps to improve our resilience and response capabilities:
- Improved observability dashboards: We are enhancing our metrics exporter dashboards to provide clearer, faster visibility into queue health and export error rates, enabling quicker detection and diagnosis of similar issues.
- Improve alert aggregation: We are implementing automatic alert grouping so that a regional incident affecting multiple clusters is better surfaced as a single correlated event and treated with appropriate severity and attention.
- Proactive alerting on ingestion error rates: We are adding dedicated alerts that fire on elevated GMP ingestion error rates, giving our team an earlier signal before exporter queues reach capacity.
These improvements are designed to reduce detection time, simplify incident triage, and strengthen our metrics pipeline against future upstream provider disruptions.
Keep exploring
More from Neo4j Aura
Neighboring incidents on Neo4j Aura's timeline and the rest of their record on OutageDeck.