Provider
OpsgenieIncident detail
Disrupted Opsgenie/JSM availability
Timeline window
to
Get alerted the next time Opsgenie breaks
Free email alerts for the handful of vendors you cannot afford to miss. No card, live in about a minute. Paid plans add Slack, Teams, Discord, and webhook delivery across your whole stack, plus higher API quotas.
Timeline
Incident updates
Every update on the official status source, oldest to newest, exactly as it appeared there.
Investigating
We are actively investigating reports of a service disruption affecting Opsgenie and Jira Service Management. We will share updates here as more information is available.
Identified
We have identified the likely cause of the issue, and our teams are diligently working on a mitigation. We will continue to share additional updates here as more information is available.
Monitoring
The issue has now been resolved, and services are operating normally for all affected customers. We will continue to monitor closely to confirm stability.
Resolved
On September 14, 2026, Opsgenie and Jira Service Management experienced a disruption, and impacted users saw delayed alerts notification and inability to view the alerts in the user interface.
The issue has now been resolved, and the service is operating normally for all affected customers.
Resolved
### Summary
On September 14, 2026, between 15:30 and 17:10 UTC, Atlassian customers using Opsgenie, Jira Service Management, and Compass experienced delays in receiving alert notifications and intermittent failures when using Ops features. The issue was triggered by long-running transactions in a core service, which led to thread pool exhaustion and caused subsequent requests to become unresponsive. The long-running transactions were caused by a configuration change released between September 9, 2026 and September 11, 2026 UTC. The team actively monitored the impact of the change on the system until early September 14, 2026 UTC, and observed no anomalies. However, increasing traffic during US working hours on September 14, 2026, caused these unexpectedly long-running transactions. The incident was detected within one minute by our automated monitoring systems, and full service stabilization occurred after isolating and rolling back the change on September 14, 2026 at 17:10 UTC.
### IMPACT
The incident primarily affected Ops features across Opsgenie, Jira Service Management, and Compass for customers hosted in the US region. During the incident, the end-to-end flow for alert creation and notification delivery experienced an average delay of 22 minutes. No alerts or notification payloads were dropped during the incident; all queued events were successfully processed and delivered as services recovered. Additionally, related Ops API endpoints and Web UI flows experienced elevated latency and intermittent errors. The disruption began at 15:30 UTC and was resolved by 17:10 UTC (total duration: 1 hour and 40 minutes).
Engineering teams were alerted immediately via backup disaster-alerting pipelines and began mitigation without delay. However, the disruption affected internal notification and coordination flows between cross-functional teams (including Customer Support and Incident Communications), which resulted in delays in publishing external Statuspage updates.
### ROOT CAUSE
The incident was triggered when a newly released configuration change made redundant downstream service calls on every page load, causing a traffic spike. The configuration change, released between September 9, 2026, and September 11, 2026, was actively monitored for impact until early September 14, 2026. However, a traffic spike during US working hours caused unexpected issues. The downstream service began rate-limiting requests, and an aggressive retry strategy without sufficient backoff held worker threads open, leading to pool exhaustion and timeouts on incoming critical traffic.
### REMEDIAL ACTIONS PLAN & NEXT STEPS
The incident was mitigated by disabling the feature flag that triggered the inter-service rate limiting.
We understand reliable access to Atlassian products is critical for your teams. We are prioritizing the following actions to prevent recurrence:
- Improve Service-to-Service Resilience
- Review and adjust rate-limiting configurations between internal services.
- Optimize retry strategies to prevent long-running transactions and keep the impact isolated to problematic point.
- Improve Architectural Isolation
- Refine the architecture of this core service to have more isolation between different transactions so that degradations in one transaction type do not impact unrelated critical flows.
- Improve Proactive Alerting
- Add early-warning alerts for thread pool saturation and elevated inter-service rate-limit responses before they impact end users.
- Implement backup alerting processes for customer support and incident communication paths
We apologize to customers whose services were impacted during this incident; we are taking immediate steps to improve the platform’s performance and availability.
Thanks,
Atlassian Customer Support
Keep exploring
More from Opsgenie
Neighboring incidents on Opsgenie's timeline and the rest of their record on OutageDeck.