Provider
CircleCIIncident detail
Elevated wait times for machine jobs
Timeline window
to
Get alerted the next time CircleCI breaks
Free email alerts for the handful of vendors you cannot afford to miss. No card, live in about a minute. Paid plans add Slack, Teams, Discord, and webhook delivery across your whole stack, plus higher API quotas.
Timeline
Incident updates
Every update on the official status source, oldest to newest, exactly as it appeared there.
Identified
Customers may be experiencing elevated wait times for machine jobs. We are working to resolve this.
Identified
Customers using Linux machine jobs are experiencing elevated wait times, averaging about 11 minutes, with the longest waits exceeding 30 minutes on the medium, arm.medium and arm.large resource classes. Our engineers have identified the issue and are working on a fix. We will provide another update by 14:00 UTC.
Identified
Customers using Linux machine jobs are experiencing elevated wait times, averaging about 40 minutes, with the longest waits exceeding 50 minutes. Most Linux machine resource classes are affected, including medium, large, xlarge, 2xlarge and their Arm equivalents. Our engineers have identified the issue and are working on a fix. We will provide another update by 14:30 UTC.
Identified
Customers using Linux machine jobs and remote Docker are experiencing elevated wait times. Wait times have started to decrease and now average about 20 minutes, with the longest waits exceeding 40 minutes on some resource classes. Our engineers have identified the issue and are working on a fix. We will provide another update by 15:00 UTC.
Identified
A fix has been deployed and wait times are decreasing, but customers using Linux machine jobs and Remote Docker are still experiencing delays. Wait times currently average about 11 minutes, with the longest waits exceeding 35 minutes on some resource classes. We are working to add capacity as quickly as possible. We will provide another update by 15:30 UTC.
Identified
Wait times continue to decrease, but customers using Linux machine jobs and remote Docker are still experiencing delays. Wait times currently average about 6 minutes. The longest waits, up to about 20 minutes, are on the 2xlarge, arm.2xlarge and gpu.nvidia.small resource classes. We are working to add capacity as quickly as possible. We will provide another update by 16:00 UTC.
Identified
Wait times continue to decrease, but customers using Linux machine jobs and remote Docker are still experiencing delays. Wait times currently average about 90 seconds, with the longest waits up to about 6 minutes on some resource classes. We are working to add capacity as quickly as possible. We will provide another update by 16:30 UTC.
Monitoring
Wait times for customers using Linux machine jobs and remote Docker have returned to normal. We are monitoring to confirm wait times remain stable while we continue to add capacity. We will provide another update by 16:30 UTC.
Resolved
Between 12:40 UTC and 16:00 UTC on October 8, customers using Linux machine jobs and remote Docker experienced elevated wait times. The issue has been resolved and wait times have returned to normal. We thank you for your patience while our team worked on implementing a fix.
Resolved
## Summary
Between October 1 and October 8, 2026, CircleCI customers experienced five separate incidents in which jobs waited longer than normal to start, or failed to start. Three of the incidents were driven by limited instance availability from our cloud provider. One was caused by a configuration issue in our scheduling infrastructure, and one by memory limits on an internal service, which we raised during the incident.
- October 1, 07:57 to 10:24 UTC: Docker Gen 2 jobs waited up to 30 minutes to start. A configuration issue prevented the Gen 2 scheduling servers from using their high-performance local disks. During a burst in traffic, the fallback disks saturated and job placement slowed.
- October 7, 13:27 to 14:06 UTC: Machine jobs waited longer than normal because we could not acquire enough virtual machines from our cloud provider.
- October 7, 14:24 to 15:21 UTC: Some jobs failed with an infrastructure error, and workflows were delayed. The service that gives each job its configuration at startup ran out of memory and restarted.
- October 7, 18:49 to 20:05 UTC: Docker Gen 2 jobs waited longer than normal because we could not acquire enough instances from our cloud provider.
- October 8, 12:48 to 16:00 UTC: Linux machine jobs and remote Docker jobs waited up to 50 minutes to start. Our cloud provider did not have enough of the instance type we request, and a defect in the logic that picks a different instance type kept us from using alternatives.
Jobs that waited during these incidents were delayed, not lost. Customers whose jobs failed with an infrastructure error during the October 7 incident from 14:24 to 15:21 UTC can rerun them.
The original status pages can be found below:
- October 1 - Delays starting Gen 2 Docker Jobs
- October 7 - Delay on starting Machine Job Tasks
- October 7 - Elevated level of infra fails on customer jobs
- October 7 - Increased task wait times for Docker Gen2
- October 8 - Elevated wait times for machine jobs
## Background
Every job that runs on CircleCI needs compute. Where that compute comes from depends on the executor.
Docker jobs run on clusters of servers that schedule work onto a fleet of instances. The scheduler is the software on those servers that decides which instance runs each job. Gen 2 Docker jobs ran on one such cluster at the time of the October 1 incident. Gen 1 Docker jobs are spread across several clusters.
Machine jobs, including remote Docker jobs, get a dedicated virtual machine. A remote Docker job is handled as a Linux machine job from submission through provisioning. Our machine provisioning service requests each machine from our cloud provider in one of the two regions we normally use for machine jobs. When the provider cannot supply the instance type we request, the service falls back: it requests a different instance type, or a different region, from a list we maintain. A job waits until a machine is available.
When any job starts, it asks a job configuration service for its configuration. That service runs as a set of pods, which are small copies of the same program. If too many pods stop, the remaining pods carry more load, and jobs cannot start until the service recovers.
These paths depend on three kinds of capacity: the throughput of our scheduling systems, the supply of instances from our cloud provider, and the memory and number of copies of the services that start jobs. A shortfall in any of them makes jobs wait or fail.
## What Happened
_(All times UTC)_
### October 1: Docker Gen 2 jobs delayed up to 30 minutes
At 07:44, disk activity on the servers that schedule Gen 2 Docker jobs rose to 90% of capacity and stayed there. At 07:57, the number of Docker jobs waiting to start began to grow, and wait times rose with it. An automated alert paged the on-call engineer at 08:08. We declared an incident at 08:42.
At 08:43, engineers identified high disk latency on the scheduling servers. They first believed the disks were too slow for the load, so they prepared a change to a faster disk type. At 09:01, the team agreed on two workstreams: move the servers to faster disks, and add a second Gen 2 cluster so that jobs would split across two clusters.
At 09:41, engineers found the root cause. Each scheduling server has fast disks attached directly to it, and the scheduler is meant to keep its working data there. These servers run on a security-hardened operating system image, and on that image the fast disks were not being used. The scheduler wrote its data to slower network-attached disks instead. The script that sets up the fast disks failed without an error, and a gap in logging on these servers hid the failure. Gen 2 traffic had also become increasingly bursty. The slower disks saturated during those bursts, and the scheduler fell behind.
Wait times peaked at 09:54, when jobs on the Gen 2 medium resource class waited up to 30 minutes. At 09:55, engineers applied a fix that set up the fast disks correctly. At 10:11, the scheduler began placing jobs faster, which confirmed the servers could keep up with new jobs. Wait times for Gen 2 medium fell to 13 minutes at 10:14, 7 minutes at 10:17, and 1 minute 30 seconds at 10:20. We moved the status page to monitoring at 10:24 and resolved the incident at 10:45.
At 11:20, the second Gen 2 cluster began running jobs.
### October 7, 13:27 to 14:06: Elevated wait times for machine jobs
We declared an incident at 13:27 after seeing elevated queueing and wait times. We posted to the status page at 13:31. At 13:31, engineers saw that we were struggling to acquire virtual machines for machine jobs, and that the number of jobs waiting to start was high mainly for machine jobs. At 13:34, wait times were elevated across executors.
At 13:48, wait times began to recover as more machines became available, and we resolved the incident at 14:06.
### October 7, 14:24 to 15:21: Job failures and workflow delays
Shortly after the previous incident was marked as resolved, we declared a new incident at 14:24 due to a sharp drop in the number of jobs submitted to our execution systems. At 14:26, engineers saw that pods of the job configuration service were restarting. Pods that restart leave fewer copies of the service to answer requests, and jobs that could not retrieve their configuration after repeated attempts failed with an infrastructure error. Workflows that were already running waited for updates from jobs that had not started.
At 14:35, engineers increased the amount of available memory to the pods and increased the number of pods, which gave the service more copies to share the load. We also increased the amount of memory available Job submissions began recovering, and the number of messages flowing through our workflow system returned to normal levels. At 14:44, we moved the status page to monitoring.
Pods in the job configuration service reached their memory limits and restarted. We increased memory allocations and pod counts, and the service recovered. At 15:13, we increased the number of instances in our Gen 2 Docker clusters so we could work through the backlog of waiting jobs quickly. We resolved the incident at 15:21.
### October 7, 18:49 to 20:05: Elevated wait times for Docker Gen 2 jobs
We declared a new incident at 18:49 after Docker Gen 2 wait times rose. At 18:51, engineers confirmed that a shortage of available instances kept us from acquiring enough of them, which raised wait times. Wait times differed by resource class size.
At 19:52, average wait times were lower than an hour earlier as more capacity became available, though brief spikes continued, and we moved the status page to monitoring. We resolved the incident at 20:05.
### October 8: Linux machine and remote Docker jobs wait up to 50 minutes
We declared an incident at 12:48 after seeing a large number of jobs waiting to start and long wait times for machine and remote Docker jobs. At 12:49, engineers saw that our cloud provider could not supply enough instances. Delays continued to climb as we investigated.
At 13:13, we began examining the logic that chooses a different instance type or region when our first choice is unavailable. At 13:24, we suspected that a defect in that logic was stopping us from using other instance types in our primary region, so we disabled our fallback region and confirmed the cause of the defect. Engineers began working on a fix and deployed it at 14:49. The service began requesting other instance types and started machines in large numbers. At 14:58, the queue of jobs waiting for a machine dropped steadily. Average wait times across resource classes continued to drop over the next hour as we worked through the backlog of waiting jobs.
At 16:00, the queue of jobs waiting for a machine cleared and the number of waiting jobs returned to normal. We moved the status page to monitoring at 16:09 and resolved the incident at 16:35.
## Future Prevention and Process Improvement
We are taking the following steps to strengthen the resilience of job start times, including supply diversification, fallback logic, and earlier detection.
We are expanding our compute regions. In the October 7 and October 8 incidents, our cloud provider could not supply enough instances in the regions we used. We are adding new regions to our fallback logic in order handle spikes for demand that outpace our preferred regions.
We are fixing the logic that picks a different instance type. On October 8, a defect kept the machine provisioning service from using other instance types in the same region. We deployed a fix during the incident and are now reviewing and refactoring this logic as a whole, with tests, so that a shortage of one instance type moves work to another.
We are expanding the instance types we offer and securing additional capacity. We are working with our cloud provider to reserve additional capacity in our existing regions, and we are adding instance types that jobs can run on. Industry demand for high-performance compute is growing. We are diversifying where and how we source capacity so our customers are insulated from supply variation.
We added a second Gen 2 Docker cluster so traffic bursts are distributed across two clusters. We are also adding alerts on scheduling server disk saturation and on how fast the scheduler places jobs, and an alert that fires when a server starts without its fast disks in use. We made the disk setup script fail with an error, and we restored log shipping from these servers so that a similar problem is visible sooner.
Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve these incidents. Please reach out to our support team with any questions or concerns.
Keep exploring
More from CircleCI
Neighboring incidents on CircleCI's timeline and the rest of their record on OutageDeck.