Provider
HarnessIncident detail
CI Linux Cloud VM Networking issues - intermittent
Timeline window
to
Get alerted the next time Harness breaks
Free email alerts for the handful of vendors you cannot afford to miss. No card, live in about a minute. Paid plans add Slack, Teams, Discord, and webhook delivery across your whole stack, plus higher API quotas.
Timeline
Incident updates
Every update on the official status source, oldest to newest, exactly as it appeared there.
Investigating
We are currently investigating this issue.
Monitoring
A fix has been implemented and we are monitoring the results.
Resolved
This incident has been resolved.
Resolved
### Summary
Between approximately 11:37 AM and 3:01 PM PDT (18:37–22:01 UTC) on September 29, 2026, customers using Harness CI self-hosted Linux cloud runners in prod1/prod2 (us-west1) experienced intermittent networking failures during pipeline execution — primarily Git/network timeouts when fetching source code.
Root Cause
This was primarily due to saturation of our NAT which was futher exacerbated by a temporary cloud provider compute capacity, preventing new VM provisioning there.
### Impact
- Customers running self-hosted Linux CI builds in us-west1 saw intermittent build/pipeline failures, not a full outage.
- Observed customer-facing errors included Git fetch timeouts, e.g. `Failed to connect to github.com port 443: Connection timed out`
Remediation
- Immediate: Rebalanced ARM build-VM load from the saturated us-west1 path toward central1 (partially limited by central1 capacity); restarted the affected Cloud NAT gateways, which stabilized pipeline networking.
Action Items
1. Secure additional NAT gateway/IP capacity (In progress)
2. Improve NAT saturation detection and alerting ahead of customer impact.
3. Evaluate cross-region capacity contingencies so a single region's compute shortage doesn't block failover mitigation.
Keep exploring
More from Harness
Neighboring incidents on Harness's timeline and the rest of their record on OutageDeck.