Provider
GitHubIncident detail
Incident with Pull Requests
Timeline window
to
Get alerted the next time GitHub breaks
Free email alerts for the handful of vendors you cannot afford to miss. No card, live in about a minute. Paid plans add Slack, Teams, Discord, and webhook delivery across your whole stack, plus higher API quotas.
Timeline
Incident updates
Every update GitHub posted, oldest to newest, exactly as it appeared on their official status page.
Investigating
We are investigating reports of degraded performance for Pull Requests
Investigating
We are investigating errors creating pull requests
Investigating
Pull Requests is experiencing degraded availability. We are continuing to investigate.
Investigating
We have applied a mitigation and are monitoring for recovery
Monitoring
The degradation affecting Pull Requests has been mitigated. We are monitoring to ensure stability.
Resolved
Between July 24, 19:17 UTC and July 24, 20:02 UTC, users were unable to create pull requests due to a database schema change. In total, 113,930 pull request creation attempts were impacted across 50,904 users, with an average error rate of 1.75% and a maximum error rate of 2.25% for all requests to Pull Requests service. Existing pull requests and other GitHub functionality were not affected. The issue was resolved by reverting the change to the affected database, upon which pull request creation immediately resumed.<br /><br />The root cause was related to a backfill workflow into the Vitess keyspace hosting Pull Request data. The backfill Vitess command encountered errors and increased VReplication lag, and the workflow was canceled at 19:17 UTC. The cancellation executed a misunderstood Vitess codepath that dropped the backing table to the target keyspace, leaving a non-existent reference that resulted in errors creating Pull Requests. The mitigation was executing a command to drop the vschema reference to the dropped table, allowing Pull Request creation to resume.<br /><br />We are adding stronger pre-flight validation to our tooling to prevent similar issues and expanding lower-environment support to provide better test coverage end-to-end before promoting them to production. We're also fixing our backfill migration tooling to protect from this specific codepath.
Keep exploring
More from GitHub
Neighboring incidents on GitHub's timeline and the rest of their record on OutageDeck.