GitHub Actions Degraded for Four Hours: 47% of Workflow Runs Failed at Peak
GitHub's initial report quantifies a four-hour Actions and hosted-runner outage: at peak, 47% of hosted-runner workflow runs failed and 71.9% of hosted jobs missed the five-minute start mark.
GitHub Actions suffered a consequential hosted-runner outage on October 5, 2026, and GitHub has now published an initial incident report with unusually specific failure data. From 18:48 to 22:49 UTC, Actions and Hosted Runners experienced degraded performance caused by a partial networking failure between one GitHub on-premises datacenter region and a subset of cloud-hosted regional database and storage services. The outage affected hosted-runner assignment, delayed workflow starts, and intermittently broke repository lists, licensing and billing pages, and some Copilot features. Self-hosted runners were not affected, according to GitHub’s resolved incident report.
GitHub Status records the following updates for October 5, 2026. Each time below is the publication time attached to that specific update on the official incident page; the incident report separately dates the start of service impact to 18:48 UTC.
| Update time (UTC, October 5) | GitHub’s reported state |
|---|---|
| 19:11 | GitHub began investigating reports of degraded Actions performance. |
| 19:15 | GitHub reported delays assigning hosted runners to Actions jobs. |
| 19:50 | GitHub was still investigating runner-assignment delays. |
| 20:39 | GitHub continued investigating job failures and delayed workflow starts. |
| 20:47 | GitHub reported degraded Actions availability. |
| 21:09 | GitHub reported repository-list, licensing and billing access issues alongside runner problems. |
| 21:22 | GitHub reported degraded Pages performance. |
| 21:31 | GitHub continued investigating degraded Actions performance. |
| 21:32 | GitHub reported applied mitigations: queued jobs were clearing and new jobs were no longer delayed. |
| 21:54 | GitHub reported Actions operating normally. |
| 22:40 | GitHub reported Pages operating normally. |
| 22:49 | GitHub marked the incident resolved. |

The measured blast radius
The most important new information is quantified failure data, not just the fact of an outage. GitHub said the roughly 90-minute period of highest impact was severe: 47.0% of workflow runs using GitHub-hosted runners failed, and 71.9% of GitHub-hosted jobs did not start within five minutes. Across the full incident, 14.3% of workflow runs and 26.5% of individual hosted jobs missed that five-minute start threshold.
| Window | Workflow runs failed | Hosted jobs not started within 5 minutes | Workflow runs not started within 5 minutes |
|---|---|---|---|
| Full incident, 18:48–22:49 UTC | Not stated | 26.5% | 14.3% |
| Roughly 90-minute peak | 47.0% | 71.9% | Not stated |
Those numbers make the incident more than a transient status-page event. For teams using AI coding agents, CI fix loops, scheduled maintenance agents, or overnight pipelines, hosted-runner startup delay is not merely an inconvenience. It changes queue depth, retry behavior, wall-clock cost, and whether an autonomous job appears stuck or actually failed. A CI agent that treats every delayed run as a failure can amplify the outage by generating duplicate runs; a job with no timeout can sit in a queue and burn budget while waiting.
The outage also illustrates the practical limits of managed CI. GitHub’s mitigation was to shift database and storage traffic to healthy regions and endpoints and add capacity to affected compute services. That recovery path worked, but it depended on operator-side failover and capacity expansion rather than user-controlled remediation. Teams with only GitHub-hosted runners had to wait for platform recovery, while teams with self-hosted runners were explicitly outside the affected runner path.
Follow-on disruption continued into October 6
The October 5 incident was not the end of the operational story. GitHub opened a separate incident at 23:47 UTC on October 5, initially describing it only as impacted performance for some GitHub services. GitHub then disclosed at 00:05 UTC on October 6 that users might have trouble accessing organization and enterprise billing and licensing pages, and resolved the incident at 01:32 UTC. GitHub said a detailed root cause analysis would be shared when available.
On October 6, a third incident reported degraded Pull Requests and Webhooks beginning around 19:57 UTC. GitHub said issues, pull requests, and webhooks briefly degraded and recovered, while enterprise migrations remained paused while the company validated causes and mitigations. The latest supplied update, at 22:40 UTC, said GitHub was deploying a mitigation and planned to resume migrations shortly thereafter. That follow-on incident should not be conflated with the original hosted-runner outage, but it shows the platform was still stabilizing related control-plane and data-plane paths a day later.
Why AI-agent operators should care
This is the kind of ordinary platform incident that can silently corrupt an autonomous workflow. An agent might see failed checks, open a fix branch, push a commit, and trigger another hosted run without knowing the first failure was infrastructure rather than code. In a high-traffic repository, that pattern can create noisy retries, duplicated CI minutes, conflicting branches, and misleading “fixes” for transient platform errors.
A safer design is to separate code failure from platform failure. Retry policies should be bounded and time-aware; agents should check GitHub’s public status and the repository’s own workflow telemetry before re-running a failed job; and expensive loops should require a human or policy gate when failure signatures resemble platform congestion. For agent fleets, the incident is also another argument for treating CI as a dependency with its own observability, not as a utility that is simply assumed to be available. Our guides to Agentic CI/CD and CI failure-to-fix loops cover the control patterns that keep autonomous repair work from thrashing during exactly this kind of outage.
The immediate service impact is resolved: GitHub says Actions and Hosted Runners recovered by 21:54 UTC, all affected services recovered by 22:49 UTC, and later billing and licensing disruption was resolved at 01:32 UTC on October 6. But the detailed root-cause analysis is still pending, and the October 6 migration pause shows the recovery was not entirely clean. For teams running agentic development pipelines, the practical lesson is to design CI autonomy for noisy shared infrastructure: bounded retries, status-aware escalation, self-hosted fallback where appropriate, and clear separation between a real regression and a platform-wide runner failure.