We are investigating reports of impacted performance for some GitHub services.
resolved
This incident has been resolved. Thank you for your patience and understanding as we addressed this issue. A detailed root cause analysis will be shared as soon as it is available.
Disruption with Copilot Code Review
Started September 4, 2026 at 8:39 PM UTC · 1h 48m
OutageMajor incident
investigating
We are investigating reports of impacted performance for some GitHub services.
investigating
Some users may be experiencing failures when using Copilot code review. We have identified the root cause and are working on a mitigation.
investigating
We are applying the mitigation and expect recovery within approximately 30 minutes.
monitoring
The degradation has been mitigated. We are monitoring to ensure stability.
resolved
This incident has been resolved. Thank you for your patience and understanding as we addressed this issue. A detailed root cause analysis will be shared as soon as it is available.
Incident with Grok Copilot AI Model Provider
Started September 3, 2026 at 2:17 PM UTC · 2h 54m
IssuesMinor incident
Affected components
Copilot AI Model Providers
investigating
We are investigating reports of degraded performance for Copilot AI Model Providers
investigating
We are experiencing degraded availability for the Grok 4.6 model in Copilot Chat, VS Code and other Copilot products. This is due to an issue with an upstream model provider. We are working with them to resolve the issue.
investigating
The Grok 4.5 model has degraded availability as well. We are working with the upstream provider to resolve the issue.
investigating
The issues with our upstream model provider have been resolved, and Grok models are once again available in Copilot products and IDE surfaces.
We will continue monitoring to ensure stability, but mitigation is complete.
resolved
This incident has been resolved. Thank you for your patience and understanding as we addressed this issue. A detailed root cause analysis will be shared as soon as it is available.
Delays in commit processing
Started September 1, 2026 at 3:00 PM UTC · 1h 1m
IssuesMinor incident
Affected components
Pull Requests
investigating
We are investigating reports of degraded performance for Pull Requests
investigating
Diffs in the PR view may be stale for several minutes. We are investigating and scaling up resources.
investigating
Time to update pull request diffs have improved to normal thresholds.
resolved
This incident has been resolved. Thank you for your patience and understanding as we addressed this issue. A detailed root cause analysis will be shared as soon as it is available.
Elevated rate of errors for OpenAI models provided by Copilot
Started August 31, 2026 at 9:15 AM UTC · 42m
IssuesMinor incident
Affected components
Copilot AI Model Providers
investigating
We are investigating reports of degraded performance for Copilot AI Model Providers
investigating
Copilot is experiencing a higher rate of errors for OpenAI models, including gpt-5.2, gpt-5.3-codex, gpt-5.4, gpt-5.4, and the gpt-5.6 family of models. Other models are not impacted.
investigating
One of our model providers has confirmed an incident on their end. We have provided them details to help identify the issue. We are starting to see recovery.
monitoring
The degradation affecting Copilot AI Model Providers has been mitigated. We are monitoring to ensure stability.
monitoring
The issues with our upstream model provider have been resolved, and gpt-5.3-codex, gpt-5.4-mini, gpt-5.4-nano, gpt-5.5, and the gpt-5.6 family of models are once again available in Copilot products and IDE surfaces.
We will continue monitoring to ensure stability, but mitigation is complete.
resolved
Between 08:37 and 09:41 UTC on August 31, 2026, GitHub Copilot experienced degradation affecting several GPT models, including gpt-5.2, gpt-5.3-codex, gpt-5.4, gpt-5.4-mini, gpt-5.4-nano, and the gpt-5.6 family (Luna, Sol, and Terra). Users encountered elevated error rates and interrupted streaming responses. Other models were not affected.
The degradation was caused by an issue with an upstream model provider. GitHub engineers detected the issue through automated monitoring, displayed in-product warnings for the affected models, and coordinated with the provider. Service returned to normal after the provider implemented a mitigation.
Incident with Copilot AI Model Providers
Started August 27, 2026 at 10:04 AM UTC · 2h 8m
OutageCritical incident
Affected components
Copilot AI Model Providers
investigating
We are investigating reports of degraded availability for Copilot AI Model Providers
investigating
We are experiencing degraded availability for the Kimi K3 model in Copilot products and IDE surfaces. This is due to an issue with an upstream model provider. While we work with them to resolve the issue, we recommend choosing another model or selecting 'Auto' to continue using Copilot.
investigating
Copilot AI Model Providers is experiencing degraded performance. We are continuing to investigate.
investigating
The issues with our upstream model provider have been mitigated, and Kimi K3 is once again available in Copilot products and IDE surfaces.
We will continue monitoring to ensure stability.
resolved
On August 27th, 2026, between approximately 09:20 and 12:14 UTC, the Copilot service experienced a degradation of the Kimi K3 model due to an issue with our upstream provider. Users encountered elevated error rates when using Kimi K3. No other models were impacted.
The issue was resolved by a mitigation put in place by our provider. GitHub is working with our provider to further improve the resiliency of the service to prevent similar incidents in the future.
Incident with Actions and Pull Requests
Started August 26, 2026 at 10:56 PM UTC · 1h 30m
IssuesMinor incident
Affected components
Pull RequestsActions
investigating
We are investigating reports of degraded performance for Actions and Pull Requests
investigating
We are investigating elevated delays and timeouts affecting Actions workflow runs triggered by pull request events. 20% of actions runs have delayed starts of more than 5 minutes and up to 4% of runs failed to trigger. We are actively working on mitigation and will provide updates as we learn more.
investigating
We've applied mitigations and are seeing recovery in Actions workflow runs and blocked pull request merges. We're continuing to monitor for sustained health of merge commit creates before resolving.
investigating
We confirmed full recovery beginning at 23:58 UTC. Actions workflow runs and pull request merges are operating normally. We will now resolve the incident while continuing to monitor service health.
monitoring
The degradation affecting Actions and Pull Requests has been mitigated. We are monitoring to ensure stability.
resolved
On August 26, 2026, from 21:55 UTC to 23:58 UTC, 2.6% of workflow runs triggered by pull request events were delayed, with the impact rising as high as 25% at its peak. Some users also experienced delays in pull request merge-commit generation, mergeability information, and merge-button availability. Actions and Pull Requests fully recovered by 23:58 UTC; the incident was resolved at 00:26 UTC after normal operation was confirmed.
Background jobs that process pull request updates and generate merge commits were impacted by timeouts reaching a single partition of git data. This resulted in a backlog in pull request merge-commit processing, delaying pull request-triggered GitHub Actions workflows and some mergeability information.
We reduced workload, shifted traffic away from affected infrastructure, and restored the affected service component to a healthy state. Together, these actions helped drain the backlog and restore normal operations.
We are working to improve resource saturation detection and to eliminate customer impact in this scenario by isolating impact, placing better bounds on retries, and strengthening backpressure to make our systems more resilient under load.
Disruption with GitHub Billing
Started August 26, 2026 at 8:40 PM UTC · 4h 11m
IssuesMinor incident
investigating
We are investigating reports of impacted performance for some GitHub services.
investigating
We are currently investigating increased errors with billing services. Customers may observe failed billing budget page loads, and users of the Copilot CLI may observe failures starting or continuing sessions.
investigating
We've applied a mitigation to unblock Copilot usage and have observed recovery for this particular impact. We're continuing to investigate and apply mitigations for the billing page disruption while monitoring to ensure Copilot remains recovered.
investigating
We are continuing to monitor the mitigation that we have applied for the billing page disruption.
investigating
Our mitigation is still holding as we continue to investigate to find the root cause.
investigating
Our mitigation continues to hold, and service conditions remain stable. We are continuing to investigate the concentrated workload responsible for the issue and are preparing additional preventative improvements. We have not identified a material change in customer impact since the previous update. We will provide another update as the investigation progresses.
investigating
No material change since the previous update. Service conditions remain stable following the mitigation, and we have not observed any further customer impact. We are actively monitoring the service while implementing targeted fixes to address the underlying root cause.
resolved
On August 26, 2026, between 20:40 UTC and 00:51 UTC on August 27, GitHub Billing experienced degraded performance affecting billing budget pages and GitHub Copilot CLI sessions. Affected customers encountered failed budget page loads or failures when starting or continuing CLI sessions. We confirmed this impact for a small number of customers (<1%).
This was caused by a concentrated workload that created processing delays in our data storage layer. Automated retries increased the load and prolonged the degradation. We mitigated the incident by rebalancing traffic within our infrastructure.
We are improving workload isolation, retry behavior, and detection of concentrated load to reduce the likelihood of recurrence and shorten our time to detect and mitigate similar incidents.
Incident with Actions
Started August 26, 2026 at 3:11 PM UTC · 2h 50m
OutageCritical incident
Affected components
ActionsPages
investigating
We are investigating reports of degraded availability for Actions
investigating
Pages is experiencing degraded performance. We are continuing to investigate.
investigating
We've identified an issue with a database primary and are failing over to a replica immediately
investigating
primary failover briefly improved performance but did not fully mitigate, we've throttled inbound traffic and are investigating upstream Vitess issues
investigating
We believe we've identified and addressed the issue and are ramping traffic back up slowly to ensure it doesn't recur. Some customers will continue to see delays as we ramp up.
investigating
Pages is operating normally.
investigating
We are continuing to observe recovery and delayed queues are burning down. Some customers will continue to see increased delays until all throttled work has been completed - we expect this within the next hour.
investigating
We are continuing to observe recovery and expect actions inbound queues to be back to normal in <30min. Work will continue to flow through the system subject to per-customer concurrency limits.
monitoring
The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.
monitoring
All inbound queues have recovered and Actions is operating as expected. 3.7% of jobs assigned to larger runners during the early stage of this incident are stuck waiting for runner assignment. Those will be canceled within the hour. Other runners are successfully processing all new jobs.
resolved
On August 26, 2026 from 15:02 to 15:45 UTC, Actions jobs failed to start. The following 2 hours until 17:40 UTC, Actions runs were delayed starting by more than 5 minutes as the system caught up with delayed load. This impact was triggered by saturation of writes to the database primary used by the service processing triggers for Actions workflows. The primary was failed over, but the system did not fully recover. The saturation was caused by growing daily peak load combined with an upstream issue in GitHub’s event processing infrastructure, https://www.githubstatus.com/incidents/hcbtzksccj2f, which caused burst amplification of already-high load. Downstream throttles that were later used to recover were set ~10% too high to protect the system.
At 15:45 UTC, throttling combined with service restarts recovered the service’s core health. Those throttles were gradually raised between 15:54 and 17:22 to restore full webhook processing for Actions runs. This ramp was deliberately slow to ensure we did not re-overwhelm the system given our original throttling was now known to be incorrectly set. The queue of webhook events was fully burned down at 17:40 UTC.
3.7% of larger-runner jobs, along with some scale-set self-hosted jobs, remained stuck in queued or “waiting for runner” state. We deployed a change to force-revoke jobs in this state, and they transitioned to failed at 18:40 UTC, about 50 minutes after incident mitigation. Releasing these jobs also freed hosted concurrency for larger-runner jobs.
Customers using concurrency groups saw longer impact due to a separate issue where runners assigned to a subset of jobs disconnected before the force-revoke mitigation was deployed, which prevented runner acquisition from progressing and left jobs in a waiting-for-runner state. This was resolved at 01:00 UTC on August 27.
Some runs triggered during the 15:02-15:45 UTC incident window encountered a bug that left them showing as queued even after service recovery. In the backend, these runs had already failed and will automatically move to canceled state 24 hours after creation. As follow-up, we are fixing the root cause of this queued state and improving our ability to bulk-cancel affected runs.
Several changes to improve the general scalability of this part of Actions were already complete and deploying to production. Rollout of those changes will be complete within the next 24 hours. Further work to improve scale, resiliency, and more graceful degradation of Actions workflows are in flight. We are also taking a repair item to accelerate clearing of stuck queued or waiting jobs in similar future cases.
Disruption with some GitHub services
Started August 26, 2026 at 3:09 PM UTC · 59m
IssuesMinor incident
investigating
We are investigating reports of impacted performance for some GitHub services.
resolved
Please refer to the combined summary in this related incident: https://www.githubstatus.com/incidents/y1t7p9fzrlj2
Actions delays in starting runs
Started August 24, 2026 at 1:33 PM UTC · 31m
IssuesMinor incident
Affected components
Actions
investigating
We are investigating reports of degraded performance for Actions
investigating
Failures while queuing and running Actions jobs for a subset of customers are now resolving. We are monitoring for full recovery.
monitoring
The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.
resolved
On August 24, 2026, between 13:33 UTC and 14:04 UTC, 3.8% of Actions runs experienced start delays over 5 minutes with 1.25% of Actions runs failing outright.
The incident was caused by a disk failure on a node hosting one of many service instances responsible for processing runner assignment events. Typically, pods on unhealthy nodes are removed and replaced automatically without impact. In this case, although the node was severely degraded and unable to perform disk operations, it continued sending healthy signals, preventing the system from immediately moving its work elsewhere. During this period, events assigned to the affected component accumulated until an automatic rebalance redirected processing to healthy components at 13:54 UTC. The queue backlog was cleared at 14:00 UTC, and processing returned to normal by 14:04 UTC.
To prevent a recurrence, we are improving detection and automated remediation for unhealthy nodes that aren’t fully offline. We are also strengthening application-level resiliency, so stalled consumers are automatically removed quickly and their work reassigned without waiting for the affected node to recover.
Elevated errors on Fable 5 due to upstream provider
Started August 24, 2026 at 7:12 AM UTC · 47m
OutageMajor incident
Affected components
Copilot AI Model Providers
investigating
We are investigating reports of degraded availability for Copilot AI Model Providers
investigating
We are experiencing degraded availability for the Fable model in Copilot products and IDE surfaces. This is due to an issue with the upstream model provider. While we work with them to resolve the issue, we recommend choosing another model or selecting 'Auto' to continue using Copilot.
resolved
On August 24th, 2026, between approximately 06:35 and 07:25 UTC, the Copilot service experienced a degradation of the Claude Fable 5 model due to an issue with our upstream provider. Users encountered elevated error rates when using Claude Fable 5, with requests sometimes failing mid-response. No other models were impacted.
The issue was resolved by a mitigation put in place by our provider. GitHub is working with our provider to further improve the resiliency of the service to prevent similar incidents in the future.
Degraded Git Operations over SSH
Started August 21, 2026 at 2:00 PM UTC · 7m
Pending
resolved
On August 21, 2026, between 14:00 and 14:07 UTC, dotcom Git operations over SSH were degraded. Successful Git operations over SSH fell by more than 95% for during the peak impact window, making clone, fetch, or push over SSH effectively unavailable to most users for approximately four minutes. Git operations over HTTPS were not affected.
The incident was caused by a software defect in our load-balancing infrastructure that was triggered by a configuration change. The defect only occurred when connections passed through multiple layers of load balancers running the new configuration, which meant it was not detected during canary testing.
We mitigated the incident by rolling back the configuration change.
We are adding regression coverage for multi-layer load-balancer configurations and improving monitoring and alerting for Git operations over SSH to reduce our time to detection and mitigation of similar issues in the future.
Intermittent failures creating agent tasks
Started August 20, 2026 at 2:43 PM UTC · 9h 54m
OutageCritical incident
investigating
We are investigating reports of impacted performance for some GitHub services.
investigating
Users may experience delays when starting tasks using Copilot Cloud Agent. We are actively investigating the issue and will provide updates as we learn more.
investigating
We have identified the problematic component and are working to fail over to a healthy instance. Further updates will be provided as we perform mitigations.
investigating
We are experiencing issues with Copilot Cloud Agent tasks, resulting in newly started tasks not properly displaying on-going progress. These Copilot Cloud Agent tasks are still being completed correctly but lack proper visibility. We are actively investigating the issue and will provide updates as we learn more.
investigating
Users are experiencing delays when starting tasks using Copilot Cloud Agent and are not be able to see the status of these tasks. Copilot Cloud Agent tasks are still being completed. We have identified the cause of the issue and are putting mitigations in place to return service to normal levels. We will provide another update about the expected recovery time shortly.
investigating
We are seeing signs of recovery for Copilot Cloud Agent task status visibility, but this recovery is slower than anticipated. We are pursuing additional mitigating measures to accelerate recovery.
investigating
We are observing gradual recovery for Copilot Cloud Agent task status visibility, with session output delayed approximately 1 hour. We have taken additional steps to accelerate the recovery and are continuing to monitor the impact.
investigating
We are continuing to observe gradual recovery for Copilot Cloud Agent task status visibility, with session output delayed by approximately 1 hour. We have taken additional steps to accelerate the recovery and are continuing to monitor the impact.
investigating
We are continuing to observe gradual recovery for Copilot Cloud Agent task status visibility, with session output delayed by approximately 1 hour. We have taken additional steps to accelerate the recovery and expect this to take effect within the next hour.
investigating
We are continuing to observe gradual recovery for Copilot Cloud Agent task status visibility. Session output continues to be delayed by approximately 1 hour as our remediation steps take effect.
investigating
We are seeing gradual recovery in Copilot Cloud Agent task status visibility as we deploy a fix for the root cause. Session output remains delayed by approximately one hour while remediation continues.
resolved
Between 13:57 UTC on August 20 and 00:37 UTC on August 21, 2026, some users of the Copilot Cloud Agent experienced delays of up to 60 to 90 minutes in seeing the status and results of their agent tasks. The agent tasks themselves continued to run and complete during this time; only the visibility of their status was delayed.
The cause was a regional outage in a third-party cloud database service that Copilot uses to store agent task status. We failed over the affected database to a healthy region, added processing capacity to work through the backlog, and restored normal operation once the underlying service recovered. No task data was lost during the incident.
To prevent repetition of similar incidents, we are removing the database configuration that made us vulnerable to this regional outage and improving our database failover procedures.
Intermittent failures in runner group and runner-related permissions pages
Started August 18, 2026 at 5:02 AM UTC · 6h 28m
IssuesMinor incident
investigating
We are investigating reports of impacted performance for some GitHub services.
monitoring
We are investigating reports of failure to load runner groups and runner-related permissions for customers using larger runners.
monitoring
We have identified the source of a communication issue between Actions services and are working toward mitigation. Customers may experience failure to load runner groups and runner-related permissions issues when using Larger Runners.
monitoring
We have applied a mitigation and are seeing recovery signals. We will continue monitoring recovery and providing updates.
resolved
On August 18, 2026, between 05:02 UTC and 11:30 UTC, customers were unable to view or manage Actions Runners and Runner Groups through the GitHub UI and API.
The issue was caused by failures in backend requests reading runner and runner group data. The failures were caused by an expired authentication certificate unique to this service. The certificate had been rotated in KeyVault, but a step to enable use at runtime had been paused to prevent recurrence of previous incidents triggered by this operation.
The impact was mitigated by completing the enablement of the new certificate in the backend system. We have added additional monitoring to this and other certificates. This service is also in the process of being replaced as part of our availability and scale work, bringing this authentication path and secret management in line with patterns across all GitHub services.
Incident with Actions
Started August 18, 2026 at 5:02 AM UTC · 6h 28m
OutageMajor incident
investigating
We are investigating reports of impacted performance for some GitHub services.
resolved
On August 18, 2026, between 05:02 UTC and 11:30 UTC, customers were unable to run jobs on Actions Larger Runners and were unable to view or manage Actions Runners and Runner Groups through the GitHub UI and API.
These issues were caused by failures in backend requests resolving essential metadata for starting Larger Runner workflow runs and for reading runner and runner group data. The failures were caused by an expired authentication certificate unique to this service. The certificate had been rotated in KeyVault, but a step to enable use at runtime had been paused to prevent recurrence of previous incidents that had been triggered by this operation.
We mitigated the issues by completing the enablement of the new certificate in the backend system. We have added additional monitoring to this and other certificates. The relevant service is also in the process of being replaced as part of our availability and scale work, bringing this authentication path and secret management in line with patterns across all GitHub services.
We are investigating reports of impacted performance for some GitHub services.
investigating
API Requests is experiencing degraded performance. We are continuing to investigate.
investigating
Actions is experiencing degraded performance. We are continuing to investigate.
investigating
We are seeing an approximate 20% error rate across numerous experiences including Pull Requests, Issues, and others. Investigations are currently under way and we will be posting updates as they become available
investigating
Pull Requests is experiencing degraded performance. We are continuing to investigate.
investigating
We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. Investigations are on-going into the root cause, and updates will continue to be provided as we investigate.
investigating
We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. Investigations are on-going and we will continue to provide updates as we discover more information.
investigating
Copilot is experiencing degraded availability. We are continuing to investigate.
investigating
Issues is experiencing degraded availability. We are continuing to investigate.
investigating
Pull Requests is experiencing degraded availability. We are continuing to investigate.
investigating
Actions is experiencing degraded availability. We are continuing to investigate.
investigating
We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. We are currently performing mitigations based on our investigation thus far and are monitoring for improvement.
investigating
Webhooks is experiencing degraded availability. We are continuing to investigate.
investigating
Pages is experiencing degraded performance. We are continuing to investigate.
investigating
Webhooks is experiencing degraded performance. We are continuing to investigate.
investigating
We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. We are currently performing mitigations and will post updates as we progress.
investigating
We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. We are still working to identify the root cause and will continue to post updates as we learn more and perform mitigation.
investigating
We identified the problematic component and have taken corrective actions. There are strong signs of recovery but we are still working to completely restore service, with error rates still remaining slightly elevated. We will post further updates as recovery continues.
investigating
The degradation affecting API Requests, Actions, Git Operations, Issues, Pages, Pull Requests and Webhooks has been mitigated. We are monitoring to ensure stability.
investigating
Git Operations is experiencing degraded performance. We are continuing to investigate.
investigating
We identified the problematic component and have taken corrective actions, but we are seeing residual impact across numerous services. We are continuing to apply additional mitigations and investigate the remaining impact.
investigating
Issues is experiencing degraded performance. We are continuing to investigate.
investigating
We identified the problematic component and have taken corrective actions, but we are seeing residual impact in the form of sporadic authentication failures. We are continuing to apply additional mitigations and investigate the remaining impact.
investigating
The degradation affecting Git Operations has been mitigated. We are monitoring to ensure stability.
investigating
API Requests is experiencing degraded availability. We are continuing to investigate.
investigating
API Requests is operating normally.
investigating
We are continuing to investigate sporadic authentication failures. We have partially disabled authentication token retries and have seen improvement, and we are monitoring impact before fully applying this mitigation.
investigating
We are continuing to investigate sporadic failures affecting Copilot authentication in some applications. Copilot usage via the GitHub CLI and GitHub App are unaffected.
investigating
Issues is operating normally.
investigating
We are continuing to apply mitigations to address sporadic Copilot authentication failures in some applications. We expect full recovery within the next 30 minutes. Copilot usage via the GitHub CLI and GitHub App are unaffected.
resolved
On August 17, 2026, from 13:28–21:15 UTC (7h 47m), GitHub.com experienced elevated errors and latency across Issues, Pull Requests, APIs, Actions, and Copilot. At peak, web/API error rates were approximately 20%, while archive and raw-content downloads reached approximately 50%. SAML/OIDC authentication, SCIM, and Team Sync were also affected, as well as Actions workflows in GHEC with Data Residency that depend on public workflow step definitions hosted on GitHub.com. Most services recovered by 16:36 UTC as our Central US datacenter recovered; Actions was degraded until approximately 18:03 UTC; and Copilot Token Service fully recovered by 21:02.
Some of the failing traffic was moved from Central US to Northern Virginia where it was served successfully until the network failure in Central US was debugged and resolved. Delayed replies to a single internal endpoint triggered a latent retry bug in VS Code that amplified traffic by approximately 10x and caused delayed recovery for the Copilot Token Service.
The immediate cause of the failure was network saturation on load balancers in Central US due to a new peak in traffic. Originally this was caused by an Istio sidecar pod reaching its concurrency limits and failing to auto scale correctly because of a misconfigured policy that watched host service but not sidecar limits. One failure cascaded to more and eventually four HAProxy nodes exhausted their flow limits, degrading the gateway auth path and causing widespread authentication latency and failures. The problem was worsened by optimistic retry logic which overloaded internal load balancers. Pausing HAProxy on those nodes simultaneously produced immediate broad recovery.
The retry storm in Northern VA was fixed by 1) temporarily reducing gateway retry logic with a PR and 2) blocking inbound Copilot Token Service token requests at the load balancers with a 403, and then gradually ramping back up traffic per-site to allow callers to succeed.
Residual Copilot authentication failures continued because client retry behavior amplified load: a failed token operation could generate many extra requests and enter a retry loop. Copilot Token Service traffic increased from a normal 7–9K RPS to 70–100K RPS. Reducing gateway authentication retries and blocking retry-triggering responses stabilized Copilot Token Service and completed recovery.
Complicating factors that impeded recovery included a number of scraping attacks on codeload endpoints.
To prevent recurrence, our follow-up actions include:
- Correcting autoscaling policies to account for service-mesh sidecar concurrency and capacity.
- Auditing Istio request, concurrency, and scaling limits across affected services.
- Reviewing retry limits and backoff behavior across gateways and clients.
- Addressing the VS Code retry behavior that amplified Copilot token traffic.
- Improving load-balancer capacity monitoring and regional failover safeguards.
Disruption with GHEC Team Sync
Started August 13, 2026 at 4:21 PM UTC · 2h 6m
IssuesMinor incident
investigating
We are investigating reports of impacted performance for some GitHub services.
investigating
GHEC Team Sync is currently degraded for enterprises with personal accounts, causing delays when syncing teams to IdP groups. We have identified the cause of the delays and are working on a mitigation. We will provide an update on our progress at 20:00 UTC.
investigating
We have deployed a mitigation. At this time GHEC Team Sync has recovered for enterprises with personal accounts. Teams syncing to IdP groups have returned to their normal cadence.
resolved
On August 13, 2026, from 15:31:21 UTC to 18:27:55 UTC, GitHub Enterprise Cloud team synchronization was degraded for enterprises using personal accounts. Organization teams experienced delays of up to 3 to 13 hours (median 8 hours) when syncing with IdP groups, resulting in delayed access grants or removals for enterprise users across 2.8% of teams.
A temporary change introduced to address a previous issue due to increased usage of this feature remained active after it was intended to be removed, causing synchronization delays during periods of high volume. We removed the temporary change and provisioned additional resources to handle the increased volume.
Incident with Webhooks
Started August 13, 2026 at 2:45 PM UTC · 51m
IssuesMinor incident
Affected components
IssuesPackagesPull RequestsWebhooksGit Operations
investigating
We are investigating reports of degraded performance for Webhooks
investigating
Pull Requests is experiencing degraded performance. We are continuing to investigate.
investigating
Issues is experiencing degraded performance. We are continuing to investigate.
investigating
Git Operations is experiencing degraded performance. We are continuing to investigate.
investigating
Packages is experiencing degraded performance. We are continuing to investigate.
investigating
We are currently investigating a brief degradation of service for Git operations (specifically pushes), issues, pull requests, package registry, and webhooks between 14:32 and 14:46 UTC. We have identified the source of the degradation and are investigating mitigation strategies to prevent recurrence.
monitoring
The degradation affecting Git Operations, Issues, Packages, Pull Requests and Webhooks has been mitigated. We are monitoring to ensure stability.
monitoring
We have temporarily disabled a background job which caused the impact. At this time the impact is fully mitigated.
resolved
Between 14:24 and 14:53 UTC on 13 August 2026, a routine background job to delete an organization overwhelmed a key shared database, causing multiple GitHub services to briefly return elevated errors and slower responses. Most affected was the webhook management API, with smaller impact to Git operations, pull requests, issues, packages, sign-in, and Copilot. Impact cleared on its own at about 14:53 UTC once the job finished; we resolved the incident at 15:36 UTC.
Affected users may have experienced a brief increase in errors and slower responses, primarily when creating, listing, or updating webhooks, with smaller impacts to pull requests, issues, packages, and Git operations. Failures peaked at about 1% for several minutes around 14:37 UTC.
To prevent future incidents, we've already shipped an update that turns on the safer deletion path for organizations, along with caps on deletion holds on databases. Building on these changes, we're auditing all bulk deletion and cleanup jobs that write to shared databases to prevent similar issues in future.
Errors with the Fable 5 Model in Copilot
Started August 13, 2026 at 2:43 PM UTC · 1h 4m
IssuesMinor incident
Affected components
Copilot AI Model Providers
investigating
We are investigating reports of degraded performance for Copilot AI Model Providers
investigating
We are experiencing degraded availability for the Fable 5 model in Copilot products and IDE surfaces. This is due to an issue with an upstream model provider. While we work with them to resolve the issue, we recommend choosing another model or selecting 'Auto' to continue using Copilot.
investigating
We are seeing modest recovery, but are still experiencing degraded availability for the Fable 5 model in Copilot products and IDE surfaces. This is due to an issue with an upstream model provider. While we work with them to resolve the issue, we recommend choosing another model or selecting 'Auto' to continue using Copilot.
investigating
The issues with our upstream model provider have been resolved, and Fable 5 is once again available in Copilot products and IDE surfaces.
We will continue monitoring to ensure stability, but mitigation is complete.
resolved
On August 13th, 2026, between approximately 14:06 and 15:47 UTC, the Copilot service experienced a degradation of the Claude Fable 5 model due to an issue with our upstream provider. Users encountered elevated error rates, peaking at 43% and averaging 12%. Users who selected Auto or alternative models were unaffected.
The issue was resolved by a mitigation put in place by our provider. GitHub is working with our provider to further improve the resiliency of the service to prevent similar incidents in the future.
GitHub outage history and incident timeline | Uptimus