Bots/crawlers are overwhelming the service. Adjusting throttling.
monitoring
Throttling adjustments are in place. Errors subsiding. Monitoring.
resolved
Errors have returned to normal levels. Fix confirmed. Resolving.
Investigating elevated repo errors
開始 2026年5月20日 13:20 UTC · 2h 35m
Issues軽微なインシデント
影響を受けたコンポーネント
Repo endpoints
investigating
We're seeing unusual traffic to our repositories that are causing elevated errors and timeouts. We've investigating possible ways to mitigate this.
monitoring
The number of open connections to each server was reaching our configured limit. We have raised the limit. Monitoring.
resolved
Configuration & scaling improvements seem to be absorbing the traffic spikes more comfortably now. Resolving.
Elevated errors for uploads and asynchronous work
開始 2026年3月25日 7:44 UTC · 47m
Issues軽微なインシデント
影響を受けたコンポーネント
Uploads
investigating
We are again seeing errors & delays for uploads. Investigating.
monitoring
Looks like some of the misconfiguration issues persisted from yesterday. A few of the queues were not processed, which prevented account updates from propagating to repositories. We have made a fix, and are now processing the backlog.
resolved
Jobs processing normally, and outstanding work has been completed. Resolving.
Job processing delays & errors
開始 2026年3月24日 16:34 UTC · 2h 6m
Issues軽微なインシデント
影響を受けたコンポーネント
Uploads
investigating
We are investigating delays in processing of uploads and other asynchronous work.
resolved
Production-specific misconfiguration caused our async workers to enter a crash-restart loop. We've fixed the issue, and we will be looking into how to catch these issues earlier in the release process. Resolving.
Elevated errors
開始 2026年3月16日 15:41 UTC · 1h 42m
Issues軽微なインシデント
影響を受けたコンポーネント
Git serverDashboardRepo endpointsAPI
investigating
We are currently investigating this issue.
monitoring
We had a temporary 10x traffic increase on one of our backend services, which lead to elevated timeouts and errors. Error rates have now returned to normal while we continue to look for the root cause of that traffic spike.
resolved
Resolving the incident to continue the investigation and fix offline.
Elevated errors
開始 2026年2月22日 4:59 UTC · 1h 52m
Issues軽微なインシデント
影響を受けたコンポーネント
Custom domainsRepo endpointsAPI
investigating
Investigating elevated error rates
monitoring
New code with an inefficient query in a critical path deployed. We have rolled it back - errors returning to normal.
resolved
Operations returned to normal. We will address the performance issue through our development cycle.
Elevated errors & timeouts
開始 2026年2月9日 11:21 UTC · 5h 0m
Issues軽微なインシデント
影響を受けたコンポーネント
DashboardUploadsCustom domainsRepo endpointsAPI
investigating
We are investigating elevated error and timeout rates.
monitoring
Unusual traffic spike seems to have temporarily overwhelmed one of our API endpoints. That spike has subsided, error/timeout rates are back to normal. We are monitoring while we continue to investigate.
resolved
We've tracked the issue down to a frequent DB update and applied a fix to reduce it. Resolving the incident while we will keep monitoring the fix.
AWS outage effects
開始 2025年10月20日 15:59 UTC · 12h 40m
Issues軽微なインシデント
影響を受けたコンポーネント
Git serverRepo endpointsAPI
identified
We are reopening the incident due new availability issues related to AWS outage.
monitoring
We are continuing to monitor the upstream issue.
resolved
Upstream issue resolved. If you’re still having availability issues, don’t hesitate to contact us directly.
postmortem
**Summary**
First, we want to thank everyone for your patience as we worked through this incident. After restoring platform availability, we noticed that some services \(e.g. uploads\) did not immediately recover. This postmortem explains what happened, what we found, and what we’re doing to prevent similar issues in the future.
**What Happened**
The platform experienced a temporary availability outage that affected several services. While core systems came back online as expected, a subset of services continued to experience availability issues after the upstream outage.
**Root Cause**
The affected services were still using stale DNS cache entries from during the outage, preventing them from reaching other internal components. While DNS caching is technically correct, our monitoring should have detected the failures, and prompted corrective action. However, our existing checks were only performing shallow availability tests and did not verify internal reachability, leading to a false positive.
**Resolution**
Once identified, clearing the DNS caches restored full connectivity and service availability. To prevent recurrence, we’ve extended our monitoring to include internal reachability probes, ensuring that similar issues are surfaced automatically in the future.
**Closing**
We understand how disruptive downtime can be, and we’re committed to learning from each incident to make our platform more resilient. Thank you again for your patience and continued trust.
AWS outage effects
開始 2025年10月20日 10:15 UTC · 26m
Pending
影響を受けたコンポーネント
Git server
monitoring
As many others, we have been affected by the AWS outage. Particularly, our Git/Build service could not accept pushes and trigger new builds. Currently, upstream errors and our errors have subsided, we are monitoring toward a resolution.
Note: We were unable to update this status page because this service was also affected.
resolved
Upstream issue has been resolved. We are back to normal. Resolving.
Elevated error rates
開始 2025年9月18日 15:37 UTC · 1h 51m
Issues軽微なインシデント
影響を受けたコンポーネント
DashboardCustom domainsRepo endpoints
investigating
We are seeing elevated error rates. Investigating.
monitoring
We've found that the issue was triggered by a recent configuration update. We've rolled back the changes, and the error rates have returned to normal levels. Checking further into the changes that caused the failure.
resolved
It appears a misplaced space in a configuration update of caused the configured library to throw exceptions for some of the requests. We've fixed the configuration and will be patching the library to ensure this is handled more gracefully. Resolving.
Timeouts for proxied npm registry
開始 2025年6月12日 19:20 UTC · 1h 40m
Issues軽微なインシデント
影響を受けたコンポーネント
Repo endpoints
identified
We are seeing increased timeouts for proxied requests to the public npm registry. We are tracking the upstream issue with npmjs.org.
resolved
Upstream incident has been resolved, and no problems otherwise. Resolving.
Platform issues causing upload failures
開始 2025年6月10日 12:51 UTC · 14h 22m
Issues軽微なインシデント
影響を受けたコンポーネント
Uploads
investigating
We are currently investigating this issue.
investigating
Our asynchronous workers are not running due to platform issues, and thus cannot process uploads. We are waiting for updates from our platform provider.
monitoring
We were able to start some of the workers. Starting to work through the async task backlog.
monitoring
Although some platform instability persists, we have processed most of the asynchronous task backlog (uploads, emails, etc.). Most functionality has been restored. We will continue monitoring the status of the platform until resolution. If you're still experiencing delays or issues with your account, please contact [email protected] with full details.
monitoring
We are continuing to monitor for any further issues.
resolved
We are resolving this as we are no longer seeing customer-facing errors. We'll continue to monitor both our and the platform's status to address any issues that may arise.
Elevated errors after deployment
開始 2024年11月20日 6:48 UTC · 1h 36m
Issues軽微なインシデント
影響を受けたコンポーネント
DashboardRepo endpointsAPI
monitoring
We've deployed an update that caused elevated errors due to a caching error that didn't surface during testing. We've rolled back the update, implemented a fix, and rolled out the fixed update. Monitoring.
resolved
The fix was effective. Service fully restored.
Custom domains and uploads outage
開始 2024年8月17日 8:59 UTC · 43m
Outage重大なインシデント
影響を受けたコンポーネント
UploadsCustom domains
investigating
Investigating custom domains and uploads outage.
monitoring
A routine infrastructure update misconfigured our load balancers. We've updated the load balancers and updated DNS settings. It may take time for changes to propagate to clients. Monitoring.
resolved
Request and error rates returning to normal levels. We will continue to monitor. Resolving.
Build failures
開始 2024年7月6日 14:12 UTC · 45m
Outage重大なインシデント
影響を受けたコンポーネント
Git server
investigating
Investigating build failures
monitoring
Shortage of resources prevented build jobs from being scheduled. Scaling up nodes seems to have fixed the issue. Monitoring.
resolved
Back to normal. Resolving.
Failing builds via Git
開始 2024年4月20日 15:32 UTC · 1h 29m
Outage重大なインシデント
影響を受けたコンポーネント
Git server
investigating
We are investigating an issue with failing builds via `git push`
resolved
Clock drift resulted in authentication failure between Git server and builder nodes. Correcting the time resolved the sporadic failures. We will investigate why nodes don't automatically synchronize their clocks. Resolving the immediate issue.
Git-push build errors
開始 2023年9月15日 19:34 UTC · 14m
Outage重大なインシデント
影響を受けたコンポーネント
Git server
investigating
Investigating "git push" build issues
identified
Payment issue with our cluster provider caused temporary API suspension. We've fixed the payment, and waiting for restoration of cluster access.
resolved
Cluster access restored. Git repo building issues have been resolved.
Elevated error rates
開始 2023年8月19日 0:34 UTC · 17m
Issues軽微なインシデント
影響を受けたコンポーネント
UploadsCustom domainsRepo endpointsAPI
investigating
We are investigating elevated error rates
monitoring
A cache instance has failed triggering automatic failover. Error rates returning to normal.
resolved
Everything returned to operating normally. Resolving.
Elevated error responses
開始 2023年8月8日 4:57 UTC · 5h 10m
Issues軽微なインシデント
影響を受けたコンポーネント
Repo endpointsAPI
investigating
Investigating latency & error spike
monitoring
We've rolled back most recent deployment. Still debugging.
resolved
We've found a caching bug that was introduced by the latest deployment. It will not be present in future builds. Resolving.
Partial service failure due to non-volatile Redis saturation
開始 2023年6月29日 16:00 UTC · 0m
Pending
resolved
Earlier today, we deployed a bug that exposed a legacy system to excessive traffic, overwhelming our core Redis instance and causing cascading failures in certain functionality. In the past, we occasionally used a legacy internal system to capture debug information for rare data states, aiding in tracking and reproducing customer issues. This information was stored in our core non-volatile Redis instance. While this approach had worked for rare conditions and low-traffic code paths, this incident occurred due to mistakenly adding such tracking to heavily-trafficked functionality.
May 29th, 12:43 UTC - Deployment of a release with the tracking bug
The release containing the tracking bug passed preflight and was deployed to production. Initially, everything appeared stable. However, the utilization of our non-volatile Redis, which usually hovers below 5%, slowly started to increase. Unfortunately, this increase went unnoticed.
May 29th, 15:55 UTC - Redis storage reached maximum utilization
When the non-volatile Redis storage reached 100% utilization, write operations began receiving "OOM command not allowed" error responses, resulting in 500 errors for certain user-facing APIs. Most read operations were successful, but not all. Regrettably, the tracking code was present in the API layer servicing the Dashboard and the CLI, causing errors for those read operations. Worse, the error rate remained low enough to not trigger any alarms.
May 29th, 23:04 UTC - Redis instance cleaned to restore service
Customer service noticed an increase in error reports and promptly notified engineering to investigate. Engineering quickly identified the issue and cleared excess data from Redis, restoring service.
May 29th, 23:40 UTC - Fix deployed to remove the tracking bug
We deployed a fix to remove the tracking bug.
Further steps
Later in the day, we removed the legacy tracking system and migrated that functionality to use standard error and metrics tracking. This step will prevent similar space issues in the future and consolidate our monitoring infrastructure. Moving forward, we will implement more alarms for storage utilization and introduce more fine-grained tracking of error rates.