GitHub incident can affect Cycode
- investigating
GitHub component: Pull Requests Original GitHub incident: https://stspg.io/ssd9z8l2g46v
- resolved
GitHub component: Pull Requests Original GitHub incident: https://stspg.io/ssd9z8l2g46v
46 Cycode incidents · 2026年2月 — official updates, affected components, duration and resolution details.
GitHub component: Pull Requests Original GitHub incident: https://stspg.io/ssd9z8l2g46v
GitHub component: Pull Requests Original GitHub incident: https://stspg.io/ssd9z8l2g46v
GitHub component: API Requests Original GitHub incident: https://stspg.io/3xn46bst0bjh
GitHub component: API Requests Original GitHub incident: https://stspg.io/3xn46bst0bjh
GitHub component: API Requests Original GitHub incident: https://stspg.io/3xn46bst0bjh
Customers may experience degraded performance in scans. Pull request and CLI scans may be affected.
The team has identified the root cause of the issue and is working on the solution.
The root cause has been resolved. The system began to stabilize itself, and all the scans are starting to get processed with regular performance.
All scan types except for SAST are fully operational. SAST continues to stabilize and will soon be fully stable.
System should be back to being fully operational.
**Root Cause** The incident was caused by a deployment of one service that ran a database index creation. The team has identified an issue with the way we perform index creations as database migrations. During a deployment the pods with newest image of the service attempted to create an index on a big table. The team has identified that the index creation took 7 minutes. However, during index creation pods were not responsive, and as a result, Kubernetes deemed them as unhealthy pods and attempted to retry those pods after 5 minutes. As a result, because the pod got killed before the index creation was fully completed, the database transaction was rolled back. Then, subsequent pods attempted to create the index again, dying after 5 minutes. This lead to the database being in unhealthy state, and the service was down. The team has rolled back the deployment, and killed all replicas that attempted to create the index. Thanks to that, the service and the database was in healthy state again. **Why safety measures did not help** Cycode provides a safety mechanism that unblocks all Pull Request scans after a specific period of time, giving each scan a maximum duration before the Pull Request is unblocked. However, because the service that is responsible for triggering and completing scans, as well as this safety net, was down, the process couldn't behave as expected. We acknowledge this gap and are working on strengthening this area of our system. **Action items** • The team is actively investigating enhancements and new safety protocols that can be put in place in order to have another safety net preventing Pull Request scans being stuck in case of any incident. • The team is investigating changes to the index creation process.
我们正在调查一个导致旧事件被重新处理的问题。 ** 影响:** 一些工作流程可能再次运行,这可能导致重复的提醒. PR扫描也可能被延迟.
我们解决了卡夫卡移民期间一个问题造成的积压案件。 这个问题导致旧事件与新事件一起处理,造成拖延,并可能使一些违规行为更新过时。 正常处理已恢复 。 然而,在查明和纠正受影响记录的同时,客户可能仍然看到一些已过时的违规行为。 PR扫描没有被延迟或再处理,工作流程也没有受到影响. 我们正在继续补救并密切监测该系统.
正常的处理已经恢复,事故正在监测中。 少数顾客可能仍然看到因事件而导致数量有限的侵权事件,其地位已经过时. 我们查明了可能受影响的环境,并正在努力纠正受影响的记录.
该系统已全面运作。 少数顾客可能仍然看到因事件而导致数量有限的侵权事件,其地位已经过时. 我们查明了可能受影响的环境,并正在努力纠正受影响的记录.
The platform is now fully operational and processing normally
自动翻译自官方事件更新。
小组通过扫描发现了一个性能退化的问题。 启动、运行和完成扫描可能会有延误。 所有扫描类型都可能会受到影响(Pull request和CLI扫描也会受到影响). 小组已查明一个部署有故障的问题,这一问题应很快解决.
小组查明了根源并解决了这一问题。 我们正在看到系统恢复稳定.
该系统现已恢复全面运作.
** 循环原因** 造成这一事件的原因是部署了一个管理数据库迁移的服务。 小组确认,这种迁移含有错误的代码,因此在试图部署服务时导致数据库超载。 因此,在恢复部署之前,这项服务部分减少。 在此期间,所有扫描处理的性能均低于预期.
自动翻译自官方事件更新。
GitHub 组件: API 请求 原"吉打"事件:https://stspg.io/vr201n49yl53
GitHub 组件: API 请求 原"吉打"事件:https://stspg.io/vr201n49yl53
自动翻译自官方事件更新。
GitHub 组件: 拉请求 GitHub事件原文:https://stspg.io/sm1tp7kfm4vj
GitHub 组件: 拉请求 GitHub事件原文:https://stspg.io/sm1tp7kfm4vj
GitHub 组件: 拉请求 GitHub事件原文:https://stspg.io/sm1tp7kfm4vj
自动翻译自官方事件更新。
我们注意到IaC Pull Request扫描显示性能下降 我们正在解决问题.
我们查明了缓慢的根源,并正在采取对策.
导致缓慢的后退几乎已经过去。 我们正在监测局势.
这个问题已经解决,我们正在继续监测局势.
自动翻译自官方事件更新。
GitHub 组件: API 请求, 拉请求 (原始内容存档于2013-10-10). GitHub 原始事件:https://stspg.io/j5c80shqm53
GitHub component: API Requests, Pull Requests Original GitHub incident: https://stspg.io/j5c80shxqm53
自动翻译自官方事件更新。
GitHub 组件: Webhooks (原始内容存档于2017-03-27) (英语). GitHub 原始事件:https://stspg.io/ syhr80rth84z
GitHub 组件: Webhooks, 拉请求 (原始内容存档于2017-03-27) (英语). GitHub 原始事件:https://stspg.io/ syhr80rth84z
GitHub 组件: Webhooks, 拉请求 (原始内容存档于2017-03-27) (英语). GitHub 原始事件:https://stspg.io/ syhr80rth84z
自动翻译自官方事件更新。
在1:41** PM EDT中,我们发现了一个在平台UI处理请求时导致错误率上升的问题. 这一问题得到迅速缓解,该平台目前正常运行。 我们的团队继续积极监测该平台,以确保服务稳定.
事件已经解决,站台正常运行.
自动翻译自官方事件更新。
我们已经注意到一部分CLI秘密扫描失败. 我们正在积极研究这一问题.
我们发现CLI版本3.17.1引入了错误行为. 将CLI版本降低到3.17.0,应该暂时解决这个问题,同时我们继续理解和解决根源.
已采用了固定装置并完全恢复了功能;我们正在继续监测,以确保一切保持稳定.
自动翻译自官方事件更新。
由于扫描负荷增加,我们看到违规状态更新出现延误,包括违规情况自动解决。 突然增加的原因已经得到缓解,拖延已经减少。 我们会一直监视局势 直到我们恢复正常
我们继续监控侦测程序 以及相关违规状态更新(包括自动解析)的延迟
自动翻译自官方事件更新。
我们已注意到多个系统组件的性能下降。 我们正在查明所有受影响的组成部分并查明根源.
我们注意到应用程序UI的多部分性能下降。 扫描以及Pull Request和CLI扫描也可能受到影响.
我们观测到Redis超时率上升。 我们正在扩大Redis集群的能力,以缓解影响并恢复稳定的服务业绩
该地区已经恢复并正常运转。 我们正在密切监测系统的业绩,以确保持续稳定
该系统现已全面运作。 不应再有退化的性能。 ** 概要** 我们观测到一段缓慢和间歇性超时的时期,影响到欧盟区域的各种系统功能,包括应用接口和拉取请求扫描. 造成这一问题的主要原因是处理系统达到其网络和内存容量极限,而单一来源的自动化活动量很大,更加剧了这一问题。 此后,我们更新了基本的基础设施,并实施了保障措施,以防止类似的高容量活动影响该系统。 这个问题现已得到全面解决,所有服务已恢复到预期的业绩水平。 ** 关键时间线(IDT)** • ** 2026年7月13日,11:44,IDT**: 在UI慢速和PR扫描延迟后检测出事件. • ** 2026年7月13日,12:19 IDT**: 确定了基础设施瓶颈;决定升级加工组群。 • ** 2026年7月13日,12:26,IDT**: 确定并停用高容量自动化工艺以减少即时负载. • ** 2026年7月13日,13:09,IDT**: 基础设施升级完成;网络吞吐量恢复到正常水平. • ** 2026年7月13日,15:35,IDT**: 所有积压案件得到清理,事件得到正式解决. ** 循环原因** 事件由多种因素共同引发:由于当前工作量配置不足,一个处理组达到了最大网络带宽和内存容量. 具体自动化工作流程产生了异常多的更新请求,使这项工作更加紧张。 此外,欧盟区域信息处理管道的配置差异使该系统无法有效处理由此产生的积压。 ** 采取的行动** • ** 基础设施升级**: 处理集群被升级为容量更高的实例类型,以提供更多的网络带宽和内存. • ** 残疾高卷来源**:对过度交通负责的特定客户身份识别器被暂时停用,以恢复系统稳定性。 • ** 恢复连接**:重新启动了受影响的服务部分,以确保它们与升级后的基础设施重新建立清洁连接。 • ** 增加处理并行性**: 受影响的信件队列中的分区数被增加,以便系统能更快地处理积压. ** 行动项目** • **加强监测**:实施新的网络和内存利用警报,在影响客户之前发现能力问题。 • ** 优化更新工作流程**: 将状态更新进程重置为批量请求,大幅降低处理系统负荷. • ** 限制税率**:采用保障措施,防止任何单一来源消耗不成比例的系统资源。 • 规范区域配置**: 进行一次审计,以确保所有区域的基础设施和信息排队设置一致.
自动翻译自官方事件更新。
我们已找出一个问题, 可能让某些产品仪表板和定制仪表板依赖违规数据(并非所有仪表板都受到影响), 我们已经开始了纠正工作, 但需要时间才能完全完成, 你可能会看到数据被更正后数的变化。 我们将会分享另一个更新.
我们在纠正影响某些产品和定制仪表板的不准确违规记录方面取得了重大进展。 • ** 现况:** 绝大多数账号的修补工作已顺利完成,并恢复了全部数据准确性. • ** 下一步:** 我们正在积极解决剩余的少数受影响账户的问题.
功能完全恢复;我们正在继续监测,以确保一切保持稳定.
自动翻译自官方事件更新。
目前,我们正在调查影响Maestro风险探索、风险AI补救、Maestro补救和图表AI服务的问题.
我们发现一个防火墙配置问题 影响了大师AI服务 已更新配置,受影响服务已恢复。 我们正在继续监测局势并努力进一步稳定局势。 在我们完成额外改进的同时,仍然可以看到一些退化的绩效.
影响大师AI服务的问题已经解决. 大师风险探索、风险AI补救、大师补救和图表AI现在已经可以使用并正常运行。 ** 概要** 2026年7月9日,在欧洲生产环境中使用Maestro服务的客户遭遇了一段时间的服务不便. 问题开始于一个配置更新,无意中改变了服务的区域路线. 这导致系统试图通过一个缺乏必要权限的网络路径连接到一个没有特定处理模式的区域。 这个问题已经完全解决,所有受影响的客户都恢复了服务。 ** 关键时间线(IDT)** • ** 2026年7月9日,12:02 IDT:** 查明了这一事件并启动了调查。 • ** 2026年7月9日,12:07 IDT:** 就服务中断发出公开通知。 • ** 2026年7月9日,下午1:00 采用了网络配置修复,恢复了初级连接。 • ** 2026年7月9日,13:39 IDT:** 实施模式倒计时后服务被全面恢复,事件被标为已解决. ** 循环原因** 服务中断是由最近更新认证和配置过程而引发的. 这一更新在系统如何确定操作区域方面引起了冲突。 具体来说,一个自动更新过程会推翻手动设置,路由流量到不同的区域终点. 这条新路径被缺失的网络安全规则所阻挡,并试图使用该特定区域不支持的处理模式,导致服务失败. ** 采取的行动** • ** 恢复网络连接:** 手动更新网络安全规则,允许通过新的区域终点安全流量. • ** 已实施模式后退:** 配置了该系统,以使用替代处理模式,确保立即提供服务,同时对长期区域配置进行调整。 • ** 最新情况来文:** 在整个恢复过程中为利益攸关方和客户保持实时更新。 ** 行动项目** • ** 设定优先配置标准:** 更新部署工作流程,防止自动化进程无声地压倒关键环境设置。 • ** 基础设施审计:** 对所有区域的网络安全规则进行全面审查,以确保一致性并预防类似的连接漏洞。 • ** 强化自动监测:** 实施端到端健康检查和合成探测器,在影响用户之前自动发现区域连接问题. • ** 改进部署政策:** 制定新准则,确保配置变化在类似生产的环境中得到更频繁的部署和验证,以降低"上市"更新的风险.
自动翻译自官方事件更新。
We are experiencing delays in infrastructure provisioning caused by cloud provider API rate limiting. We are actively investigating the issue with our cloud provider.
Please refer to the AWS Health Status page for details on the related incident: [https://health.aws.amazon.com/health/status](https://health.aws.amazon.com/health/status "https://health.aws.amazon.com/health/status")
Mitigation: We temporarily scaled up the managed node group to get pods scheduled while we wait for AWS to fully resolve the underlying issue.
We are starting to see stabilization and a reduction in API errors. However, we continue to closely monitor the situation.
AWS has confirmed that the issue has been fully mitigated and we are currently not observing any related issues.
** 调查 -- -- 与违规和定制板有关的问题(Prod-US)** 我们目前正在调查我们** Prod-US**环境中一个违反规定的情况。 因此,依赖违规数据的自定义仪表板面板也可能无法渲染出或显示出错. 我们的工程团队正在积极研究根源,我们将在学习更多知识时在这里提供最新情况。 我们为不方便而道歉.
已针对影响普罗德-美国的违规和定制仪表板的问题采取了补救措施。 我们正在积极监测环境,以确保全面恢复服务.
核心问题已经解决,违规和定制仪表板现在应该正常运行。 我们的小组正在积极监测数据同步情况,以解决与新违规行为相左的问题。 一旦同步完成,我们将提供最后更新.
我们正在继续监测UI中新违规行为的数据同步进程。 虽然功能已经恢复,但所有最新数据可能需要长达**6小时**才能完全赶上和准确反映。 一旦同步完成,我们将提供最后更新.
数据同步完成,近期所有违规事件均成功入驻UI. 违规和自定义仪表板正常运行,事件已完全解决. 我们感谢你在我们努力恢复全面服务时的耐心.
功能完全恢复;我们正在继续监测,以确保一切保持稳定.
自动翻译自官方事件更新。
We have identified the source of an issue and currently deploying the fix. At the same time we scaled our scanning platform up to accelerate scanning
The fix was deployed. The queue is decreasing and we're monitoring it
The system has processed all jobs with higher priorities. There is a queue of lower priority jobs that should not impact overall Cycode scanning performance
**Summary** During the incident, customers experienced significant delays and temporary disruptions across SAST, SCA, CCA, and Secret repository scans and push events. The issue was caused by a surge in reachability scanner jobs that overwhelmed the processing queue, compounded by scanner pods requesting excessive CPU and memory, infrastructure resource limits being reached, and inefficiencies in job prioritization and retry logic. As a result, processing capacity was improperly consumed and a large job backlog accumulated. A series of corrective updates were deployed to stabilize the environment, and the processing environment has since returned to expected performance levels. **Impact** Customers experienced delays for SAST, SCA, CCA, and Secret repository scans and push events, with some requests delayed by several hours and a peak queue size of over 64,000 jobs. Lower priority scans such as Trivy, Syft, and CCA were most affected, though high-priority jobs were eventually processed without further delay. **Key Timeline (IDT)** • **21.06.2026, 17:16 IDT**: A surge in reachability scanner jobs caused the CycodeX queue to grow rapidly. • **22.06.2026, 10:07 IDT**: The issue was identified by an on-call engineer. • **22.06.2026, 12:55 IDT**: We increased the scanning platform resources to process more jobs. • **22.06.2026, 14:10 IDT**: A fix that lowered new reachability scanners was deployed to production. • **22.06.2026, 18:43 IDT**: Existing reachability scanners' priority was lowered. • **23.06.2026, 09:12 IDT**: Scans with higher priority were processed. Only lower priority scans remained, including CCA. • **23.06.2026, 13:51 IDT**: A fix that reduced communication overload to Kubernetes was deployed. The scanning platform started processing scan jobs much faster. • **23.06.2026, 17:34 IDT**: The queue was fully processed. **Root Cause** The issue was triggered by a combination of factors: 1. **Reachability scanner job surge** -- A surge in reachability scanner jobs caused the CycodeX queue to grow rapidly, which led to resource bottlenecks in the cluster and a peak queue size of over 64,000 jobs. 2. **Excessive pod resource requests** -- Due to configuration bugs, scanner pods requested excessive CPU and memory, which prevented efficient scheduling and amplified the resource bottlenecks in the cluster. 3. **Infrastructure resource limits** -- AWS VPC subnet IP and EKS API limits were reached, restricting the cluster's ability to scale and schedule new work. 4. **Job prioritization and retry inefficiencies** -- Inefficiencies in job prioritization and retry logic meant lower priority scans (Trivy, Syft, CCA) competed for capacity and were most affected, while the backlog continued to grow. **Actions Taken** • Increased cluster and node pool capacity. • Fixed job prioritization to deprioritize reachability scans. • Capped resource requests for scanner pods to enable efficient scheduling. • Deployed additional fixes to the scanning platform. • Opened AWS support tickets to address resource limits. • Restored monitoring and logging. • Cleared the job backlog; the queue now processes new jobs as they arrive. **Action Items** • Improve monitoring to better understand the behavior of the processing environment. • Improve scanning optimization and prioritization for all scan types.
**Problem**: SAST (Static Application Security Testing) scans for pull requests were running slowly **Impact**: Some users experienced slow pull request scans potentially delaying code reviews and deployments.
The issue was resolved. The system is fully stable now