Prod2断断续续无法使用
- investigating
我们目前正在调查这一问题.
- resolved
这一事件已经得到解决.
自动翻译自官方事件更新。
88 Harness incidents · 2026年3月 — official updates, affected components, duration and resolution details.
我们目前正在调查这一问题.
这一事件已经得到解决.
自动翻译自官方事件更新。
管道行刑被卡在了Prod1. 我们正在调查这个问题.
一项措施已经执行,我们正在监测结果.
这一事件已经得到解决.
自动翻译自官方事件更新。
我们目前正在调查这一问题.
这一问题已经确定,一个解决办法正在实施之中.
一项措施已经执行,我们正在监测结果.
这一事件已经得到解决.
□ 总结 2026年8月27日到8月28日,客户遭遇了一个问题,一些管线,部署,以及相关资源在"Harness UI"和"API"中似乎没有发现,尽管基础数据仍然完好无损. 这个问题发生在计划中的内部基础设施更新期间,这影响到内部平台服务之间的沟通。 因此,取决于账户、组织和项目范围解决能力的请求无法成功完成,导致未找到的不正确答复被退回给现有实体的客户。 工程查明了问题,回滚了变化,恢复了正常服务. 事件中没有丢失或删除客户数据. □ 根源 这个问题是Production计划内部服务线路更新过程中引入的配置错误所引发. 负责解决账户、组织和项目背景的内部平台服务在应用变更后无法验证其他Harness服务的请求。 由于在许多实体读取和进行管道相关行动之前需要采取这一验证步骤,失败的请求向客户表明,不存在正常存在的资源错误。 这一问题仅限于受影响的生产环境,通过恢复变革和恢复先前的服务通信途径而得到解决。 □ 影响 * 一些客户认为,在UI和API中似乎没有现有的管道、部署和相关实体。 * 一些与管道有关的业务,包括执行进展、网络触发启动、预定触发评价和实体列名,暂时中断。 * 这个问题影响到现有实体的可用性和可见性,但没有删除数据或改变客户的配置。 * 没有发生未经授权的访问,也没有发现客户数据丢失。 □ 补救 * ** 立即:** 恢复了基础设施配置更新,并恢复了之前的工作服务通信路径. * ** 恢复验证:** 核实受影响实体的检查、管道作业和依赖性API通常在回滚后运作。 * ** 常设:** 纠正了与更新相关的配置处理,因此类似的问题不会在未来推出时干扰服务对服务的认证。 □ 动作项目 为了防止这些问题再次发生,哈妮丝将 1. 改进配置验证,加强部署前测试,在转移生产流量之前核实内部服务通信。 2. 加强对内部认证失败的监测和警报,以便及早发现问题。 3. 改进对错误的处理,使客户不太可能发现依赖性故障,因为资源没有发现错误.
自动翻译自官方事件更新。
我们目前正在调查IACM管道中报告的Prod-1、Prod-2、Prod-4和EU1带组的问题.
我们正在继续调查这一问题.
我们恢复了造成所有各组问题的变化.
这一事件已经得到解决.
自动翻译自官方事件更新。
我们目前正在调查这一问题.
我们正在调查关于地物管理和实验(FME)用户界面未能加载的报告. 试图进入FME控制台的客户可能会遇到出错或无响应的页面. 地物旗的评价和SDK流量被认为没有受到影响. 稍后将进行进一步更新.
一项措施已经执行,我们正在监测结果.
这一事件已经得到解决.
□ 总结 *从2026年8月23日**23:42 UTC**起,多家FME客户报告装载FME UI失败. * CDN提供的FME UI Artiffacts由于保留政策而过期,导致FME UI无法加载. * 任何国旗请求通过API更改,更改发送,数据管道继续不间断地工作. □ 根源 * FME UI由CDN提供。 UI的文物由于保留政策被逐出,导致UI无法为所有用户加载. □ 影响 * FME UI无法为所有生产环境中的用户装入。 什么没有受到影响? * SDK功能和运行时旗评价 * API行政电话 * 客户旗帜配置数据 * 未发生数据损失 □ 补救 *FME UI通过部署在CDN恢复了功能. * 在事件结束前,所有生产环境的恢复都得到证实。 □ 动作项目 * 改进资产保留政策,使目前有效的版本永远不会被逐出.
自动翻译自官方事件更新。
我们目前正在调查这一问题.
缓慢可造成以下症状之一: - 管道没有启动 - 延迟执行 - 由于超时而取消管道
我们的云提供商正面临一个积极事件,我们正在采取后续行动.
我们正在继续调查这一问题.
我们的云提供商已经确认 正在发生的 影响多个区域的事件。 利用管道没有因此出现故障,尽管一些用户可能继续经历缓慢. 我们正在密切监测局势,并将在获得更多信息后提供最新情况.
我们正在全面观察继我们的云层提供者所执行的修复措施之后出现改进的延误。 我们正在继续密切监测局势,并将酌情提供进一步的最新情况。 我们发现一些CI的死因是针对几个顾客的 我们正在调查
这一事件已经得到解决.
总结 2026年8月20日,从协调世界时下午3:00左右开始,哈内斯平台在所有生产环境中都经历了广泛的性能退化. 通常在两分钟内完成的管道处决需要7至10分钟。 连续交付、连续一体化、管道管弦以及地物管理和实验都受到影响。 Google Cloud平台在我们西部1地区经历了多产品事件,影响了Bigtable,Compute Engine,Google Kubernetes Engine,以及恒定盘I/O. Harness生产基础设施运行于该地区的恒定盘上. 退化将数据库操作的延迟度从大约2米提高到了95分之10分以上,这反过来又造成信息排行滞后,并传播到每个依赖于及时访问数据库的服务。 # 影响 这是一种退化,而不是停产。 管道在整个过程中继续顺利地执行和完成;它们缓慢而不是失败。 没有丢失数据,也没有因这一事件而放弃客户工作。 # 旋转原因 # 受影响环境中的利用性生产基础设施运行于我们西部1 地区的Google Cloud平台持续磁盘上。 当存储层退化时,效应在可预见的链中通过平台传播: ** 我们西面的活性磁盘一/O降解。 ** Google Cloud平台经历了影响Bigtable,计算引擎,Google Kubernetes引擎的多产品事件,并遭遇了持久性磁盘性能. 这是在哈内斯无法控制的供应商环境中发生的基础设施故障。 # ** 预防行动** 虽然哈尼斯无法防止云供应商基础设施故障. 以下行动旨在更快地发现一个,并更有能力采取行动。 行动** | -- -- 继续按此事件进行目标跨区域数据库故障的例行预先测试,以保持故障准备状态得到核实,而不是假设 * 评估跨区域延迟是不可接受的未来情况的全部多区域故障准备情况
自动翻译自官方事件更新。
我们目前正在调查这一问题.
这一问题已经确定,一个解决办法正在实施之中.
一项措施已经执行,我们正在监测结果.
这一事件已经得到解决.
页:1 2026年8月20日,在协调世界时10:24至14:55之间,一个分集FME写作失败. 由FME UI所制作的并使用 Harness 访问符\(PATs和SATs\)所制作的书写没有受到影响. 运行时旗评价继续正常进行. 这个问题通过在共享治理服务中恢复近期的认证变更而得到缓解,受影响写作在协调世界时14:55恢复正常. 现状:[https://status.harness.io/incidents/rhthgm7d5dkz](https://status.harness.io/incidents/rhthgm7d5dkz) 互联网档案馆的存檔,存档日期2013-12-02. 根原因 : 共享治理服务如何认证入境电话导致一些FME写作被拒绝. 这些文章使用了服务对服务证书,即治理部门在变更后无法再核实。 FME向客户端呈现治理失败为HTTP 499,当治理政策有意否认改变时使用的相同状态. 因为499是一个有效的,预期的反应 在这个否认路径, 失败看起来不像一个停用 在我们的警报, 和事件是从客户报告 而不是内部发现。 影响 * FME 的一个子集在窗口中写出失败,主要是使用遗留的 Split API 密钥或更改请求的排程来编写的. * FME UI制成的书信没有受到影响。 * 使用 Harness 访问符\ (PATs 和 SATs\) 写作没有受到影响。 * 运行时旗评价继续正常进行。 * 未发生数据损失。 写入失败没有应用 。 补救 恢复治理服务认证变更. 受影响作者立即恢复正常. 行动项目 为了防止这些问题再次发生, *在因治理无法评价而导致写作失败时,Harness会返回一个明显的错误\(不是499\),所以它不会与有意的政策否定相混淆. * 增加关于治理评价的提醒,而不是依赖客户端状况代码。 * 扩大对政策评价的认证支持。 * 扩大其他写作情景的自动覆盖.
自动翻译自官方事件更新。
我们目前正在调查这一问题.
我们正在继续调查这一问题.
这一问题已经确定,一个解决办法正在实施之中.
一项措施已经执行,我们正在监测结果.
我们正在继续监测任何其他问题.
这一事件已经得到解决.
** 概要** 2026年8月19日协调世界时12:35至17:29之间,哈内斯应用安全服务遭遇了严重的中断,同时影响到了SaaS生产区和US1区的客户-发报控制台和数据摄取管. ** 循环原因** 为几乎所有其他组件提供运行时间设置的内部配置服务被超载并进入了重复重启周期. 由于如此多的服务依赖于它,其效果是广泛的:控制台页面,如保护政策,姿态视图,活动日志,API库存,以及自定义政策等都未能加载或超时,下游处理在等待无法得到的配置时也停滞了. # 客户影响** 详细情况** | -- -- -- -- -- -- -- QQ Console \ (UI\) 影响 QQ 多页无法加载或超时, 包括保护政策、 态势事件页和仪表板和透视页内的态势视图、 活动日志查询、 API 库存屏幕、 自定义政策、 敏感数据视图和部件 。 | 安全遥测处理严重退化,在某些路径中完全停止。 消费者在正常化、分组、异常检测、生成和相关加工阶段的滞后。 | 数据损失 中断时摄入的一组遥测数据被永久放弃. | ** 缓解** 一些中间减缓措施增加了CPU和内存,放宽了健康检查门槛,重新启动了数据库,一个更大的连接池改善了这一问题。 使这两个受影响区域的新特点失效,其吞吐量迅速而持久地得到恢复。 事件于协调世界时17:29解决. # ** 预防行动** 已承诺采取下列行动,并进行内部跟踪,直至完成。 引发这一事件的特征仍然被禁用,在以下作品完成并验证之前不会被重新启用. 行动** | -- -- | QQ 通过调取缓存删除和保留等参数来优化代码, QQ 为服务范围界定访问模式添加一个目的创建的数据库索引QQ * 补救管道恢复语义,使消费者在位置标记器丢失后能够安全地重放,而不是跳过积压的 * * * 授权分阶段推出改变下游请求模式的配置:低数量组,然后是中数量组,然后是高数量组 QQ 为配置服务添加后压和货币保护: 线路断开, 限定队列, 以及超时隔离 QQ * 增加可观察性,采用更详细的衡量标准
自动翻译自官方事件更新。
We are currently investigating a Harness component that is experiencing issues. We are working to identify the cause and restore normal operations as soon as possible.
We are continuing to investigate this issue.
The issue has been identified and a fix is being implemented.
This incident has been resolved.
## Summary Customers on Prod1, Prod2, and Prod3 \(US\) clusters experienced failures when loading SEI 2.0 dashboards on August 6, 2026, from 7:22 AM PDT to 9:03 AM PDT. Customers calling the SEI 2.0 API also experienced similar failures. No customer data was lost, and ingestion of all integration data continued to work uninterrupted. SEI customers using 1.0 were not impacted. ## Root Cause The incident was caused by resource exhaustion on the nodes serving queries. This resource degradation developed in a pattern that did not cross our existing alerting thresholds early enough to provide sufficient warning or allow mitigation before customer impact occurred. ## Impact Customers on Prod1, Prod2, and Prod3 \(US\) clusters were unable to load SEI 2.0 dashboards during the incident window. **Duration:** August 6, 2026, 7:22 AM PDT – 9:03 AM PDT \(~1 hour 41 minutes\) ### What was not impacted? * Data ingestion and processing * SEI 1.0 customers * Integrations and metadata flows No customer data was lost. ## Remediation Upon identifying the root cause, our team took immediate corrective action by adding capacity to restore the affected systems. Services were fully recovered, and all dashboards resumed normal operation at 9:03 AM PDT. ## Action Items To prevent from such issues happening again, Harness is/has Proactively added capacity updates have been applied to prevent this issue from recurring #### Enhanced Monitoring and Alerting Additional monitoring and alerting have been put in place to detect anomalies early, focused on a leading indicator, which in this case was thread pool exhaustion, before they can impact dashboard availability and data rendering. #### System Patch in Progress We are working with our vendor to apply a patch to remediate this and similar issues completely.
我们正在监视被卡住的管道 随着我们不断监测这些部门,新的处决正在过去.
我们正在监视被卡住的管道 对于那些仍然看到被卡住的管道的顾客,我们请求你们中止并重新触发.
这一事件已经得到解决.
□ ** 摘要** 2026年8月6日\(早上的PDT\),一些在Prod2生产环境中运行管线的客户观察到管线处决停止了进展——这些阶段没有取得进展,也没有产生进一步的输出或状态更新. 受影响的客户报告了这一问题。 机能工程师查明了原因,减轻了影响,管道处决恢复了正常运行. 这个问题是由自通管道表达引起的. 一个Git Webhook触发了一个管道,该管道引用了Webhook有效载荷的内容,有效载荷本身还载有同一表达的其他副本。 因此,每一轮表达式决议产生更多的表达式,每次将工作量增加一倍。 这耗尽了处理处决案件和分配给同一案件的其他处决案件的服务案件的资源,但当案件处于这种状态时,无法取得进展。 □**影响** 2026年8月6日, * 一些客户对Prod2的行刑拖延了中期执行,没有取得进一步的进展。 * 受影响的处决没有产生新的步骤产出或状况更新,必须在减轻影响后中止和重新进行。 * 行为仅限于受影响的服务机构正在处理的处决——其他机构继续正常执行。 没有数据损失**。 管道定义、执行历史和存储状态没有受到影响。 在整个事件中,Prod2上的多数管道继续顺利地执行;主要影响是一些飞行中的处决无法完成,一旦问题得到缓解,就需要重新进行。 □** 轮回原因** 利用管道支持在运行时解决的表达式——例如,一种插入触发管道的Git Webhook有效载荷内容的表达式. 在这种情况下,Git承诺电文包含有效载荷表达式本身的文字文本,两次,管道引用了同样的有效载荷表达式。 由于承诺信息是webhook有效载荷的一部分,解析表达法插入了整个有效载荷——包括承诺信息中包含的表达法的两字拷贝. 这些新插入的副本随后作为待解决的表达方式处理,每次通行证再插入两个完整的有效载荷副本。 因此,所处理的价值以及处理价值所需的工作在每次通过时都翻了一番,并且呈指数增长而不是趋同。 Harness有一个保障措施,旨在完全阻止这种情况:表达式分辨率被一个最大筑巢深度所约束,超过这个深度,分辨率会停止而管道会以明显的出错而失效. 这一保障的缺陷意味着,在这一具体的自我优惠案件中没有适用这一限制,因此,解决办法继续不受限制。 表达式分辨率在启动管道步骤的线条上运行。 由于每次传出都逐渐消耗了更多的内存和CPU而从未完成,履行这项工作的服务实例停止了进展,分配给该实例的每一个执行都停滞不前——这是客户报告的。 □** 军事** 利用技术完成了以下即时缓解步骤: * 指明了输油管和对径行解析负责的表达模式。 * 停止了受影响的服务,使其不再继续工作。 其余的健康案例通常会得到处理并排队处决。 * 确认管道处决恢复正常,并结束了事件。 这些行动恢复了正常的管道执行行为并解决了对客户的影响. □ ** 行动项目** 为了减少再次出现的风险并改进检测,以下行动处于不同的实施阶段: * 修正表达深度和循环检测保障中的缺陷,使自我特惠表达方式被抓住并迅速以明显的错误而失败,而不是消耗不受约束的资源. * 防止有效载荷表达方式脱离触发有效载荷含量,从而完全去除自我偏好路径。 * 将最大表达式嵌入深度收紧,除了现有的深度限制外,还评价明显的环检测. * 加强生产前环境中的自动测试,以复制自我偏好表达模式,并核实保障措施检测和阻止它们。 * 在准备处决时增加对这种模式的监测,以便主动发现.
自动翻译自官方事件更新。
We are currently investigating this issue.
We are continuing to investigate this issue.
We have identified the issue and started to implement the fix , prod2 is restored.
We are continuously rolling out fixes to all Clusters, Prod EU1 has been resolved.
We are continuously rolling out fixes to all Clusters, Prod3 has been resolved.
This incident has been resolved.
# Executive Summary On August 4, 2026, between approximately 3:36 PM and 9:00 PM IST, customers using Infrastructure as Code Management \(IaCM\) on Prod0 and Prod1 were unable to access the Variable Sets settings page. The page rendered blank with no error message, and customers with Variable Sets attached to their workspaces could not view or manage them for the duration of the incident. Prod2, Prod3, and EU1 were not affected. Separately, during the same window, a scheduled maintenance action caused the IaCM settings tab to temporarily disappear across all environments. This was identified and reversed within the incident bridge call before significant customer impact occurred. We deployed a hotfix that restored full access to the Variable Sets page on Prod0 and Prod1 the same evening, and we are implementing permanent safeguards described below to prevent this class of issue from recurring. # Impact * Customers with the Variable Sets feature enabled on Prod0 and Prod1 were unable to view or manage Variable Sets for approximately 5–6 hours. * No data was lost or corrupted, this was a UI routing failure only; underlying Variable Sets data and configuration were not affected. * Prod2, Prod3, and EU1 were not affected by this issue. * A secondary issue, a scheduled feature flag operation caused the IaCM settings tab to temporarily disappear across all environments during the incident bridge call. This was identified and reversed within minutes. External customer exposure for this secondary issue is still being confirmed. # Root Cause A platform routing change released on July 18, 2026 updated how IaCM settings pages are resolved in the user interface. As part of that change, any settings page that had not been explicitly re-registered in the new routing structure became unreachable. The Variable Sets page had not been re-registered under the new routing structure, making it inaccessible in the environments where the routing change had been deployed — Prod0 and Prod1. Because the failure occurred at the routing layer rather than within the page itself, the page rendered blank with no visible error rather than showing a clear failure message. # Remediation ## Immediate We deployed a hotfix that re-registered the Variable Sets page in the updated routing structure, restoring access for all affected customers on Prod0 and Prod1. ## Permanent We are adding automated end-to-end tests that navigate to settings pages with relevant feature flags enabled, configured as a required gate in our release pipeline. We are also documenting and enforcing the routing constraint through static analysis so that settings pages are never inadvertently left out of the routing structure during future platform changes. # Action Items To prevent such issues from happening again, 1. Enhance automated tests, that navigate to settings pages with relevant feature flags enabled, configured as a blocking gate in the release pipeline, so this class of regression is caught before it reaches production. 2. Establish an explicit checklist step for future platform-wide architectural changes that verifies all existing settings pages remain accessible in the updated routing structure before the change is promoted to production.
我们目前正在调查这一问题.
一项措施已经执行,我们正在监测结果.
这一事件已经得到解决.
自动翻译自官方事件更新。
我们目前正在调查这一问题.
这一问题已经确定,一个解决办法正在实施之中.
一项措施已经执行,我们正在监测结果.
这一事件已经得到解决.
# ** 概要** 2026年7月25日至8月4日期间,Harness Prod 2和Prod 3集群的管线执行仪表板和概览页面显示的数据介于实时后. 管道本身继续在整个过程中正常地建造、部署和执行;问题仅限于如何迅速将执行记录复制到数据库,为报告和仪表板查看服务。 ** 没有丢失客户数据。 ** 每个受影响的记录都长期保存,一旦消除了基本限制,就会被重放到分析数据库。 Harness于2026年8月1日将受影响的集群迁移到可横向扩展的、排队支持的复制组件版本,并完成了所有受影响账户的定向数据回填。 # 旋转原因 # Harness维持一个变化数据采集组件,不断将主操作数据存储器的管道执行记录复制到单独的时间序列数据存储器中,优化用于仪表板和报告查询. 盘子完全从分析数据库读取。 当复制工作落后时,仪表板能够准确但更古老地看待世界,而执行本身则不受影响. 这是由于另一个共享相同复制路径的Harness平台模块的数据库写出量急剧而持续地增加,超过了仍在Prod 2和Prod 3中运行的更古老的单层组件的吞吐量上限. 积压工作形成并增加。 # ** 预防行动** 为了防止这些问题,已经或承诺采取下列行动。 行动** | -- -- QQ 精细调谐复制时滞提示, 以通知超过定义阈值的任何延迟 QQ 在标准平台监测板上添加一个复制后置面板, 通过每个模块率限制或实体在复制流上过滤,减少从共同租户模块中的写入放大
自动翻译自官方事件更新。
我们目前正在调查这一问题.
这一问题已经确定,一个解决办法正在实施之中.
一项措施已经执行,我们正在监测结果.
这一事件已经得到解决.
# **Summary** On July 31, 2026, artifact uploads performed through pipeline in the EU1 cluster began failing with an authentication error. Uploads initiated manually \(outside of a pipeline\) were not affected, and the ability to retrieve existing artifacts \(downloads\) was also unaffected — this was isolated to the specific pipeline upload path in one cluster. # **Impact** * Artifact uploads performed through pipeline in the EU1 cluster failed with an authentication error for approximately 4 hours and 34 minutes. * Retrieving existing artifacts \(downloads\) was not affected. * Manually uploading artifacts outside of a pipeline was not affected. * Other clusters/regions were not affected by this issue. # **Root Cause** The component responsible for handling pipeline-based artifact uploads is distributed as a container image. In the EU1 cluster, this image is retrieved from an internal registry that mirrors a public image source; in other clusters, the same image is retrieved directly from the public source. A publishing error in our release process caused a new build of this component to be published using a version label that was already in use, rather than being assigned a new, unique version. As a result, two different images ended up associated with the same version label in the public source. Our internal registry mirrors images from the public source via an automated replication process. Because of how that replication was triggered, it copied the original \(earlier\) image associated with that version label rather than the corrected one. This meant the EU1 cluster — which pulls from the internal mirror — ended up running a different, defective image than other clusters, which pull directly from the public source and therefore received the corrected image. The defective image contained an authentication issue that caused pipeline uploads to fail. # **Mitigation** * Reverted the affected account to the last known-good version of the upload component, immediately restoring pipeline uploads. * Published a corrected, permanent version of the component to resolve the issue across all clusters. # **Next steps** * Fix the upload step to remove the underlying container-related defect that made this failure mode possible. * Update our release pipeline for this component so that publishing an image can never overwrite an existing version — every publish must create a new, distinct version going forward.
自动翻译自官方事件更新。
目 录 - 我们间歇地面临网络连通性问题,因为我们的Build VM无法连接到外部资源. 我们目前正在调查这一问题.
一项措施已经执行,我们正在监测结果.
我们正在继续监测任何其他问题.
这一事件已经得到解决.
## Summary Starting on August 4, 2026, CI runners in the us-west1 and us-central1 regions intermittently experienced connection timeouts of approximately 134 seconds when reaching external services such as GitHub and Bitbucket over outbound network gateways. ## Impact * CI runners in the affected regions intermittently experienced connection timeouts of approximately 134 seconds when reaching external services \(e.g., GitHub, Bitbucket\) over our outbound network gateways. * The issue was intermittent rather than constant — connections succeeded under normal load, and failures clustered during periods of high outbound traffic volume. * No data was lost or corrupted. This was a network-connectivity and capacity issue, not a data-integrity issue. * us-west1 and us-central1 were the affected regions; other regions were not impacted by this issue. ## Root Cause Our load balancer distributes outbound traffic across multiple NAT gateways using a hashing method based on connection details \(source/destination address and port\). For any single connection, these details stay constant for that connection's lifetime. We had a sustainted traffic surge for a few seconds which congested the gateways ## Action Items To prevent such issues from happening again Harness will, Increase outbound connection capacity on our NAT gateways by provisioning additional external network interfaces, giving each gateway a substantially larger pool of connections it can serve concurrently..
自动翻译自官方事件更新。
我们目前正在调查这一问题.
一项措施已经执行,我们正在监测结果.
这一事件已经得到解决.
自动翻译自官方事件更新。
我们目前正在调查这一问题.
我们正在继续调查这一问题.
这一问题已经确定,一个解决办法正在实施之中.
我们正在继续努力解决这个问题。
一项措施已经执行,我们正在监测结果.
我们正在继续监测任何其他问题.
这一事件已经得到解决.
# Summary During a recent production deployment, a defect in our internal deployment tooling caused two critical services to run with incorrect, non-production configuration values in our production environment This led to a related set of four distinct symptoms: incorrect configuration behavior, intermittent login/access failures, a filestore access issue affecting one customer environment, and delayed pipeline status updates in the UI. We have identified and are implementing a permanent fix for the underlying configuration defect, and have already put in place resource and capacity changes that resolve the UI delay symptom. At no point during this incident were pipeline executions themselves lost, corrupted, or left in a stuck state. Where execution behavior was affected, it was limited to delays in status visibility, not in the underlying processing. # Incident Details ## Incorrect Production Configuration Values Applied Our engineering team confirmed a defect in the Service Manager deployment pipeline that caused certain production services to be deployed using configuration values intended for a different environment, rather than the correct production configuration. **Root Cause** The service responsible for fetching configuration overrides during deployment queries an internal API that returns a maximum of 1,000 results per request. The total number of services in the environment recently grew beyond that limit. As a result, any service beyond the first 1,000 returned was not included in the response, and the deployment pipeline silently fell back to default configuration values for those services. This is a confirmed pagination defect in the deployment tooling, not an issue with the configuration values themselves. **Resolution** Engineering has confirmed the mechanism and is implementing a permanent fix to remove this limit-related gap in the deployment pipeline. ## Intermittent Login / Access Failures During the Service Manager deployment referenced above, some users experienced intermittent login or access failures. Under normal operation, previously running instances should continue serving traffic without interruption while a new deployment is in progress. In this incident, that fallback behavior did not occur as expected, contributing to access failures during the deployment window. ## Filestore Access Issue A filestore access issue was identified that was specific to the Prod-3 environment and affected a single customer's environment. **Root Cause** This is related to an IAM / storage-bucket permission configuration on Service Manager, potentially triggered by rollback activity. ## Delayed Pipeline Execution Status Updates in UI Some users observed that the pipeline execution graph in the UI was slow to refresh and did not reflect the latest status promptly. Importantly, this was a visibility delay only: there was no impact to actual pipeline executions, and no executions were stuck or failed as a result of this issue. **Root Cause** The pipeline execution graph relies on a message stream \(the orchestration log\) to receive status updates. During the incident window, consumer processing of this stream fell behind \(high consumer lag\), which delayed how quickly status updates reached the UI. This was caused by the fact that the underlying database was in the middle of a planned scaling operation at the same time, and a traffic spike during that window further exacerbated the delay. Users experienced this as apparent pipeline slowness, even though the underlying executions were running normally. **Resolution** We have increased resource capacity for the affected components to maintain more than 50% spare headroom going forward, reducing sensitivity to similar load spikes. This change has been implemented and is currently being validated as part of longer-term hardening for this part of the platform. # Impact Summary * Service Manager and License Manager ran with incorrect configuration values in the Prod-1 and Prod-3 environments. * Some users experienced intermittent login or access failures during the affected deployment window. * One customer environment in Prod-3 experienced a filestore access issue. * Users across affected environments saw delayed pipeline execution status updates in the UI; underlying pipeline executions continued to run correctly and were not lost, stuck, or corrupted. # Preventive Actions The following corrective and preventive actions have been identified. | **Corrective / Preventive Action** | | --- | | Correct the pagination limit in the configuration-lookup service so that all services are returned and evaluated, regardless of total count. | | Add safeguards so that a service which cannot retrieve its configuration fails safely \(e.g. alerts and blocks the deployment\) rather than silently falling back to non-production defaults. | | Increase Postgres and messaging-pipeline resource headroom \(target: greater than 50% spare capacity\) to reduce sensitivity to concurrent load and scaling events. | _We recognize the impact this incident had across multiple areas of the platform and appreciate your patience as we work through a complete resolution._
自动翻译自官方事件更新。
我们正在调查一个影响AIDI仪表板的问题。 用户在访问仪表板时可能遭遇加载时间或间歇性故障. 我们的团队正在积极努力查明根源,并恢复正常业绩。 一旦获得更多信息,我们将提供最新资料.
这一问题已经确定,一个解决办法正在实施之中.
一项措施已经执行,我们正在监测结果.
这一事件已经得到解决.
□ 总结 Prod1,Prod2,和Prod3集群上的客户在2026年7月22日访问AIDI 2.0仪表板时经历了断断续续的部件负载故障并增加了负载时间. 并非所有部件都同时受到影响,问题表现为零星故障,而不是完全停产。 没有客户数据丢失。 SEI 1.0客户没有受到影响。 □ 根源 随着时间的推移,一个例行的数据库维护程序未能运行在分析数据库中的某些表格上,导致这些表格积累了大量用于跟踪被删除记录的内部元数据。 当数据库计划对这些表格进行查询时,它把这些累积的元数据都装入内存,导致受影响的节点的内存使用再三激增. 这些再三的突起引发了自动安全机制,在发现过量的内存压力后会重启一个节点,受影响的节点也因此开始以循环方式重启. 这使得事故期间AIDI 2.0仪表板上的查询性能断断续续地退化. □ 影响 Prod1,Prod2,和Prod3集群上的客户可能会在AIDI 2.0仪表板上发生间歇部件负载故障或加载时间增加. ** 期限:** 2026年7月22日,07:58 PDT – 16:16 PDT \ (~8小时18分\),断断续续的部件故障;系统从08:25 PDT起被重新启动并接受活监测. 什么没有受到影响? * 数据摄入和处理 * SEI 1.0客户 * 整合和元数据流动 没有丢失客户数据。 □ 补救 发现问题后,于08:25重新启动了受影响的数据库节点,恢复了最初的稳定性. 我们继续密切监测该系统,当在此之后仍然观察到断断续续的退化时,我们又应用了几项新的补救措施: * 调整了数据库配置设置,以限制用于处理累积元数据的内存量,并调制了查询规划设置来降低内存压力. * 开展清理工作以减少受影响表格上累积元数据的积压。 * 提高受影响数据库节点的能力,以提供额外的头室。 这些变化逐渐稳定了系统,事件于16:16PDT被完全解决. □ 动作项目 为防止再次发生这种情况,我们正在执行以下措施: 1. 我们升级了后端,其中包括处理过多删除文件造成的内存突起的基本改进。 2. 我们为新引入的表格推出了自动压缩工作,以防止文件累积工作向前推进.
自动翻译自官方事件更新。
我们目前正在调查这一问题.
这一问题已经确定,一个解决办法正在实施之中.
一项措施已经执行,我们正在监测结果.
这一事件已经得到解决.
总结 2026年7月17日,经过例行的代码部署后,旧代办版本\(858xx和以下\)的客户开始经历延迟的CI,利用我们全球的建设排队能力,在Harness Cloud托管的建筑上建立. 受影响的建筑物在继续前在“等待基础设施”阶段意外地暂停了约8分钟,而不是在预期次秒内进行。 总体建设缓慢是间歇性的。 # 影响 * 所有CI建筑都有可能被推迟;影响最显著的是利用全球建筑排队功能建造由Harness Cloud托管的基础设施。 * 受到影响的建筑物在继续前曾出现长达约8分钟的无法解释的停顿,随后由于没有预留的计算档位——这是作为缓慢的建筑物而不是建造故障提供给用户的,因此"冷起"速度较慢. * 更新的代表版本\(858xx及以上\)的账目没有受到影响。 * 没有建筑物因这一问题而彻底失败,也没有丢失数据。 # 根源 根本原因是内部代码的更改,无意中打破了一个特定的构建排队记录在建立我们的数据库时,一旦在之前的代码版本下排队,就会从我们的数据库中读回,从而遇到新部署的版本. 我们清理了受影响的记录,恢复了基本的代码修改,从而解决了眼前的影响。 我们正在执行若干保障措施,防止这类问题再次发生。 接下来的步骤 我们评估出现类似情况的风险,因为以下行动正在消失。 导致这一事件的具体代码路径已经恢复,我们正在执行结构性保障措施,以便这一一般问题类别不再发生,无论在代码库中可能发生在哪里。 ** 更正/预防行动** | -- -- QQ 在存储于我们数据库的所有内部数据类别中添加明确,稳定的标识符,这样未来的内部代码重组无法破坏系统读取先前存储的记录的能力. | 在我们的生产前环境中引入回滚和后向相容性测试,专门设计在达到出产前抓住这一类问题. |
自动翻译自官方事件更新。
我们目前正在调查这一问题.
这一问题已经确定,一个解决办法正在实施之中.
一项措施已经执行,我们正在监测结果.
我们正在继续监测任何其他问题.
这一事件已经得到解决.
总结 2026年6月19日至7月17日,使用无人机-git克隆插件的内置Git Clone步骤——以及任何管道步骤——在ARM64 Kubernetes上以错误exec/usr/local/bin/clone:exec格式错误来构建基础设施失败. AMD64\(英特尔/AMD\)建立,Windows建立,VM无容器二进制路径没有受到影响. 根本原因是我们内部的图像出版过程存在缺陷,导致ARM64被标记为无人机的图像实际上包含AMD64二进制. 我们在同一天查出并缓解了这个问题,报告将无人机引擎图像恢复到最后一个已知的好版本。 无需客户采取行动或改变配置。 # 根源 2026年6月19日,一次安全整治调整了无人机-git图像的构建方式. AMD64建筑得到正确更新,但ARM64建筑管线并没有直接构建ARM64文件——它通过文本替换来修改了AMD64建筑文件,并将其编译在ARM64基础设施上. 6月19日的更改更改了AMD64文件,使得替代默默地不开放而不是失败,因此管道发布了一个被标记为ARM64的图像,其Git Clone和Git LFS的二进制仍然被编译为AMD64. # 影响 *受影响:内置的克克隆人步骤,以及任何使用无人机-克克隆人插件的管道步骤,运行于ARM64 Kubernetes上,于2026年7月6日至7月17日期间跨越所有账户建设基础设施. * 症状: 由 exec/usr/local/bin/clone:exec格式错误在 Git Clone 步骤下构建失败. *未受影响:AMD64\(Intel/AMD\)Kubernetes和VM构建,Windows构建,VM无容器执行路径,以及我们的硬化图像变体. # 缓解 我们把所有受影响的服务中使用的无人驾驶飞机的图像版本恢复到最后已知的好版本。 这完全解决了ARM64执行失败;不需要修改客户配置. 接下来的步骤 防止这些问题再次发生。 * 重建ARM64图像发布管道,以建设我们专用的ARM64直接构建文件,而不是改造AMD64构建文件. *加强自动的后推后验证到每个图像放出:验证二进制架构与图像标记相匹配,并在一个图像被视为可再放出之前进行功能烟雾测试. * 扩大自动化测试覆盖范围,包括ARM64 Kubernetes构建情景。 * 更正后的发布经验证后,将受影响的中间图像版本从发行中移除.
自动翻译自官方事件更新。