US07上的挂板缓冲和出错
- identified
我们正在调查断断续续的误差 以及撞击美国07飞机板的延迟性.
- resolved
这是早先影响Dashboard的问题的再现。 永久固定装置已经实施,Dashboard自此保持稳定. 我们相信,这个问题现已得到全面解决。 我们为任何影响道歉.
自动翻译自官方事件更新。
56 Braze incidents · 2024年11月 — official updates, affected components, duration and resolution details.
我们正在调查断断续续的误差 以及撞击美国07飞机板的延迟性.
这是早先影响Dashboard的问题的再现。 永久固定装置已经实施,Dashboard自此保持稳定. 我们相信,这个问题现已得到全面解决。 我们为任何影响道歉.
自动翻译自官方事件更新。
我们的工程组正在调查 US07Dashboard上的间歇性延迟和错误.
我们查明了导致间歇性仪表板出错的问题,而且时间紧迫,工程师正在努力解决这一问题.
工程师实施固定,达什板稳定了30多分钟. 我们正在继续监测这一问题,以确保全面解决.
Dashboard继续保持稳定,问题现已解决.
自动翻译自官方事件更新。
We are currently investigating an issue with the US 01 and US 03 clusters.
Between 11:30 ET and 12:55 ET, some customers experienced elevated error rates. Braze engineers identified the issue and resolved it by rolling back a recent change and applying database optimizations. Error rates returned to normal shortly after each fix was applied, and they have remained stable for the past hour. This incident is now resolved.
我们目前正在调查欧盟02集团的退化业绩,这导致对少数客户的海流处理出现延误。 没有丢失数据,大多数海流数据继续正常流动. 我们的小组正在积极努力以达成一项决议,并将在得到决议后提供最新情况.
我们继续调查欧盟02年的退化业绩。 我们查明了造成这种长期性的根本原因,其根源是多个海流连接器在例行维修后未能顺利重新启动,但在我们的监测系统中呈现出健康状态。 这些特定的服务器在这个状态下会积存大量积压. 我们完成了一次优雅的重新启动,并调动了更多的资源来更快地处理积压的项目。 此时~95%的数据是正常流出,而高达~5%的连接器,撞击了一子客户,却继续经历暂时性. 我们看到,由于配置变化,吞吐量大有改进,将在1小时后或当我们有更新的时间表时提供最新情况.
我们继续取得进展,清理欧盟02联盟的剩余积压,改进我们的业绩并减少总体延迟。 我们期望在下个小时内赶上,尽管少数事件可能需要稍长的时间才能充分进行。 剩下的影响 限制在我们总流量的不到一半。 一旦我们完全赶上,我们会提供另一个更新.
此事已决. 截至15:18 ET / 19:18 UTC,我们已完全处理好积压的工作,我们的EU02电流连接器都在正常延迟范围内再次处理. 没有丢失数据,所有海流事件都已发送到目的地.
自动翻译自官方事件更新。
在大约7:30 AM ET时,我们意识到AWS在US-WEST-2区域网络中断,影响到6:55 AM ET至8:00 AM ET之间的网络连接。 自那时起,阿军已经解决了根本问题。 在撞击窗口期间,一些客户可能遭遇间歇性延迟和电子邮件发送不便. 邮件现在正常发送。 随着队列深度的提高,一些客户可能遭遇了信息发送的剩余延迟. 这些拖延从今天的2:40 PM ET解决了. 这种干扰源于AWS基础设施,并非由Braze平台的任何部件所造成. 关于AWS事件的更详细情况,请参阅AWS状态网页:https://health.aws.amazon.com/health/status
自动翻译自官方事件更新。
Braze engineers have identified an issue affecting the Braze AI Operator, where users may see the following error message: "Connection unstable - Check your Wi-Fi or internet signal and try again". We are actively investigating the issue and will provide updates as more information becomes available.
A fix is being deployed for the issue affecting the Braze AI Operator, where users were seeing the error message: "Connection unstable - Check your Wi-Fi or internet signal and try again." We are seeing error rates improve and are continuing to monitor closely to confirm full resolution. We will provide a further update once we can confirm the issue is fully resolved.
The issue impacting the Braze AI Operator, where users were seeing the error message: "Connection unstable - Check your Wi-Fi or internet signal and try again," has been fully resolved. The fix has been deployed and error rates have returned to normal levels.
Meta, a third-party, is currently experiencing an issue impacting WhatsApp services. During this time, WhatsApp messages may experience latency. Campaign and Canvas creation or modification may experience errors and latency when utilizing WhatsApp templates. Audience sync to Facebook may also experience errors and latency during this period.
Meta's services appear to be stabilizing. We are continuing to monitor the situation and will provide a further update once full service has been restored.
Meta has recovered from the earlier issue impacting WhatsApp services, and full service has now been restored. WhatsApp message delivery, Campaign and Canvas creation or modification using WhatsApp templates, and Audience sync to Facebook are all operating normally.
Starting at 1:00 PM ET, customers on our US-01 and EU-01 clusters are experiencing elevated latency affecting campaign and Canvas processing, data ingestion, and message sending. Our engineering team is actively investigating, and we will provide updates as the situation develops.
Our engineering team continues to investigate the elevated latency. We will provide another update in 30 min or sooner if there are any changes.
Our engineering team has implemented a change, and we are beginning to see improvements in latency across campaign and Canvas processing, data ingestion, and message sending on US-01 and EU-01. We are monitoring the situation closely.
Latency and error rates have remained at normal levels across all services following the implementation of our fix. This incident is resolved.
Starting at 11:45 PM ET, Braze customers on our US05 have been experiencing latency across campaign/canvas processing, data processing, and message sending due to an ongoing issue with an upstream provider (AWS). We are working on a fix. Please see https://health.aws.amazon.com/health/status for more details.
We are continuing to experience delays in Campaign & Canvas Processing, Data Processing, and Message Sending. Additionally, Dashboard navigation is intermittently erroring. We are continuing to work on a fix while we await updates from AWS. Please see https://health.aws.amazon.com/health/status for more details.
We have implemented a fix to mitigate the impact of the ongoing AWS incident in Availability Zone (use1-az4) in the US-EAST-1 region. At this point, latency and error rates have returned to normal levels for all services. Additionally, AWS is reporting early signs of recovery. Please see https://health.aws.amazon.com/health/status for more details.
Latency and error rates have remained at normal levels for all services, post-implementation of our fix, and this incident is resolved. We will continue to monitor the overall system health until AWS reports full recovery. Please see https://health.aws.amazon.com/health/status for more details.
We have identified a third party issue causing Agent Console auto models to error. Our engineering team is actively working to remediate the issue.
Our engineering team have released a fix and Agent Console auto models are now recovering. We are continuing to monitor the situation.
Following continued monitoring by our engineers, Agent Console auto models have remained stable. This incident is now resolved.
Braze Engineers have identified an issue causing SMS delivery delays across networks in multiple European countries. Braze Engineers confirmed this is related to an ongoing Twilio incident: https://status.twilio.com/incidents/shhrhbqxflt1 We expect to provide another update as soon as more information becomes available.
Braze Engineers continue to monitor the ongoing SMS delivery delays affecting multiple networks across European countries. This remains related to an active Twilio incident, and our team is closely tracking their progress. Twilio has identified the cause and is actively working to resolve the issue. We will provide another update as soon as the situation changes. For real-time updates from Twilio, please refer to: https://status.twilio.com/incidents/shhrhbqxflt1
Braze Engineers continue to monitor the ongoing SMS delivery delays affecting multiple networks across European countries. This remains an active Twilio incident, and our team is closely tracking their progress. Twilio has identified the cause and is actively working to resolve the issue. They state that delivery delays to multiple networks continue in Romania, Serbia, Poland, Bosnia and Herzegovina, and Cabo Verde. The incident has been resolved, and latency has returned to normal in all other countries. Braze's next update on the issue will be 10 UTC on 9 May 2026. For real-time updates from Twilio, please refer to: https://status.twilio.com/incidents/shhrhbqxflt1
SMS sending delays have been resolved for the majority of countries and sending has returned to normal. SMS delivery delays continue to impact Romania, Serbia, Poland, Bosnia and Herzegovina, and Cabo Verde. For real-time updates from Twilio, please refer to: https://status.twilio.com/incidents/shhrhbqxflt1
Latency has returned to normal operational bounds and this incident is now resolved. Twilio's Statuspage will continue to provide updates on the remaining impact to Romania, Serbia, Poland, Bosnia and Herzegovina, and Cabo Verde: https://status.twilio.com/incidents/shhrhbqxflt1
We are currently investigating an issue with the EU01 cluster.
Braze engineers are actively investigating the issue and working to restore services.
Braze engineers have identified the issue. Services remain degraded but are improving.
Braze engineers are processing the remaining backlog of messages. Data processing remains in a degraded state due to backlog and latency issues.
Braze engineers are continuing to work through the remaining messages in the backlog. We are now seeing improvements in processing speeds, though data processing remains in a degraded performance state while the backlog is fully cleared.
The message backlog has been cleared, and data processing has returned to normal operational bounds. All services continue to remain stable, and this incident is considered resolved.
Between 16:00 and 18:14 ET, Braze Engineers identified and resolved an issue causing intermittent SDK errors on the US01 cluster. As of 18:19 ET, services remain stable, and any failed SDK requests during the impact window were automatically retried.
We are currently experiencing an issue where cases sent to [email protected] may not be created as expected. Our team is actively investigating the issue. In the meantime, if you require assistance, please submit your request via the Braze Support Portal or reach out to your Customer Success Manager. We will provide an update as soon as more information becomes available.
We have identified the root cause of the issue is Google Workspace incident (https://www.google.com/appsstatus/dashboard/incidents/224ozRqzW4sFBDK8hLnT). This is causing delays in emails being sent and received.
We are continuing to investigate this incident. Both Google Workspace and Salesforce's Email-to-Case functionality are reporting issues related to the processing of email. New cases created by emailing support@braze as well as responses on existing cases are not being processed. In the meantime please submit your request by clicking "Get Help" in the top-right corner of your Braze workspace. Alternatively, your CSM may also be able to support.
We have begun receiving responses in cases that were delayed by this incident. Case creation via email is now being processed. Support is actively working through the backlog of requests and follow-ups created during the disruption.
The issue affecting email processing for support cases has been resolved. New cases sent to [email protected] and replies to existing cases are now being processed as expected. Our Support team is actively working through the backlog of requests created during the disruption. Thank you for your patience while we worked through this.
Between 17:00 and 18:35 EST, Braze Engineers identified and resolved an issue causing intermittent SDK errors on the US01 cluster. As of 18:40 EST, services remain stable, and any failed SDK requests during the impact window were automatically retried.
Starting around 16:45 UTC, Braze Engineers began observing an issue where EU02 end users may experience errors when accessing or attempting to unsubscribe via the email preference center.
As of 17:47 UTC, Braze Engineers have implemented a fix and are monitoring service health for improvement.
No further issues have been experienced, and all services are operating normally. The incident is now resolved.
Braze Engineers are investigating an issue affecting users' ability to reach the Braze Dashboard for the US07 Cluster, starting around 13:41 EST. At this time, no other services are impacted, and messaging sending is not impacted. Engineers are actively working to restore service.
As of 15:12 ET, Braze engineers have implemented a fix and are continuing to monitor service health for improvement.
No further issues have been experienced, and all services are operating normally. The incident is now resolved.
Braze Engineers are investigating an issue affecting users' ability to reach the Braze Dashboard for the US05 Cluster, starting around 16:15 EST. At this time, no other services are impacted, and messaging sending is not impacted. Engineers are actively working to restore service.
As of 16:47 ET, Braze engineers have implemented a fix and are continuing to monitor service health for improvement.
No further issues have been experienced, and all services are operating normally. The incident is now resolved.
Braze Engineers detected an increase in latency on email outbound messaging for customers on Bird (Sparkpost) non-EU Braze clusters. This latency build up started at 13:53 ET. Braze has engaged Sparkpost, who have declared an incident and are actively investigating, you can follow their Statuspage for updates, and we'll update here in 30 minutes, or when we learn more. https://status.sparkpost.com/incidents/ndjn00g5vwm3
As of 14:34 ET, Sparkpost has identified the root cause to be a networking issue on their US cluster. We are monitoring performance for improvement.
As of 14:41 ET, Sparkpost has declared this incident resolved after implementing a fix. Braze Engineers have confirmed that real time outbound email latency has returned within normal operational bounds as of 14:34 ET. Any messages that failed during this window, were automatically retried. This incident is considered resolved.
We have received reports of errors when navigating the EU01 and EU02 Dashboards. Braze engineers are observing a small increase in errors and are actively investigating the issue.
We are continuing to investigate the issue.
We are continuing to investigate this issue with urgency and have escalated to our network service provider.
We've identified an issue isolated impacting our edge network impacting customers accessing the dashboard from Europe. We have escalated the issue to our network provider and are investigating with urgency.
The errors have stabilised to normal operating levels. We are continuing to monitor the situation closely and are actively investigating with our networking provider.
We can confirm services have continued to remain stable and the issue is now considered resolved.