US07 のダッシュボードのレイテンシとエラー
- identified
US07 Dashboard に影響する断続的なエラーやレイテンシを調査しています.
- resolved
これは、ダッシュボードに影響する以前の問題の再発でした。 恒久的な修正が実施され、ダッシュボードは以来安定しています。 問題は解決しました。 誠に有難うございます.
公式のインシデント更新を自動翻訳しています。
56 Braze incidents · 2024年11月 — official updates, affected components, duration and resolution details.
US07 Dashboard に影響する断続的なエラーやレイテンシを調査しています.
これは、ダッシュボードに影響する以前の問題の再発でした。 恒久的な修正が実施され、ダッシュボードは以来安定しています。 問題は解決しました。 誠に有難うございます.
公式のインシデント更新を自動翻訳しています。
エンジニアリングチームは、US07 Dashboardの断続的なレイテンシーとエラーを調査しています.
当社は、断続的なダッシュボードのエラーとレイテンシーとエンジニアが問題を解決するために取り組んでいる問題を特定しました.
エンジニアは修正を実装し、ダッシュボードは30分以上安定しています。 問題の監視を続け、完全な解像度を確保しています.
Dashboard は引き続き安定しており、問題は解決されています.
公式のインシデント更新を自動翻訳しています。
We are currently investigating an issue with the US 01 and US 03 clusters.
Between 11:30 ET and 12:55 ET, some customers experienced elevated error rates. Braze engineers identified the issue and resolved it by rolling back a recent change and applying database optimizations. Error rates returned to normal shortly after each fix was applied, and they have remained stable for the past hour. This incident is now resolved.
現在、EU02クラスターで劣化した性能を調査し、お客様の小さなサブセットで現在の処理の遅延を引き起こしています。 データが失われず、現在のデータの大部分は正常に流れ続ける。 当社のチームは、解決に向けて積極的に取り組んでおり、それらが利用できるようにアップデートを提供します.
EU02の劣化性能を調査し続けています。 私たちは、根本的なレイテンシーの根本的な原因を特定しました。これは、複数の電流コネクタからステンドされ、定期的なメンテナンスの後、優雅に再起動に失敗しましたが、私たちの監視システムで健康な状態に表示されます。 これらの特定のサーバーは、この状態で重要なバックログを作成します。 私たちは、優雅な再起動を完了し、バックロックされたアイテムをより迅速に処理するために追加のリソースをスピンしました。 この時点では、当社のコネクタの~5%まで、お客様のサブセットに影響し、レイテンシを引き続き経験する一方で、データの95%は正常に流れています。 設定変更により、スループットが大幅に改善され、1時間または更新されたタイムラインがある場合に更新を提供します.
EU02で残りのバックログをクリアし、パフォーマンスを改善し、全体的なレイテンシを減少させていきます。 次回は、イベントの数が少々増えることもありますが、次の時間内に巻き込まれる予定です。 残りの衝撃は、全体のトラフィックの50%未満に限定されます。 完全にキャッチアップしたら、別のアップデートを提供します.
この事件は解決しました。 15:18 ET / 19:18 UTCと同様に、ジョブのバックログを完全に処理し、EU02電流コネクタは、通常のレイテンシ範囲内で再び処理されます。 データが失われず、すべての現在のイベントが目的地に送られてきました.
公式のインシデント更新を自動翻訳しています。
午前7時30分頃に米国WEST-2地域におけるAWSネットワークの普及率は、6時55分~8時の間にネットワーク接続に影響を及ぼしました。 AWS は、根本的な問題の解決を続けてきました。 インパクトウィンドウでは、一部の顧客は、電子メール配信で断続的な遅延と遅延を経験している可能性があります。 メールでのお問い合わせ キューの深さが増加すると、一部の顧客はメッセージ配信で残留遅延が発生する可能性があります。 この遅延は、今日の午後2時40分に解決しました。 この混乱は、AWS インフラストラクチャ内で発生し、Braze プラットフォームの任意のコンポーネントによって引き起こされるものではありませんでした。 AWSインシデントの詳細については、AWSステータスページhttps://health.aws.amazon.com/health/status をご覧ください
公式のインシデント更新を自動翻訳しています。
Braze engineers have identified an issue affecting the Braze AI Operator, where users may see the following error message: "Connection unstable - Check your Wi-Fi or internet signal and try again". We are actively investigating the issue and will provide updates as more information becomes available.
A fix is being deployed for the issue affecting the Braze AI Operator, where users were seeing the error message: "Connection unstable - Check your Wi-Fi or internet signal and try again." We are seeing error rates improve and are continuing to monitor closely to confirm full resolution. We will provide a further update once we can confirm the issue is fully resolved.
The issue impacting the Braze AI Operator, where users were seeing the error message: "Connection unstable - Check your Wi-Fi or internet signal and try again," has been fully resolved. The fix has been deployed and error rates have returned to normal levels.
Meta, a third-party, is currently experiencing an issue impacting WhatsApp services. During this time, WhatsApp messages may experience latency. Campaign and Canvas creation or modification may experience errors and latency when utilizing WhatsApp templates. Audience sync to Facebook may also experience errors and latency during this period.
Meta's services appear to be stabilizing. We are continuing to monitor the situation and will provide a further update once full service has been restored.
Meta has recovered from the earlier issue impacting WhatsApp services, and full service has now been restored. WhatsApp message delivery, Campaign and Canvas creation or modification using WhatsApp templates, and Audience sync to Facebook are all operating normally.
Starting at 1:00 PM ET, customers on our US-01 and EU-01 clusters are experiencing elevated latency affecting campaign and Canvas processing, data ingestion, and message sending. Our engineering team is actively investigating, and we will provide updates as the situation develops.
Our engineering team continues to investigate the elevated latency. We will provide another update in 30 min or sooner if there are any changes.
Our engineering team has implemented a change, and we are beginning to see improvements in latency across campaign and Canvas processing, data ingestion, and message sending on US-01 and EU-01. We are monitoring the situation closely.
Latency and error rates have remained at normal levels across all services following the implementation of our fix. This incident is resolved.
Starting at 11:45 PM ET, Braze customers on our US05 have been experiencing latency across campaign/canvas processing, data processing, and message sending due to an ongoing issue with an upstream provider (AWS). We are working on a fix. Please see https://health.aws.amazon.com/health/status for more details.
We are continuing to experience delays in Campaign & Canvas Processing, Data Processing, and Message Sending. Additionally, Dashboard navigation is intermittently erroring. We are continuing to work on a fix while we await updates from AWS. Please see https://health.aws.amazon.com/health/status for more details.
We have implemented a fix to mitigate the impact of the ongoing AWS incident in Availability Zone (use1-az4) in the US-EAST-1 region. At this point, latency and error rates have returned to normal levels for all services. Additionally, AWS is reporting early signs of recovery. Please see https://health.aws.amazon.com/health/status for more details.
Latency and error rates have remained at normal levels for all services, post-implementation of our fix, and this incident is resolved. We will continue to monitor the overall system health until AWS reports full recovery. Please see https://health.aws.amazon.com/health/status for more details.
We have identified a third party issue causing Agent Console auto models to error. Our engineering team is actively working to remediate the issue.
Our engineering team have released a fix and Agent Console auto models are now recovering. We are continuing to monitor the situation.
Following continued monitoring by our engineers, Agent Console auto models have remained stable. This incident is now resolved.
Braze Engineers have identified an issue causing SMS delivery delays across networks in multiple European countries. Braze Engineers confirmed this is related to an ongoing Twilio incident: https://status.twilio.com/incidents/shhrhbqxflt1 We expect to provide another update as soon as more information becomes available.
Braze Engineers continue to monitor the ongoing SMS delivery delays affecting multiple networks across European countries. This remains related to an active Twilio incident, and our team is closely tracking their progress. Twilio has identified the cause and is actively working to resolve the issue. We will provide another update as soon as the situation changes. For real-time updates from Twilio, please refer to: https://status.twilio.com/incidents/shhrhbqxflt1
Braze Engineers continue to monitor the ongoing SMS delivery delays affecting multiple networks across European countries. This remains an active Twilio incident, and our team is closely tracking their progress. Twilio has identified the cause and is actively working to resolve the issue. They state that delivery delays to multiple networks continue in Romania, Serbia, Poland, Bosnia and Herzegovina, and Cabo Verde. The incident has been resolved, and latency has returned to normal in all other countries. Braze's next update on the issue will be 10 UTC on 9 May 2026. For real-time updates from Twilio, please refer to: https://status.twilio.com/incidents/shhrhbqxflt1
SMS sending delays have been resolved for the majority of countries and sending has returned to normal. SMS delivery delays continue to impact Romania, Serbia, Poland, Bosnia and Herzegovina, and Cabo Verde. For real-time updates from Twilio, please refer to: https://status.twilio.com/incidents/shhrhbqxflt1
Latency has returned to normal operational bounds and this incident is now resolved. Twilio's Statuspage will continue to provide updates on the remaining impact to Romania, Serbia, Poland, Bosnia and Herzegovina, and Cabo Verde: https://status.twilio.com/incidents/shhrhbqxflt1
We are currently investigating an issue with the EU01 cluster.
Braze engineers are actively investigating the issue and working to restore services.
Braze engineers have identified the issue. Services remain degraded but are improving.
Braze engineers are processing the remaining backlog of messages. Data processing remains in a degraded state due to backlog and latency issues.
Braze engineers are continuing to work through the remaining messages in the backlog. We are now seeing improvements in processing speeds, though data processing remains in a degraded performance state while the backlog is fully cleared.
The message backlog has been cleared, and data processing has returned to normal operational bounds. All services continue to remain stable, and this incident is considered resolved.
Between 16:00 and 18:14 ET, Braze Engineers identified and resolved an issue causing intermittent SDK errors on the US01 cluster. As of 18:19 ET, services remain stable, and any failed SDK requests during the impact window were automatically retried.
We are currently experiencing an issue where cases sent to [email protected] may not be created as expected. Our team is actively investigating the issue. In the meantime, if you require assistance, please submit your request via the Braze Support Portal or reach out to your Customer Success Manager. We will provide an update as soon as more information becomes available.
We have identified the root cause of the issue is Google Workspace incident (https://www.google.com/appsstatus/dashboard/incidents/224ozRqzW4sFBDK8hLnT). This is causing delays in emails being sent and received.
We are continuing to investigate this incident. Both Google Workspace and Salesforce's Email-to-Case functionality are reporting issues related to the processing of email. New cases created by emailing support@braze as well as responses on existing cases are not being processed. In the meantime please submit your request by clicking "Get Help" in the top-right corner of your Braze workspace. Alternatively, your CSM may also be able to support.
We have begun receiving responses in cases that were delayed by this incident. Case creation via email is now being processed. Support is actively working through the backlog of requests and follow-ups created during the disruption.
The issue affecting email processing for support cases has been resolved. New cases sent to [email protected] and replies to existing cases are now being processed as expected. Our Support team is actively working through the backlog of requests created during the disruption. Thank you for your patience while we worked through this.
Between 17:00 and 18:35 EST, Braze Engineers identified and resolved an issue causing intermittent SDK errors on the US01 cluster. As of 18:40 EST, services remain stable, and any failed SDK requests during the impact window were automatically retried.
Starting around 16:45 UTC, Braze Engineers began observing an issue where EU02 end users may experience errors when accessing or attempting to unsubscribe via the email preference center.
As of 17:47 UTC, Braze Engineers have implemented a fix and are monitoring service health for improvement.
No further issues have been experienced, and all services are operating normally. The incident is now resolved.
Braze Engineers are investigating an issue affecting users' ability to reach the Braze Dashboard for the US07 Cluster, starting around 13:41 EST. At this time, no other services are impacted, and messaging sending is not impacted. Engineers are actively working to restore service.
As of 15:12 ET, Braze engineers have implemented a fix and are continuing to monitor service health for improvement.
No further issues have been experienced, and all services are operating normally. The incident is now resolved.
Braze Engineers are investigating an issue affecting users' ability to reach the Braze Dashboard for the US05 Cluster, starting around 16:15 EST. At this time, no other services are impacted, and messaging sending is not impacted. Engineers are actively working to restore service.
As of 16:47 ET, Braze engineers have implemented a fix and are continuing to monitor service health for improvement.
No further issues have been experienced, and all services are operating normally. The incident is now resolved.
Braze Engineers detected an increase in latency on email outbound messaging for customers on Bird (Sparkpost) non-EU Braze clusters. This latency build up started at 13:53 ET. Braze has engaged Sparkpost, who have declared an incident and are actively investigating, you can follow their Statuspage for updates, and we'll update here in 30 minutes, or when we learn more. https://status.sparkpost.com/incidents/ndjn00g5vwm3
As of 14:34 ET, Sparkpost has identified the root cause to be a networking issue on their US cluster. We are monitoring performance for improvement.
As of 14:41 ET, Sparkpost has declared this incident resolved after implementing a fix. Braze Engineers have confirmed that real time outbound email latency has returned within normal operational bounds as of 14:34 ET. Any messages that failed during this window, were automatically retried. This incident is considered resolved.
We have received reports of errors when navigating the EU01 and EU02 Dashboards. Braze engineers are observing a small increase in errors and are actively investigating the issue.
We are continuing to investigate the issue.
We are continuing to investigate this issue with urgency and have escalated to our network service provider.
We've identified an issue isolated impacting our edge network impacting customers accessing the dashboard from Europe. We have escalated the issue to our network provider and are investigating with urgency.
The errors have stabilised to normal operating levels. We are continuing to monitor the situation closely and are actively investigating with our networking provider.
We can confirm services have continued to remain stable and the issue is now considered resolved.