资源审计日志延迟了活动[eu-west-1][by company 00006]
- investigating
我们目前正在调查这一问题.
- investigating
一些客户(第14行)可能因为审计日志中的系统出错而看到审计日志中缺失的资源变化事件. 没有客户信息数据受到影响,只有审计日志.
- resolved
这一事件已经得到解决.
自动翻译自官方事件更新。
51 Front incidents · 2024年5月 — official updates, affected components, duration and resolution details.
我们目前正在调查这一问题.
一些客户(第14行)可能因为审计日志中的系统出错而看到审计日志中缺失的资源变化事件. 没有客户信息数据受到影响,只有审计日志.
这一事件已经得到解决.
自动翻译自官方事件更新。
We are currently investing an issue where a subset of customers are seeing offline O365 channels.
We are attempting a mitigation step to get impacted channels back online.
Front continues to investigate paths to restore access for affected Office365 channels. Office reports internal errors, which we are trying to work around.
Front is partnering with Microsoft to investigate the channel sync failures and will continue to update here as we make progress. Manually re-authorizing the channel in Front's channel settings has worked in some cases, but not all. Please contact [email protected] for assistance.
Front is continuing to partner with Microsoft to investigate the channel sync failures. The next update will be within the next 8 hours. Manually re-authorizing the channel in Front's channel settings has worked in some cases, but not all. Please contact [email protected] for assistance.
Front is continuing to partner with Microsoft to investigate the channel sync failures. The next update will be within the next 6 hours. Manually re-authorizing the channel in Front's channel settings has worked in some cases, but not all. Please contact [email protected] for assistance.
Front is continuing to partner with Microsoft to investigate the decreasing volume of channel sync failures. The next update will be within the next 12 hours. Manually re-authorizing the channel in Front's channel settings has worked in most cases, but not all. Please contact [email protected] for assistance.
Front is continuing to partner with Microsoft to investigate the channel sync failures. The next update will be within the next 12 hours. Manually re-authorizing the channel in Front's channel settings has worked in most cases, but not all. Please contact [email protected] for assistance.
Front is continuing to partner with Microsoft to investigate the channel sync failures. The next update will be within the next 12 hours. Manually re-authorizing the channel in Front's channel settings has worked in majority of cases, but not all. Please contact [email protected] for assistance.
Front is continuing to partner with Microsoft on root cause of remaining channel sync failures. We added the ability to automatically bring these channels back online to prevent further customer impact.
Counters (numbers in the navigation bar of the app) are updating with a delay for some customers [us-west-1]
This incident has been resolved.
There was a delay in message delivery of up to an hour. Messages are now being delivered without delay. We are continuing to monitor
This incident has been resolved.
Some customers are experiencing delays in inbound messages being received from Gmail and Office channels (typically 2-4 minutes of delay). We are currently scaling up our systems to address this.
Delays have been reduced, but remain slightly elevated. We are continuing to investigate.
Office delays have been resolved. Gmail delays have reduced further (to approximately 1 minute).
Delays are resolved for both Gmail and Outlook channels. We're continuing to monitor to ensure we're at a stable steady-state.
This incident has been resolved.
We are currently investigating this issue.
The issue has been identified and a fix is being implemented.
We are continuing to work on a fix for this issue.
A fix has been implemented and we are monitoring the results.
This incident has been resolved.
We are currently investigating an issue that is causing delayed message import. Impacting all channels all regions
This incident has been resolved.
Front is currently inaccessible. We have identified the cause and are working on a resolution.
We are continuing to work on a fix for this issue.
Application and API access have been restored and we are continuing to monitor the system.
This incident has been resolved. A fix has been implemented.
Microsoft Office365 is experiencing an outage affecting some Front channels: "Users may be seeing degraded service functionality or be unable to access multiple Microsoft 365 services." Front is unable to send or receive messages for affected customers. Their status page is available at https://status.cloud.microsoft/.
Microsoft appears to have recovered, and we're no longer seeing any impact on Office 365 channels. All pending messages have been delivered. We're continuing to monitor the situation closely.
Microsoft is still reporting service degradation, but it's not impacting Front. We'll continue monitoring the situation in case it changes.
We are currently investigating the issue. [us-west-1], [us-west-2]
We are continuing to investigate this issue.
The issue has been identified and a fix is being implemented.
A fix has been implemented and we are monitoring the results. We are seeing recovery but continuing to close out remaining recovery items.
We are continuing to monitor and close our our remaining recovery items for [us-west-2] customers. Other regions are recovered.
Front is operational for all customers. We are continuing to backfill any missed messages and application webhooks.
All backfills complete
On Friday, Dec 19, at 14:50 UTC \(6:50am PST\), customers backed in our US-West-2 data center experienced dramatically increased API latency, resulting in the website failing to load and messages being queued in the backend. This continued until 18:10 UTC \(10:10am PST\). During this time no messages were lost, though there may have been a significant delay for messages to appear in customer inboxes. All queued messages were delivered by 21:00 \(1:00pm PST\). Customers based in Front’s EU-West-1 and US-West-1 datacenters may have experienced some delays during this time, as some systems are interdependent, but this impact was intermittent and uncommon. The root cause of this issue was the failure of a caching system. There are several database systems that support the Front application, which are supported by caching to improve performance. A recent change increased the size of some objects in the cache layer. This is not inherently wrong, and did not have any immediate impact. On Friday the 19th the caching layer in US-West-2 crossed a new threshold of data volume which triggered a large number of evictions, particularly of other data that is necessary for most application activity. Besides putting additional load on the databases, there was simply not enough room in the cache for all the data we needed to store there. This caused a high amount of thrashing that significantly increased latency for all systems.
We are currently investigating an issue that is causing errors loading conversations. We believe it is limited in scope to a subset of customers who use our slack integration
We believe we have identified the issue, currently working on implementing a fix
We are rolling out the fix now
We've rolled out a fix to the issue, all conversations should now load. Please reach out to support if you continue to experience any issues
This incident has been resolved.
We have identified an issue where delegated inboxes are not appearing in the navigation panel. We are working on a fix for this issue and will provide another update once there is an ETA for the resolution.
We are continuing to work on a fix for this issue.
We have restored delegated inboxes. If delegated inboxes are not currently visible, please restart the application.
We are currently investigating an outage.
A rollback has started recovering the failures in us-west-1. We are monitoring and further investigating root cause
Recovering messages for all regions but restoring back to operational state. Continuing to monitor
Fully recovered for us-west-1 and eu-west-1 regions. Continued recovery for us-west-2 customers
Us-West-2 customers are now recovered. We are continuing to monitor
Investigating continued outages for [by_company_00017] customers only.
Continuing investigation for [by_company_00017] customers only.
Recovering app for [by_company_00017] customers, delay for rules and calendar remains.
This incident has been resolved.
At 19:05 UTC \(11:05 am PST\), Front began a routine application deployment. Front typically deploys new versions of the application multiple times per day as part of normal operation. This app release included a change in the way we connect to an internal caching system. The change passed all our testing requirements, but due to an environment-specific configuration on the caching system it immediately caused a problem in our customer environment. Front was alerted within minutes of the increased error rate and immediately started investigating. At 19:20 UTC we initiated a rollback of the deployment across all regions. Because of the progressive nature of the rollback, some customers saw the application recover quickly after while for others it took up to 30 minutes. By 19:50 all rollbacks were complete and most customers were back to normal. However, for one slice of Front customers the rollback triggered a second issue that kept the application from recovering. For clarity, Front divides customers into one of 40 “cells” that isolate them from one another. This architecture is designed to limit the scope of certain failures, like this one. But for the 2% of customers on the affected cell we had to take additional steps to recover. This second issue was triggered by a spike in database load caused by the stampede of traffic from customers returning to the application. In this cell a particular workload caused high database contention that ultimately timed out and then started over, which prevented us from escaping the problem. Traffic continued to back up and we were unable to progress through all the messages. Once we identified that the problematic workload was in the evaluation of rules, we were able to block all rules from processing. This allowed the database and the application to immediately recover, coming back online for users at 21:39 UTC. Rule evaluations continued to back up, and we were able to slowly restart processing to make sure the database didn’t reenter a bad state. All rule evaluations were caught up by 22:20 UTC. We would like to apologize for the disruption this outage has caused and for the unusually long duration. In the immediate aftermath of the incident we have identified a number of actions we can take to prevent this kind of event from occurring again. The first step is to ensure we have a test environment that appropriately represents the configuration of the production customer environment for the caching system. This issue should have been detected before we ever deployed it. The second task is to investigate the rollback system to see if we can safely improve the speed, so that if we have a similar issue in the future we can recover even faster.
Some Canadian IP addresses are experiencing app slowness and/or degraded email message performance.
We are continuing to investigate this issue.
We are still investigating the issue and searching for the root cause.
We are continuing to investigate this issue, and are yet to identify a root cause Some impacted users have been able to workaround the issue by using a VPN to re-route their traffic, or by tethering to a mobile device. We will update this page as soon as we have more information
We've received reports this morning that users are no longer encountering issues. We are continuing to monitor, please reach out to support if you are experiencing degraded performance
We've had no further reports of issues. Please reach out to Front support if you experience degradation in the app
App is not loading in specific regions and timeouts have been spiking for some API calls
A cache node was lost and automatically recovered in 6 minutes. All systems should now be operational.
We've have received reports of the google sign in flow failing to redirect back to the mobile app successfully, and have been able to reproduce the issue We are currently investigating the problem
We have isolated the issue to iOS devices, Android customers should be unaffected Continuing to investigate
We've identified the issue, and are still working on a fix. In the meantime we have identified a workaround that appears to be working on some clients - by navigating directly to https://app.frontapp.com/oauth/google-signin?client_type=universalLink in Safari on iOS and signing you should get redirected to the app in a signed in state
We've released a new version of the app (5.69.1) that fixes the behavior in some cases, but does not fully address the issue. We are in progress on a full fix. For clients still unable to login we have identified a workaround that appears to be working on some clients - by navigating directly to https://app.frontapp.com/oauth/google-signin?client_type=universalLink in Safari on iOS and signing you should get redirected to the app in a signed in state
We've implemented a fix - you can download the latest version from the app store (5.69.2) to test. Please let us know if you experience any issues on this version
We have confirmed reports that the latest version of the iOS app (5.69.2) addresses the issue. If you run into login issues, ensure you've updated to the latest version, and reach out if the problem persists.
Front is currently investigating an issue in which trying to load some conversations results in an error message.
Conversations that leveraged Suggest Reply between 21:30 and 21:50 UTC (2:30pm – 2:50pm PDT) may cause an error when loading. Front has identified the root cause and is deploying a fix that will allow these messages to display.
A fix for the affected conversations has now been deployed, and initial tests show all conversations should now be visible. Please make sure to refresh your browser or app if you don't see this right away.
Front has now completed the fix and cleanup associated with this incident. If you continue to see errors loading conversations, please make sure you've done a browser or app refresh, and if the issue persists please contact support. Summary: Conversations for which Front attempted to provide a Suggested Reply between 21:30 – 21:50 UTC may have been unreadable until the fix was completed at 22:49 UTC. During that initial time, suggested replies were generated with an additional field that caused an unrecoverable error in the frontend. No data or messages were lost or corrupted during the incident. Suggested replies for conversations that arrived during that time may need to be regenerated.
We have an ongoing issue causing delays importing new messages from all gmail channels. We have identified the issue and are currently working on a fix
Fixes are currently rolling out, we are monitoring backlog of unimported messages
Continuing to monitor our backlog of messages, the majority of channels have recovered - but still have a small segment still catching up
All channels should now be recovered
We are investigating reports of Google SSO and Office SSO login issues on our mobile apps (iOS and Android).
We are continuing to investigate this issue.
The issue has been identified and a fix is being implemented.
This incident has been resolved.
We are currently investigating reports of increased latency affecting a subset of customers in the EU region. Impact: Some customers in EU may experience slower response times or intermittent delays when using our app. Not all customers in this region are affected.
We are continuing to investigate this issue.
The issue has been identified and a fix is being implemented.
Response times are stabilizing, though we continue to actively monitor the situation to ensure the issue is fully resolved.
This incident has been resolved.