Some users may be experiencing intermittent connection issues. Our team is actively investigating.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
Users being redirected to last opened email
Início 14 Du 2024 da 15:38 UTC · 14m
Pending
investigating
We are currently investigating this issue.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
Sporadic Dyspatch Application Outages
Início 11 Gwengolo 2024 da 20:40 UTC · 6m
OutageMajor incident
Componentes afetados
APIDashboard
monitoring
Some users may be experiencing issues accessing Dyspatch. We've implemented a solution and are monitoring it.
resolved
This incident has been resolved.
Application and API unavailable
Início 8 Ebrel 2024 da 19:18 UTC · 13h 42m
OutageMajor incident
Componentes afetados
APIDashboard
investigating
Some users may be experiencing availability issues on the Dyspatch application and API. The Dyspatch engineering team is investigating.
identified
The issue has been identified and a fix is being implemented.
identified
We are continuing to work on a fix for this issue.
monitoring
A fix has been implemented and we are monitoring the results. Thank you for your patience.
resolved
This incident has been resolved.
postmortem
# **Post Mortem** - **April 8 2024 Dyspatch Outage Intro**
On April 8, 2024, Dyspatch was unavailable between the hours of 12:30PM and 01:00AM Pacific time due to an issue that occurred during a routine upgrade of Dyspatch's infrastructure. This post mortem aims to analyze the root causes of the outage, assess its impact on our services, and outline steps Dyspatch is taking to prevent similar incidents in the future.
## **Timeline \(Pacific Time\)**
**11:35 -** We begin the upgrade
**12:10 -** The production cluster intermittently returns 503s for users. Dyspatch's services cannot communicate with each other.
**12:17 -** We attempt to rollback the changes.
**12:30 -** We identify the problem: the internal authentication mechanism our services use to communicate securely is out of sync across services.
**12:30 - 17:30 -** We try several strategies to bring production online.
**17:30** - To avoid further impact to our production environment, work begins on our staging environment.
**18:17 -** We identify that previous changes were made to our staging environment without getting applied to our production environment.
**21:16 -** Staging is online. We begin applying the changes from our staging environment to our production environment.
**00:56 -** Dyspatch is available again.
## **Why did this happen? What did we learn?**
During the outage we ran into several challenges trying to restore service. We discovered that a previous update to a critical component of our infrastructure was applied only to our staging environment. It was quickly determined that the issue was an authentication misalignment between Dyspatch's services which meant that our various services could not communicate with each other. We learned that we did not have a way to generate new credentials without taking the services that manage our cluster offline. After we determined that critical services had to be taken offline we switched to testing on our staging environment to prevent data loss in our production environment.
Ultimately a difference in our production and staging environment had knock-on effects affecting our ability to rollback and recover quickly.
## **What are we doing about it?**
There are several actions we intend to take to prevent similar issues from happening:
1. We immediately aligned our staging and production environments to ensure that any infrastructure testing done in staging will be the same when applied to our production environment. The root cause of this outage came from a difference in environments and this ensures that we can be confident when testing required infrastructure changes.
2. We plan to invest in tooling to help us automatically catch and audit any drift between our environments. Catching the difference beforehand would have prevented this incident.
3. We are investing in tooling and processes to help us rebuild our cluster more reliably and quickly. We had to spend time migrating changes from our staging environment to our production environment when trying to restore Dyspatch.
## **Summary**
Finally, we want to apologize. We know Dyspatch is important for supporting our customers' communications. Your patience and support mean a great deal to us and we appreciate everyone who reached out to our team. Like with any operational issue, we will spend time in the coming days and weeks to understand the details of the event and make improvements mentioned above to our infrastructure and processes.
Mobile preview display issues
Início 30 Du 2023 da 09:30 UTC · 0m
Pending
resolved
Mobile previews in the email builder were not working. Desktop previews and email editing in general were unaffected. A fix has been implemented and deployed.
Sporadic Dyspatch Application Outages
Início 24 Du 2023 da 03:28 UTC · 53m
OutageMajor incident
Componentes afetados
APIDashboard
investigating
Some users may be experiencing issues accessing Dyspatch. We are currently investigating.
identified
The issue has been identified and a fix is being implemented.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
Email service provider exports not working
Início 16 Du 2023 da 21:31 UTC · 1h 10m
Pending
investigating
Email exports with integrations like Braze, Iterable, et cetera are not completing for some users. The Dyspatch engineering team is investigating.
identified
The issue has been identified and a fix is being implemented.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
PDF Download and Preview Offline
Início 17 Here 2023 da 01:09 UTC · 1d 14h
IssuesMinor incident
identified
We have identified an issue with PDF downloads and previews which may affect some users.
resolved
This incident has been resolved.
Templates not loading
Início 20 Gwengolo 2023 da 14:21 UTC · 14m
OutageMajor incident
Componentes afetados
Dashboard
investigating
Some customers are experiencing issues loading some templates. We are investigating.
identified
The issue has been identified and a fix is being implemented.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
Google OAuth Login Issues
Início 30 Eost 2023 da 22:45 UTC · 8h 3m
Pending
investigating
Some customers may be experiencing issues logging into Dyspatch using Google OAuth. Our engineering team is currently investigating.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
Image upload failures
Início 30 Eost 2023 da 02:10 UTC · 21m
IssuesMinor incident
Componentes afetados
Dashboard
investigating
Our upstream image upload management provider is currently experiencing an outage and our engineering team is looking into it.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
Dyspatch Availability Issues.
Início 19 Mezheven 2023 da 17:52 UTC · 8m
OutageCritical incident
Componentes afetados
APIDashboard
investigating
The Dyspatch dashboard and API are currently unavailable. The Dyspatch engineering team is investigating.
monitoring
A solution has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
Image uploads not working
Início 14 Cʼhwevrer 2023 da 02:03 UTC · 14h 55m
IssuesMinor incident
Componentes afetados
Dashboard
identified
The Dyspatch team has identified an issue with an upstream service provider and we are working with them to resolve the issue. If you need any images uploaded to your account please email [email protected] with your images.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
Dashboard availability issues
Início 9 Cʼhwevrer 2023 da 19:32 UTC · 51m
IssuesMinor incident
Componentes afetados
Dashboard
identified
Some customers may have issues loading the dashboard. Our engineers are aware and working towards a solution.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
Dashboard Availability Issues
Início 24 Eost 2021 da 23:38 UTC · 36m
Pending
Componentes afetados
Dashboard
investigating
We're currently experiencing some intermittent issues with the Dyspatch dashboard. The engineering team is investigating.
identified
The issue has been identified and a fix is being implemented.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
Increased application latency
Início 23 Mezheven 2021 da 18:15 UTC · 6h 2m
IssuesMinor incident
Componentes afetados
Dashboard
investigating
We've identified performance issues with the Dyspatch app and our engineers are investigating.
identified
Our engineers have identified the root cause and are working towards a solution.
identified
We have a plan in place to improve the performance of the application. At 3PM PDT (22:00 UTC) Dyspatch will be unavailable while we perform emergency maintenance.
monitoring
Our engineers have deployed some changes and we're now monitoring.
resolved
This incident has been resolved. Dyspatch speeds have returned to normal.
Image library upload issues
Início 10 Cʼhwevrer 2021 da 20:14 UTC · 5h 13m
IssuesMinor incident
Componentes afetados
Dashboard
monitoring
Due to an issue with an upstream provider, some customers may be experiencing problems with uploading to the image library. We're working with them to resolve the issue.
identified
Due to an issue with an upstream provider, some customers may be experiencing problems with uploading to the image library. We're working with them to resolve the issue.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
Degraded dashboard performance
Início 27 Genver 2021 da 23:19 UTC · 43m
IssuesMinor incident
Componentes afetados
Dashboard
investigating
Some customers may notice degraded performance and timeouts in the Dyspatch dashboard. Our engineers are investigating.
identified
The issue has been identified and a fix is being implemented.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
API and Dashboard Availability Issues
Início 16 Kerzu 2020 da 18:16 UTC · 1h 48m
Pending
Componentes afetados
APIDashboard
investigating
The Dyspatch engineering team is investigating reports of API and dashboard slowdown.
investigating
We are continuing to investigate this issue.
identified
Our engineering team has identified the issue and are working toward a fix. We will continue to update you as more information becomes available.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
API and Dashboard Availability Issues
Início 31 Mae 2018 da 23:32 UTC · 3h 34m
IssuesMinor incident
Componentes afetados
APIDashboard
investigating
The Dyspatch engineering team is investigating reports of API and dashboard slowdown.
monitoring
A fix has been implemented and we are monitoring the results.