Prowadzimy śledztwo w sprawie opóźnień i błędów przy importowaniu miejsc pracy za pośrednictwem narzędzia importu. Niektórzy użytkownicy mogą doświadczyć w tym czasie nieudanych lub powolnych procesów importowych.
investigating
9: 15 AM EST - Badanie
Prowadzimy śledztwo w sprawie opóźnień i błędów przy importowaniu miejsc pracy za pośrednictwem narzędzia importu. Niektórzy użytkownicy mogą doświadczyć w tym czasie nieudanych lub powolnych procesów importowych.
9: 41 AM EST - rozwiązane
Ten incydent został rozwiązany. Funkcje Import Tool wróciły do normy i wszystkie opóźnione przywozy zostały przetworzone. Przepraszamy za wszelkie niedogodności.
Czas trwania: 26 minut
Home pageCustomer TrackerDispatch Web PlatformAndroidNotificationsiOSOutbound SMS ServiceOutbound Email ServiceOperator APIMapsPublic APIRouting and Itinerary Optimization
investigating
We are currently experiencing service disruptions due to a Microsoft Azure outage affecting Cigo Tracker’s infrastructure.
Our team is actively monitoring the situation and working closely with Microsoft to restore full functionality as soon as possible.
We will continue to provide updates as new information becomes available.
Thank you for your patience and understanding.
investigating
We are continuing to investigate this issue.
identified
We’re aware of a Microsoft Azure service disruption that began around 16:00 UTC (12:00 PM EST) and is currently affecting the availability of some Cigo Tracker services.
The issue is related to Azure Front Door (AFD), which is also impacting access to the Azure Portal and other Azure-hosted systems globally.
Microsoft has acknowledged the issue and is actively working on mitigation steps, including failover options for affected infrastructure. Their latest update was posted at 16:57 UTC (12:57 PM EST) on October 29, 2025, confirming that investigations are ongoing.
Our team continues to monitor the situation closely and will provide additional updates as Microsoft releases more information or once services stabilize.
For live updates directly from Microsoft, you can check the Azure Status Page here:
🔗 https://azure.status.microsoft/en-us/status
Thank you for your patience and understanding.
identified
To ensure our clients’ operations remain uninterrupted, we have activated our failover.
You can access the Cigo Tracker platform directly at https://app.cigotracker.com
.
Please note that our main website (home page) will remain temporarily inaccessible during this period. This decision was made intentionally to prioritize platform uptime and maintain delivery operations for all clients.
Microsoft is actively deploying their last known good configuration and has begun recovering nodes to restore full connectivity. While progress is ongoing, Microsoft has indicated that protective safeguards are extending the overall recovery time.
We will keep the system status as Degraded until we have collected sufficient data to confirm full stability, at which point we will mark the system as Operational.
You can monitor Azure’s live updates here: https://azure.status.microsoft/en-us/status
All times are reported in Eastern Standard Time (EST).
Thank you for your patience and understanding as we continue to monitor and provide updates in real time.
identified
We are continuing to work on a fix for this issue.
monitoring
We have successfully bypassed the Azure Front Door issue. While Azure continues to experience a global outage, Cigo Tracker remains fully operational through our failover.
Please note that Cigo Tracker is currently accessible only via https://app.cigotracker.com
.
We will continue to closely monitor system performance and provide updates as the situation with Microsoft Azure progresses.
Thank you for your continued patience and understanding.
resolved
Microsoft has confirmed that all affected Azure services have been fully restored.
Cigo Tracker systems were back to normal operation as of 16:59 EST.
Following service stabilization, we have now reverted from our backup environment back to our primary infrastructure. All systems are operating normally, and performance remains stable.
If you wish to receive the official Microsoft incident report once it is published, please email us at [email protected]
, and we will make it available to you as soon as we receive it.
Thank you for your patience and understanding throughout this Azure service disruption.
SSL Certification Issue Affecting Import Tool and Confirmation Module Notifications
Początek 15 września 2025 18:12 UTC · 2h 12m
OutagePoważny incydent
Dotknięte komponenty
Dispatch Web PlatformOutbound SMS ServiceOutbound Email Service
identified
We recently experienced unexpected issues with SSL certification that disrupted database connections for several of our internal services.
The Import Tool was initially affected but has since been fully restored. During our investigation, we also identified that the issue impacted the notifications subsystem for the Confirmation module, which relies on SMS and email sending.
While the Import Tool is now operating normally, we are still working to restore full connectivity for Confirmation module notifications, along with a few other automated functions that may currently be degraded.
Our engineering team has identified the root cause and is implementing safeguards to prevent similar issues during future SSL certification rotations.
We sincerely apologize for the disruption and will continue to provide updates until all affected services are fully restored.
identified
We have resolved the issues affecting the Import Tool and the Confirmation module notifications. Both are now fully operational.
Our team is currently working on finalizing the recovery of a few remaining automated functions that may still be degraded.
We will provide another update once all affected services are fully restored.
resolved
The SSL certification issue that impacted several internal services, including the Import Tool and the Confirmation module notifications, has been fully resolved. All affected services are now fully operational.
We have implemented safeguards to prevent similar issues during future SSL certification rotations.
Thank you for your patience and understanding throughout this incident.
postmortem
### Summary
On **September 15th, 2025**, we experienced service disruptions affecting the Import Tool and Confirmation module notifications. The root cause was a misidentified SSL certificate dependency during the scheduled **Azure Database for MySQL maintenance**, which included the rotation of the MySQL server’s SSL certificates.
### Impact
* The Import Tool was unavailable for a period of time.
* Confirmation module notifications \(via SMS and email\) were disrupted.
* Some other automated functions experienced degraded performance.
### Root Cause
While we were aware of the scheduled SSL certificate rotation as part of Azure’s maintenance process, our evaluation process incorrectly identified which certificates were being used by certain subsystems. As a result, several services lost connectivity when the certificates were rotated.
### Mitigation and Resolution
* Our engineering team corrected the certificate configuration for the affected subsystems.
* The Import Tool and Confirmation module notifications were restored to full functionality.
* Safeguards are being added to ensure that future SSL rotations do not cause similar disruptions.
### Contributing Factors
At the onset of the incident, we initially believed the outages were linked to updates and fixes deployed on Thursday and Friday of the previous week. This led to rollbacks that did not address the issue. Although unrelated to the outage, some of those updates remain rolled back and will be re-published in the coming days.
### Next Steps
* We are reworking our dependency system analysis to more accurately map SSL and package dependencies across all services.
* We are improving our monitoring and review processes for certificate rotations to ensure full coverage across all subsystems.
* We will re-publish rolled-back updates after additional validation.
### Conclusion
We sincerely apologize for the operational problems this incident caused. We take these issues seriously and are committed to ensuring that similar outages do not reoccur in the future.
Incident: Web Application Outage
Początek 18 lipca 2025 16:28 UTC · 2h 42m
OutagePoważny incydent
Dotknięte komponenty
Customer TrackerDispatch Web PlatformAndroidNotificationsiOSOperator APIPublic API
investigating
Our web application is currently unavailable. We’re actively investigating the cause of the outage.
Earlier today, we performed some database changes which have since been reverted. We're currently assessing whether these changes contributed to the issue and working to restore full functionality as quickly as possible.
We sincerely apologize for the inconvenience this is causing and appreciate your patience while we work to resolve the problem. Further updates will follow shortly.
monitoring
The web application is now accessible again at https://app.cigotracker.com. We are closely monitoring system performance to ensure stability and confirm that all services are functioning as expected.
We were also seeing some issues accessing the Cigo Tracker landing page (https://cigotracker.com), but this was also just reestablished.
Thank you for your continued patience. We’ll provide another update once the incident is fully resolved.
resolved
The web application outage has been resolved, and all systems are currently operating normally.
While we've successfully isolated the scope of the issue, there are still some contributing factors that remain under investigation. Our team is continuing to analyze the event to determine the full root cause of today's outage and to implement measures that prevent recurrence.
We sincerely appreciate your patience and apologize again for the disruption.
Degraded Platform Performance Due to Ongoing System Update
Początek 23 czerwca 2025 20:08 UTC · 52m
OutagePoważny incydent
Dotknięte komponenty
Customer TrackerDispatch Web PlatformAndroidiOSOutbound SMS ServiceOutbound Email ServiceOperator APIPublic API
investigating
We are currently investigating an issue related to an ongoing system update that is impacting platform performance and availability. Users may experience slow response times or intermittent access issues. Our engineering team is working to identify the root cause and restore full functionality as quickly as possible. We will share updates as progress is made.
identified
We’ve identified the cause of the issue. While platform performance has recovered, the underlying system update did not complete successfully. Our team is conducting a thorough review of the root cause and will provide further updates as we determine the next steps.
resolved
The issue has been fully resolved. We reverted the deployment that caused the disruption, and the related changes will be revised and rescheduled for a future release. Thank you for your patience and understanding throughout the investigation.
Network Degradation Impacting Cigo Services
Początek 19 marca 2025 00:44 UTC · 2h 51m
IssuesDrobny incydent
Dotknięte komponenty
Customer TrackerDispatch Web PlatformOperator APIPublic API
monitoring
We want to inform you that we recently experienced network degradation affecting our services due to an ongoing issue within Microsoft's Azure infrastructure in the East US region.
What Happened?
According to Azure, between 13:09 UTC and 18:51 UTC, a fiber cut impacted network capacity in the region, leading to intermittent connectivity loss and increased latency. While Azure has since mitigated the issue, we observed disruptions in our own services between 7:20 PM and 8:27 PM (Eastern Time), specifically affecting connections between the Cigo Tracker web app and our Redis service.
Current Status
As of 8:27 PM UTC, network latencies have returned to normal, and service stability has been restored. However, to ensure a prompt and complete resolution, we have escalated this matter to Azure's Operations Support with a critical priority.
We appreciate your patience and will continue monitoring the situation closely. If you experience any further issues, please reach out to our support team.
resolved
After monitoring the situation over the past few hours, we can confirm that the immediate impact of the outage has been fully mitigated. Our services are stable, and network performance has returned to normal.
We will provide a more detailed post-mortem once we receive a conclusive report from Microsoft's Azure Operations Support (OSS) team.
Thank you for your patience and understanding. We sincerely apologize for any inconvenience this may have caused.
postmortem
## 📣 Incident Summary – March 18–19, 2025
**Service Impact on Cigo Tracker due to Azure Regional Outage**
On March 18, 2025, Cigo Tracker experienced intermittent service disruption due to a regional outage within Microsoft Azure’s East US data center region. Below is a summary of the root cause, impact, and the steps being taken to prevent future occurrences.
### 🕒 What Happened?
Azure’s East US region suffered two separate impact windows:
* March 18, 13:37 to 16:52 UTC
* March 18, 23:20 to March 19, 00:30 UTC
The incident was triggered by a **third-party fiber cut** during external drilling work, which caused reduced network capacity in one of Azure’s Availability Zones. A **tooling failure** during Azure's recovery efforts later reintroduced traffic prematurely, leading to congestion and a second round of intermittent connectivity issues.
### 🔍 Root Cause
1. **Fiber Cut**: A construction-related accident physically damaged fiber cabling serving the East US datacenter, degrading network capacity.
2. **Router Maintenance**: A key router in the same zone was already under repair, limiting redundancy.
3. **Tooling Error**: Azure’s automated recovery system failed to fully isolate damaged infrastructure, inadvertently reintroducing traffic too early.
4. **Congestion Spillover**: The unexpected traffic load caused congestion to spread beyond AZ03 into neighboring zones.
### 🎯 Impact on Cigo Tracker
While the Azure issue only affected a subset of inter-zone traffic in East US, this included infrastructure we rely on, resulting in **intermittent connectivity issues** for some customers during the incident windows. Core services were restored once Azure manually completed isolation and fiber recovery work.
### 🛠 Resolution Timeline
* **13:37 UTC, Mar 18** – Outage begins due to fiber cut
* **13:55 UTC** – Initial mitigation starts; traffic rerouted
* **16:52 UTC** – First impact window ends
* **23:20 UTC** – Second outage begins due to tooling error during recovery
* **00:30 UTC, Mar 19** – Final mitigation complete
* **06:50 UTC** – Full restoration of all infrastructure
### ✅ What Azure is Doing to Prevent Recurrence
* Fixing tooling failures that allowed reintroduction of unready capacity \(by May 2025\)
* Accelerating a capacity upgrade for the East US datacenter \(by July 2025\)
* Architecting better safeguards to prevent impact from spreading across zones \(by February 2026\)
We apologize for the inconvenience caused. Please rest assured that our team is working closely with Azure and continuing to invest in the resiliency of our platform.
If you have any questions or would like help designing a more resilient setup, feel free to reach out to our support team.
Thank you for your continued trust
DNS and Network Downtime for cigotracker.com and cigopay.com
Początek 15 listopada 2024 03:15 UTC · 0m
Pending
resolved
On November 14th, at approximately 10:15 PM, during a planned network maintenance session, we undertook a migration of core network routing resources. Unfortunately, this migration inadvertently impacted our DNS records, causing DNS Probe errors for some customers attempting to access our services.
To address the issue, our operations team promptly remapped the affected DNS records and adjusted network restriction rules to restore connectivity. The total downtime was less than one hour.
The root cause of this incident was an unforeseen interaction between the migration process and automatic changes to both networking rules and DNS records. While this was an unanticipated challenge, we’re grateful for the swift recovery efforts of our operations team.
The issue has been fully resolved, and measures have been implemented to mitigate the risk of similar occurrences in the future.
If you are still experiencing connectivity issues, we recommend:
1. Allowing additional time for DNS propagation to reach your network.
2. Performing a DNS flush on your network or operating system.
If the issue persists, please contact us at [email protected], and we’ll be happy to assist you.
We sincerely apologize for the inconvenience caused and appreciate your patience and understanding during this time.
Email Consumer Disruption and Queue Management Issue Impacting Notification Service
Początek 11 sierpnia 2024 15:00 UTC · 0m
Pending
resolved
Incident Summary:
On August 11th, a significant issue occurred with the Email Consumer in the Notification Service. The Email Consumer stopped functioning at approximately 8:00 PM EDT, leading to a series of cascading issues within the Notification Service. This incident did not impact the rest of the Cigo Tracker service.
Timeline of Events:
- 8:00 PM EDT: The Email Consumer ceased functioning.
- 8:00 PM - 7:00 AM EDT: Messages began queuing at a significantly reduced rate.
- 7:00 AM - 10:00 AM EDT: The rate at which messages were queued increased substantially.
- 11:00 AM - 2:08 PM EDT: The Email queue reached its capacity, resulting in the rejection of new messages. During this period, both email and SMS message publishing were halted, despite the expectation that only the Email queue should have been affected.
Impact:
- The system experienced a complete halt in both email and SMS message publishing between 11:00 AM and 2:08 PM EDT, affecting communication and potentially leading to delays in message delivery.
Root Cause:
- There is a suspected code-logic error that may be causing incorrect checks on the queues, leading to the stoppage of both email and SMS publishing when the Email queue reaches its capacity.
Actions Taken:
- An alert was sent to the appropriate internal communications channel when the Email Consumer failed. However, due to a configuration issue, the notification was not received by the intended recipients at the time of the failure.
- The issue was remediated at around 2:00 PM EDT by restarting the Email Consumer and fixing the Email queue. Adjustments have been made to the channel configuration to include additional team members and ensure quicker responses in the future.
Next Steps:
- Investigate the root cause of the Email queue failure.
- Determine why reaching the maximum queue size for emails also impacted the SMS queue.
- Ensure that the monitoring and alerting systems are fully operational and that all team members are notified promptly of any critical failures.
- Conduct a thorough review of the queue management logic to identify and rectify any underlying issues.
Conclusion:
We are taking appropriate actions to prevent this issue from recurring. Our team is committed to ensuring the reliability of the Notification Service, and we will continue to monitor and improve our systems to provide the best possible service to our customers.
Azure Outage
Początek 30 lipca 2024 12:24 UTC · 8h 57m
IssuesDrobny incydent
Dotknięte komponenty
Home pageCustomer TrackerDispatch Web PlatformAndroidNotificationsiOSOutbound SMS ServiceOutbound Email ServiceOperator APIMapsPublic APIRouting and Itinerary Optimization
investigating
We are currently experiencing issues accessing our services due to an outage reported by Microsoft Azure.
Microsoft Azure has reported an issue impacting access to the Azure portal and Azure services in general. We suspect that this may be related to general DNS issues affecting their services, which in turn is impacting access to our website and services.
Our team is actively monitoring the situation and working closely with Azure to resolve the issue as quickly as possible.
We will provide updates as soon as we have more information. Thank you for your patience and understanding as we work through this issue.
If you have any questions or need further assistance, please contact our support team at [email protected].
identified
We have identified the issue as being related to Azure’s network infrastructure. Multiple engineering teams at Microsoft are engaged to diagnose and resolve the issue.
Azure's latest update ( https://azure.status.microsoft/en-us/status ):
"We are investigating reports of issues connecting to Microsoft services globally. Customers may experience timeouts connecting to Azure services. We have multiple engineering teams engaged to diagnose and resolve the issue. More details will be provided as soon as possible."
This message was last updated at 13:13 UTC on 30 July 2024.
We will provide updates as soon as we have more information.
Thank you for your patience and understanding as we work through this issue.
monitoring
A fix has been implemented by the Microsoft Azure team. As of their latest update:
"We have implemented networking configuration changes and have performed failovers to alternate networking paths to provide relief. Monitoring telemetry shows improvement in service availability from approximately 14:10 UTC onwards, and we are continuing to monitor to ensure full recovery."
This message was last updated at 14:54 UTC on 30 July 2024.
On our end all connectivity issues seem to be resolved so far, but we are monitoring until Azure confirms full resolution.
monitoring
The latest update from Microsoft Azure indicates that the majority of the issue should now be mitigated:
"An unexpected usage spike resulted in Azure Front Door (AFD) components performing below acceptable thresholds, leading to intermittent errors, timeout, and latency spikes. We have implemented network configuration changes and have performed failovers to provide alternate network paths for relief. Our monitoring telemetry shows improvement in service availability from approximately 14:10 UTC onwards.
As we investigate reports of specific services and regions that are still experiencing intermittent errors, we believe that our network configuration changes have successfully mitigated the impacts of the usage spike, but that these changes are causing some side effects to certain services. We are updating our mitigation approach to minimize these side effects, and applying these following Safe Deployment Practices - beginning in Asia Pacific regions and then expanding in phases.
We will provide an update on our continued mitigation efforts by 19:00 UTC, or sooner if we have progress to share."
This message was last updated at 17:59 UTC on 30 July 2024
You can refer to Azure's latest updates here: https://azure.status.microsoft/en-us/status
resolved
Microsoft Azure has claimed that the issue has been fully resolved.
For details, you can review the complete Azure outage history here for today's partial outage:
Mitigation Statement - Azure Front Door - Issues accessing a subset of Microsoft services (Tracking ID: KTY1-HW8) [ https://azure.status.microsoft/en-us/status/history/ ]
Everything seems to be fully operational on our end now.
If you are still experiencing any issues, please reach out to our support team at [email protected]
SMS and Email Notifications Service Outage
Początek 30 maja 2024 14:18 UTC · 46m
OutagePoważny incydent
Dotknięte komponenty
Outbound SMS ServiceOutbound Email Service
investigating
We are currently investigating an emerging issue with our SMS and Email notifications service. Our team is actively investigating the issue to identify the root cause and implement a solution. We will provide updates as we make progress. We apologize for any inconvenience this may have caused and appreciate your patience as we work to resolve this issue promptly.
identified
We've identified the issue with our messaging service that started at 9:40 AM (Eastern Time). Rest assured, all scheduled and event notifications sent before this time were successfully delivered. Our team is actively working on a fix, and we will post an update as soon as the problem is resolved. Thank you for your patience.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
The incident has been resolved, and SMS and email notification services are now fully restored. We are currently analyzing the issue to prevent it from happening again. Thank you for your patience and understanding.
SMS Notifications Service Outage
Początek 17 maja 2024 14:35 UTC · 45m
OutagePoważny incydent
Dotknięte komponenty
Outbound SMS Service
investigating
We are currently experiencing an outage with our SMS service that began last night. Our team is actively investigating the issue to identify the root cause and implement a solution. We will provide updates as we make progress. We apologize for any inconvenience this may have caused and appreciate your patience as we work to resolve this issue promptly.
identified
The issue has been identified, and a fix has been implemented. We determined that the system stopped sending SMS messages yesterday at 2:28 PM EDT. Our SMS notification system is now operational and is currently processing the backlog at a rate of 300 SMS per second. We will continue to monitor the system to ensure a full recovery.
monitoring
The backlog of SMS notifications has been fully processed, and our SMS service is back to normal operation. We are continuing to monitor the system to ensure everything is working well.
resolved
This incident has been resolved. We identified the root cause of the problem and will work on improving our system's health checks to ensure such issues are identified and mitigated more rapidly in the future. Once again, we apologize for the inconvenience this has caused and appreciate your patience throughout this process.
Homepage Accessibility Issue
Początek 7 maja 2024 14:13 UTC · 6h 22m
IssuesDrobny incydent
Dotknięte komponenty
Home page
investigating
As of 8:55 AM EDT today, we are experiencing an unplanned outage affecting our homepage (https://cigotracker.com). Our technical team is actively investigating the cause of this issue.
Impact:
Please note that this outage does not affect the overall availability of our web application. Users can continue to access our web application through the following link: https://app.cigotracker.com/site/login. The operator applications on Android and iOS, as well as all other components of our web application, are operating normally.
Next Steps:
We are working diligently to identify and resolve the issue. Updates will be provided as more information becomes available.
investigating
We are actively investigating the issue causing delayed loading on the homepage. As a temporary measure, we have redirected traffic from cigotracker.com to the login page of our web application.
Meanwhile, we are thoroughly examining our operational logs and conducting trace analysis to pinpoint the cause of the outage on the landing page.
Thank you for your patience.
resolved
This incident has been resolved.
Cigo Pay Tip Submission Error (April 23rd, 6 PM EDT - April 24th, 1 PM EDT)
Początek 23 kwietnia 2024 22:00 UTC · 0m
OutagePoważny incydent
resolved
On Tuesday, April 23rd, at approximately 6 PM EDT, we implemented a network configuration adjustment in our backend routing rules. This adjustment inadvertently affected Cigo Pay, specifically causing errors when users attempted to submit tip transactions. Although Cigo Pay remained technically online, users encountered difficulties with the final submission of tips.
Reports of this error began to surface on Wednesday morning. Upon investigation, we identified that the network configuration change was the root cause of the issue. We promptly rectified the configuration error by 1 PM EDT on the same day.
We sincerely apologize for any inconvenience this incident may have caused. To prevent similar occurrences in the future, we have implemented enhancements to our monitoring systems. Additionally, we regret the oversight in failing to create an incident report at the time of the problem.
Thank you for your patience and understanding as we worked to resolve this issue. If you have any further questions or concerns, please don't hesitate to reach out.
Unplanned Platform Downtime (Cloud Vendor Outage)
Początek 21 stycznia 2024 06:40 UTC · 6h 44m
OutagePoważny incydent
Dotknięte komponenty
Customer TrackerDispatch Web PlatformAndroidNotificationsiOSOutbound SMS ServiceOutbound Email ServiceOperator APIMapsPublic APIRouting and Itinerary Optimization
investigating
We are currently experiencing extended database downtimes stemming from Microsoft Azure's database instances. This has been occurring intermittently since 9:45 PM (EST), and the issue has been recurring at a higher frequency since 12 AM (EST). Our team is actively investigating the root cause of this service disruption. We appreciate your patience as we work to resolve this issue promptly.
identified
Our ongoing investigation into the connectivity disruption impacting a subset of our database servers and various Azure services has identified an issue on Microsoft's end. The Microsoft Operational Systems Support team has acknowledged the problem and is actively addressing it.
We are diligently awaiting further updates from their team and will keep you informed as soon as new information becomes available. Your patience during this time is sincerely appreciated.
monitoring
It appears that Microsoft Azure has successfully implemented a fix, and our database server connections are now operational.
However, we are currently awaiting official confirmation from the Microsoft Operational Systems Support team to validate that the issue has been fully mitigated.
Thank you for your continued understanding.
monitoring
Our team is actively monitoring our database server instances to guarantee service availability. While things are looking positive with the recent improvements, we are awaiting official confirmation from the Azure team to ensure that the problem has been fully resolved.
resolved
The incident has been successfully resolved. We're now awaiting the Root-Cause Analysis report from Microsoft Azure's team. Once received, we'll compile a post-mortem of the event to provide you with a comprehensive overview.
postmortem
We want to provide you with an update on the January 21st incident that impacted our services. Here's a breakdown of the situation:
**Incident Summary:** On January 20th, 2024, at around 9 PM EST, an internal maintenance process by the Azure OSS team resulted in a configuration change to Azure Resource Manager. Unfortunately, this led to repeated failures of the Azure Resource Manager's node upon startup.
**Root Cause:** The configuration change triggered a negative feedback loop, overwhelming the remaining Azure Resource Manager nodes and causing a rapid drop in availability. This, in turn, affected our backend storage, leading to random failures on data plane API calls. These failures, specifically, disrupted the functionality of our database server, leading to intermittent crashes, particularly during the timeframe of 12 AM to 2 AM.
**Resolution:** The Azure engineering team worked to address the issue, and we’re able to fully resolve it around 4 AM EST on January 21st, 2024.
**Preventive Measures:** To prevent similar incidents in the future, we are closely reviewing our internal processes and working collaboratively with the Azure OSS team to implement additional safeguards.
We sincerely apologize for any inconvenience this may have caused, and we appreciate your understanding as we continue to enhance our systems to provide you with a more reliable experience.
Intermittent platform availability issue
Początek 22 listopada 2023 20:14 UTC · 1h 51m
OutagePoważny incydent
Dotknięte komponenty
Customer TrackerDispatch Web PlatformNotificationsOutbound SMS ServiceOutbound Email ServiceOperator APIMapsPublic APIRouting and Itinerary Optimization
investigating
We are investigating the cause of the intermittent platform availability issue. We will post an update once we have more information.
identified
We experienced connectivity issues between our web platform and one of our caching database services. We have identified a potential cause and implemented corrective measures. The situation appears to have been resolved, and the platform is now stable.
Our team is currently conducting a thorough analysis to understand the root cause.
Further updates will follow as we gather more information.
resolved
We encountered a temporary glitch in domain name resolution, leading to connectivity issues with our caching database and causing a partial service outage. We've taken immediate measures to address the issue, and it also appears to be self-healing.
The disruption occurred from 3:04 PM to 3:16 PM (ET) and briefly from 3:25 PM to 3:28 PM (ET).
We apologize for any inconvenience this may have caused. Rest assured, we're proactively engaging with our cloud hosting vendor to conduct a thorough analysis and prevent future occurrences.
Thank you for your understanding.
Degraded web and mobile app performance due to server connectivity issues
Początek 13 września 2023 14:59 UTC · 5h 43m
OutagePoważny incydent
Dotknięte komponenty
Customer TrackerDispatch Web PlatformNotificationsOutbound SMS ServiceOutbound Email ServiceOperator APIPublic APIRouting and Itinerary Optimization
investigating
We are presently in the process of identifying the underlying reasons for the diminished performance of our web and mobile applications, which appears to be stemming from connectivity problems with our servers. Our initial examination has not uncovered any issues originating from our side. Consequently, we have initiated contact with our cloud hosting provider, Microsoft Azure, to collaborate with their Operations Support System (OSS) team in order to further investigate this matter.
monitoring
We wanted to provide you with an update on the recent server connectivity issues that were impacting our web and mobile app performance. Based on our observations, it appears that the server connectivity issues have been resolved, and our systems are now showing signs of stability.
However, we are still actively monitoring our systems to ensure that everything remains in good working order. We understand the importance of a comprehensive analysis to prevent future occurrences, and to that end, we are eagerly awaiting further information from the Microsoft Azure Operations Support System (OSS) team. We hope that their expertise will help us pinpoint the root cause of the issue, allowing us to take any necessary preventive measures going forward.
We appreciate your patience and understanding as we continue to work on this matter, and we will keep you updated as soon as we receive more information from the Microsoft Azure OSS team. If you have any questions or concerns in the meantime, please don't hesitate to reach out to us.
resolved
We are pleased to inform you that we are closing this incident with the following important notes:
1. Platform Performance: Our platform's performance has returned to its normal levels since approximately 1:28 PM (Eastern Time). We've closely monitored the situation, and the intermittent connectivity errors that were affecting our services have now disappeared.
2. Ongoing Investigation: While the immediate issue has been resolved, we continue to work closely with the Azure Operations Support System (OSS) team to conduct a comprehensive root cause analysis. Our joint efforts aim to identify the underlying reasons for the incident.
3. Future Updates: As soon as we gather more information and insights from our collaboration with the Azure OSS team, we will provide a post-mortem update on this incident. This update will offer a detailed account of our investigation results and outline our plan of action to prevent a recurrence of this issue.
We sincerely apologize for any inconvenience or disruption this incident may have caused to our customers' operations. Our team is committed to ensuring the reliability and performance of our services, and we appreciate your patience and understanding throughout this process.
If you have any further questions or require additional information, please do not hesitate to reach out to us.
postmortem
### Incident Timeline
**Incident Duration:** Approximately 6 hours, from 11:20 UTC to 17:30 UTC on September 13, 2023
### Timeline of Events
#### Incident Identification \(September 13, 2023\)
* **11:20 UTC / 7:20 ET:** Customers started experiencing issues with degraded web and mobile app performance, including higher latency and disconnects.
#### Incident Response \(September 13, 2023\)
* **13:00 UTC / 9:30 ET:** We initiated an internal investigation into the performance degradation and suspected connectivity issues with our servers.
#### Mitigation and Communication \(September 13, 2023\)
* **17:28 UTC / 13:28 ET:** Full platform performance was restored, and intermittent connectivity errors disappeared. However, we continued monitoring the situation.
#### Incident Closure and Ongoing Investigation \(September 13, 2023\)
* **19:00 UTC / 15:00 ET:** The incident was officially closed as platform performance returned to normal levels.
### Root Cause Analysis
The root cause of this incident was identified as a faulty device in Azure Frontdoor. This device continued transmitting traffic from the edge sites for an extended period of time, leading to congestion and packet drops. The prolonged transmission from the faulty device resulted in higher latency, disconnects, and failed service responses.
### Mitigation
Microsoft Azure mitigated the issue by routing traffic away from the problematic device to a healthy one. This action restored normal service operations.
### Preventive Measures
To prevent future occurrences, we are committed to implementing the following measures:
1. **Collaboration with Azure:** We will maintain a strong collaboration with Microsoft Azure's OSS team to ensure a proactive approach to identifying and addressing potential issues promptly.
2. **Traffic Monitoring:** Regular monitoring of traffic patterns will be implemented to detect anomalies and address them swiftly.
3. **Redundancy and Failover:** We will explore redundancy options and failover mechanisms to minimize the impact of similar incidents.
### Conclusion
We sincerely apologize for the inconvenience and disruption this incident may have caused our customers during the impact window of 11:20 UTC to 17:30 UTC \(7:20 ET to 13:30 ET\) on September 13, 2023. We appreciate your patience and understanding throughout the incident resolution process. Our commitment to providing reliable and performant services remains unwavering, and we will continue to work diligently to improve our systems and prevent future incidents.
If you have any further questions or require additional information, please do not hesitate to reach out to us. Thank you for your continued support.
Gateway connection errors
Początek 20 kwietnia 2023 17:36 UTC · 3h 17m
OutagePoważny incydent
Dotknięte komponenty
Customer TrackerDispatch Web PlatformAndroidNotificationsiOSOutbound SMS ServiceOutbound Email ServiceOperator APIMapsPublic APIRouting and Itinerary Optimization
investigating
We are currently investigating a partial outage that is causing our web services to be unavailable intermittently. When our service fails to load, some of our users may observe a 502 gateway error page.
We apologize for the inconvenience and are working with Microsoft Azure's operational support systems team to resolve the problem as soon as possible.
monitoring
It appears that the issue was caused by the infrastructure of our cloud vendor, which has since stabilized.
Our team is awaiting confirmation from Azure's Operations Support Systems (OSS) team and we are closely observing the performance and availability of our web application.
resolved
The incident that took place from 12:48 PM to 1:31 PM EDT (UTC -4) has been successfully rectified and the networking infrastructure has now returned to a stable-state.
While we await Microsoft Azure's OSS team to furnish us with a RCA (Root Cause Analysis) in response to additional logs and information provided by us, we take this opportunity to apologize for any operational disruption that may have been caused by the partial outage.
We are grateful for your patience and understanding in this matter.
Microsoft Azure networking outage
Początek 25 stycznia 2023 07:00 UTC · 0m
OutageKrytyczny incydent
resolved
The Cigo Tracker platform and backend systems that feed the mobile app were impacted by an outage within Microsoft Azure's networking infrastructure. The outage also impacted Microsoft 365, Teams, and Outlook.
Due to this widespread outage, our systems were either fully or partially unavailable from 2:05 AM to 3:45 AM EST (7:05 AM to 9:45 AM UTC).
The root cause was identified by Azure, and mitigated. This seems to have primarily impacted our customers in Europe and in the Middle East during their core hours of operation.
We apologize for any inconvenience this may have caused.
You can read more about the incident (Tracking ID VSG1-B90) here: https://status.azure.com/en-us/status/history/
This outage also made it to numerous online publications:
- https://www.reuters.com/technology/microsoft-teams-down-thousands-users-india-downdetector-2023-01-25/
- https://www.theguardian.com/technology/2023/jan/25/microsoft-investigates-outage-affecting-teams-and-outlook-users-worldwide
- https://techcrunch.com/2023/01/25/microsoft-teams-outlook-service-outage
DNS connectivity issues
Początek 7 września 2022 17:07 UTC · 3h 25m
IssuesDrobny incydent
Dotknięte komponenty
Customer TrackerDispatch Web PlatformAndroidNotificationsiOSOutbound SMS ServiceOutbound Email ServiceOperator APIMapsPublic APIRouting and Itinerary Optimization
identified
Our platform is currently experiencing a partial outage. A portion of our users are fully impacted and are unable to access the platform, for other response times are slow, and the rest of our users are unaffected.
The report from the Microsoft Azure team is as follows (as of 2022-09-07 1:05 PM EDT):
Connectivity issues
We are aware of connectivity issues to the Azure Portal and customers using Azure Front Door. We will provide more information as it is known.
This message was last updated at 17:00 UTC on 07 September 2022
identified
New update from Azure. They are still investigating the issue and we remain on hold:
Starting at 16:10 (UTC) on 07 Sep 2022, customers using Azure Front Door could be experiencing connectivity issues. This could also be impacting customers’ ability to access the Azure Management Portal. We are investigating a spike in traffic as a potential cause. Retries are likely to be successful.
identified
The Microsoft Azure team has confirmed that they identified the potential cause as "a spike in traffic".
They have further clarified the following:
"While we are not currently observing any traffic spikes currently, we are working on remediating the residual impact. We are recovering a number of nodes that are showing intermittent connectivity issues. For customers who are experiencing connectivity issues, retries are likely to be successful. Most customers should be seeing recovery at this stage."
From our monitoring systems, our DNS health seems to be recovering, but there is still a 5% fluctuation that may impact some customers in some regions.
monitoring
Microsoft Azure's team has confirmed that they are working on a full system recovery following a large spike in traffic that disrupted their network infrastructure:
"We are recovering intermittent connectivity issues. Traffic managed by Azure Front Door service is being recovered by systematically going through the regions where we are observing resource impact and enforcing traffic management on the same. Once the recovery process is completed, the service should be able to resume handling traffic normally."
We are monitoring the health of our Front Door and are seeing a steady recovery.
monitoring
Microsoft Azure's team hasn't reported any additional updates, but based on our monitoring, traffic seems to have returned to normal in the last 30 minutes (3:30 PM EDT).
We are standing by for a full resolution confirmation to be confirmed by the Azure team.
Thank you for your patience and understanding.
resolved
The problem on Microsoft Azure's end seems to have been mitigated. Our workloads have returned to normal.
Issues with routing and route optimization
Początek 31 sierpnia 2022 17:15 UTC · 40m
OutagePoważny incydent
Dotknięte komponenty
Routing and Itinerary Optimization
identified
The issue has been identified with one of our third party providers, and we are working with them to resolve it as soon as possible.
monitoring
A fix has been implemented and we are monitoring the results.
monitoring
We are continuing to monitor for any further issues.