We are aware of and investigating an issue with some clients' systems being extremely slow or not functioning at all.
We understand this is a serious issue for our clients, and are working hard to get the issue resolved as quickly as possible. We apologize for the inconvenience and will update this incident as it develops.
investigating
We are still investigating the issue.
But we want to let you know you can still view your inventory by clicking View in the top left and then selecting "View Store Inventory Offline Backup". This report shows all your racked orders so you can still do cash pickups.
monitoring
Response times have returned to normal levels. We'll continue to monitor the system and provide an incident report.
monitoring
Response times are still at normal but at slightly elevated levels. There is a large backlog in the reports queue that should clear overnight. Reports may be inaccurate for the rest of the day (3/5/24).
resolved
This incident has been resolved.
SMRT Service Outage
Début 18 décembre 2023 à 18:55 UTC · 1d 1h
OutageIncident majeur
Composants affectés
POS
investigating
We are aware of an issue which may be preventing some clients & users from accessing SMRT or using certain production functions.
We understand this is a serious issue for our clients, and are working hard to get the issue resolved as quickly as possible. We apologize for the inconvenience, and will update this incident as it develops.
identified
We've identified the issue causing some or all SMRT services to be unavailable, and have begun to implement a series of fixes. You should already see improvements in access, but performance may still be slow. We're continuing our work on a full resolution, and apologize for the continued inconvenience.
monitoring
The SMRT POS is recovering from an outage and is now fully functional. Conveyors should unload within 30 min, and reports should be up to date in a few hours. We sincerely apologize for the disruption this has caused your business.
monitoring
The primary cause of the issue yesterday which led to a partial service outage has been resolved, and most servers are back online. Today, you should find SMRT to be fully functional, although reporting and some non-production systems may still be less responsive than usual. We continue to work to bring our systems back to full capacity, but be advised that some systems may still show some instability. We will continue to prioritize production-related server requests so that your operations may continue, and we will provide further updates as the situation evolves or is resolved.
resolved
Our system is now back to normal. Please let us know immediately if you think you are experiencing anything else.
You can sign up to receive real time alerts about any issues the system may experience at https://status.smrtsystems.com/. Feel free to forward this link to anyone you think should get alerts.
We understand how frustrating delays and/or interruptions are to your normal process and are very sorry you experienced them. We are always focused on providing you with the very best software in the world for your business and will continue to devote even more resources to make sure everything operates as you expect with SMRT.
Sincerely,
SMRT Systems
Report Responsiveness Impact Due to Performance Update
Début 15 décembre 2023 à 18:11 UTC · 2d 20h
IssuesIncident mineur
Composants affectés
Reporting / KPI System
investigating
Dear Customers,
We're currently undergoing a performance update which may affect report responsiveness today. You might experience slower report loading times. These delays should resolve by the end of the day. We apologize for any inconvenience caused.
Thank you for your patience.
Sincerely,
SMRT Support
investigating
We are continuing to investigate this issue.
investigating
This issue has been resolved.
resolved
This incident has been resolved.
7/17/23 - Route Messaging Outage
Début 17 juillet 2023 à 21:00 UTC · 0m
IssuesIncident mineur
resolved
On July 17, 2023, between 5pm-9pm, SMRT experienced a partial outage incident with its messaging component. This incident interfered with outgoing route messages. The proximate cause has been resolved, and we're working to ensure that it doesn't happen again.
SMRT Slowdowns
Début 1 juin 2023 à 20:51 UTC · 3h 48m
IssuesIncident mineur
Composants affectés
POSReporting / KPI System
investigating
We are currently investigating an issue causing SMRT to run unusually slow in some circumstances. The issue may also be affecting data syncing, which may cause some reports not to update promptly. Our developers are actively working on this issue and we'll have it resolved as soon as possible.
resolved
The slowdown issue has been resolved.
8-24-22 AWS Outage Causing Racking and Other Issues
Début 25 août 2022 à 00:48 UTC · 2h 27m
IssuesIncident mineur
Composants affectés
POSReporting / KPI System
monitoring
There's currently an AWS ECS outage causing SMRT queue delays. Users may be experiencing delays with racking, Metal Progetti Assembly, and other areas of the system that rely on our queue system. You can continue to use the system but be aware that there will be a delay for all events that use our queue system.
AWS has identified the issue and is working to resolve it. Once they have resolved the issue our queues will process.
This means that racking events, MP unloads and events, ready order texts, etc. will process after the issue is resolved.
monitoring
There was a recent update from AWS.
Amazon Elastic Container Service - Increased Rates of Insufficient Capacity Errors
5:49 PM PDT We have identified the root cause for the increase in insufficient capacity error rates for launching new Fargate tasks and pods. Customers using ECS with Fargate and EKS with Fargate are impacted, starting at 1:15 PM PDT. We continue to work towards full resolution of the issue, however we are experiencing some delays with full recovery. We are working multiple, parallel paths to make additional capacity available. You will still see some task launches succeeding during this event. Running tasks and pods are not impacted. Customers using ECS with EC2 or EKS with EC2 are not impacted by this issue. We will provide an another update in the next 30 minutes.
Here's the link to their status page (see the second issue Operational issue - Amazon Elastic Container Service (N. Virginia))
https://health.aws.amazon.com/health/status
monitoring
New update from AWS.
6:49 PM PDT We have identified the cause of the decreased capacity and understand why the Fargate task launch success rate is only 70% at this point. We are working on multiple parallel actions to address the underlying issues and have identified one area in particular that should help us make faster progress towards recovery. We have started work on this and have an indication on progress by 7:00 PM PDT. Once we have that progress data we will be able to provide an ETA for recovery. We are also making a change to the rate at which ECS launches tasks as part of ECS services to reduce load on Fargate and to speed up recovery. For customers with prepared and rehearsed plans for moving to a different region should exercise those if they are in a place to do so. Customers can also switch to using the EC2 with ECS and EKS as a mitigation since ECS with EC2 and EKS with EC2 are not impacted by this event.
We have seen a significant decline in the job count in our queues. That said, they are still backlogged and we'll continue to provide updates until the issue is fully resolved.
monitoring
New update from AWS.
7:45 PM PDT We have identified the cause of the decreased capacity and understand why the Fargate task launch success rate is only 70% at this point. Our remediation actions are making slower progress than expected, so we are working on additional actions to further reduce load on Fargate. The work started in the previous update is still progressing but we do not yet have a projected ETA for when it will complete or when we will see recovery. Customers can switch to using the EC2 with ECS and EKS as a mitigation since ECS with EC2 and EKS with EC2 are not impacted by this event.
Our queue backlog is down over 5x from its peak a couple of hours ago. There's still a moderate backlog but most of the queue has been processed. We should be back to normal within the next hour if this pace holds. Thank you for your patience.
resolved
AWS has yet to announce that the issue is resolved. However, our queues are back to normal levels and everything should be functioning as normal.
This is a function of reduced system load due to the late hour and the 70% recovery of AWS ECS. If you experience any issues tomorrow morning please reach out to support. Our dev team will be standing by and monitoring the queues even though we are back to normal.
Thanks for your patience,
SMRT Systems
Phone Service Interruption
Début 28 mars 2022 à 15:27 UTC · 16m
Pending
Composants affectés
POS
identified
We are working with our telephone company and hope to have this issue resolved shortly. Please use our email ticketing for support.
resolved
This incident has been resolved.
Issues with card Processing
Début 1 novembre 2021 à 13:11 UTC · 45m
IssuesIncident mineur
Composants affectés
POS
identified
We are currently having issues with swiped cards using the Magtek Dynapad, we have identified the issue and are currently working on a solution. In the meantime it seems like manual entry of card number will allow you to process the transaction.
resolved
This incident has been resolved.
Google Maps syncing issue
Début 30 septembre 2021 à 15:16 UTC · 3h 44m
IssuesIncident mineur
Composants affectés
POS
investigating
We are currently experiencing issues with our Google Maps integration, this may cause issues optimizing routes and adjusting or adding customer addresses.
identified
We have identified the issue and are working on a resolution, we expect the issue to be resolved in the next 10-15 minutes.
identified
We are continuing to work on this issue, we apologize because it is taking longer than expected.
resolved
This incident has been resolved, we apologize for any inconvenience.
Intermittent outages with the SMRT app.
Début 9 septembre 2021 à 17:01 UTC · 55m
Pending
Composants affectés
POS
investigating
We apologize again and we are investigating the issues with the service going up and down.
investigating
We are continuing to investigate this issue.
resolved
The incident has been resolved
SMRT Outage
Début 9 septembre 2021 à 16:09 UTC · 6m
OutageIncident majeur
Composants affectés
POS
investigating
SMRT Systems is experiencing an outage and we are investigating it now.
resolved
We apologize for the inconvenience but all systems are now operational.
Customer lookup issue occurring again
Début 15 juillet 2021 à 16:11 UTC · 6h 58m
IssuesIncident mineur
Composants affectés
POS
investigating
Our developers are aware of the issue and are researching the cause.
investigating
We are continuing to investigate this issue.
identified
The issue has been identified and a fix is being deployed.
monitoring
A fix has been implemented and the customer search function is now working as expected. We will continue to test the rest of the system and monitor the situation.
resolved
Dear SMRT Customers,
Here's an update on the issues that some of you experienced today.
Yesterday we received an alert saying we were low on storage for our prod6 elasticsearch cluster (this cluster runs search and reporting for some of our customers).
We updated our cluster via a “blue-green deploy” which basically means Amazon Web Services (AWS) spins up 6 more nodes, totalling 12 nodes online at once. All data from the old servers is copied over to the new servers, and then the old nodes are killed.
This was completed without issue yesterday, but today we saw 1 node drop. What happens when a node drops is that AWS starts a new one. The new node then reads data from the rest of the cluster to restore itself.
This is fine when a single node drops because we have 2 copies of all data over the 6 nodes.
The issue today occurred because the new AWS nodes dropped multiple times without any obvious reason. There are a couple of issues with this:
1. The cluster is now operating on fewer nodes, thus slowing down search / reports. In addition to causing random errors
2. Since data is stored twice throughout the cluster, if 2 nodes holding both copies drop at the same time, some data is lost & has to be restored from a backup or be resynced.
The second issue occurred a couple of times today. Our dev teams in both Sweden and San Francisco have been in contact with AWS about the nodes dropping while simultaneously restoring data from backups.
When restoring from backup, the newest set of data from today isn’t synced so we will be running re-syncs shortly. Initially, we did this directly after restoring from our backups, the issue is that the nodes kept dropping & so we had to redo the backup again.
AWS and our developers are currently upgrading our cluster further and they will both be monitoring the status of the cluster throughout the night to ensure there are no further issues tomorrow.
Rest assured you will have no data loss and that this is our number 1 priority. Thank you very much for your patience throughout the day, we understand the pain you felt.
If you do experience anything tomorrow or have any questions/concerns about this issue with the elasticsearch AWS cluster please reach out to support.
Sincerely,
The Entire SMRT Team
Search and report indexing down
Début 15 juillet 2021 à 14:32 UTC · 44m
OutageIncident majeur
Composants affectés
POS
identified
Searching for new customers and report updates are currently unavailable. We are working on fixing it. Searching for existing customers, and by order number is not affected.
identified
We are continuing to work on a fix for this issue.
resolved
This incident has been resolved.
June 25th Incident
Début 25 juin 2021 à 21:08 UTC · 8m
Pending
investigating
Some customers on our P6 Cloud are reporting intermittent speed issues and a white screen. Refreshing is fixing the issue in some cases. Developers are investigating the cause.
resolved
This issue has been resolved. Please contact support if you continue to experience any issues.
Twilio SMS outage
Début 26 février 2021 à 13:44 UTC · 2h 42m
IssuesIncident mineur
Composants affectés
POS
identified
Our SMS provider, Twilio, is currently experiencing an outage.
This is causing some disruptions to SMRT but the system should remain functional.
Please see Twilio's status panel for more real time alerts: https://status.twilio.com/
identified
While the Twilio outage is in progress, SMRT has just disabled all texting as a temporary measure to prefer routing of notifications to email instead. We will enable texting again once Twilio is operating normally. We will send you another update when we do so.
Please follow Twilio's status page: https://status.twilio.com/
resolved
The issue has been resolved and SMRT is operating normally again.
Amazon AWS Outage
Début 25 novembre 2020 à 15:37 UTC · 18h 45m
Pending
Composants affectés
POS
identified
We're investigating elevated reports loading SMRT. It appears our datacenter provider, AWS, is having intermittent issues.
identified
We are continuing to work with Amazon Web Services to resolve this issue. It appears to be a global outage with their infrastructure:
https://downdetector.com/status/amazon/
identified
We're still working on resolving these issues, unfortunately all 3 of our AWS datacenter locations are down in us-east-1 (Virginia).
identified
AWS is still experiencing a severe outage that's affecting many companies including some big names like Roku, Adobe, and Ring to name a few.
Here's an article explaining the outage. Amazon has not given a clear answer as to what caused the issue.
https://techcrunch.com/2020/11/25/amazon-web-services-outage-takes-a-portion-of-the-internet-down-with-it/
identified
Latest from Amazon:
"We continue to work towards recovery of the issue affecting the Kinesis Data Streams API in the US-EAST-1 Region. For Kinesis Data Streams, the issue is affecting the subsystem that is responsible for handling incoming requests. The team has identified the root cause and is working on resolving the issue affecting this subsystem.
The issue also affects other services, or parts of these services, that utilize Kinesis Data Streams within their workflows. While features of multiple services are impacted, some services have seen broader impact and service-specific impact details are below."
monitoring
The system is back up but with degraded performance. We're working with AWS to get us back to normal.
monitoring
AWS Outage Update:
While the system is back online Amazon Web Services is still experiences issues but is on the mend. Expect slower performance than usual for the rest of the day and reports to take longer than normal to update.
Here's the latest from Amazon:
"We continue to work towards recovery of the issue affecting the Kinesis Data Streams API in the US-EAST-1 Region. We also continue to see an improvement in error rates for Kinesis and several affected services, but expect full recovery to still take up to a few hours. For Amazon Cognito, the issues affecting APIs and authentication for user and identity pools has now recovered. For AutoScaling, delays in launching new instances has now recovered, however some scaling operations are still delayed due to delayed CloudWatch metrics. For EventBridge, we have seen partial recovery for the issue affecting delivery of Events.
We are actively working toward full recovery for all affected services, and will continue to provide updates regularly as we have new information to share."
monitoring
AWS Outage Update:
The latest from Amazon:
"We have restored all traffic to Kinesis Data Streams from Internet-facing endpoints, and we are continuing to incrementally restore all requests to Kinesis Data Streams using VPC Endpoints. We are also beginning to observe the incremental recovery of CloudWatch metrics functionality for new incoming metrics, and working towards full recovery. The backlog of metrics will take additional time to populate.
We will continue to keep you updated on our progress."
SMRT knows how hard today was and we thank you for your paitience! This was the worst AWS issue since at least 2017. Interestingly enough we had one customer demoing an offline inventory toll and they were able to handle pickups throughout the outage. This update will be available to all SMRT customers before the new year.
resolved
This incident has been resolved.
SMRT System Outage
Début 17 juillet 2020 à 21:26 UTC · 17m
OutageIncident majeur
Composants affectés
POSReporting / KPI System
investigating
We are experiencing a temporary system outage and are working to resolve the problem ASAP! We will keep you updated as needed through this alert system.
resolved
We apologize for the short interruption but it seems like everything is back up and running at this time.
Outage
Début 9 juin 2020 à 22:42 UTC · 1h 47m
OutageIncident majeur
Composants affectés
POS
identified
One of our cloud service providers, IBM, is down. We’ve contacted their support and are attempting to reroute traffic. Please stand by for further updates.
identified
IBM has provided us the following feedback on the current issue:
“We are actively investigating this issue, and will return back with an update as soon as possible. We are treating this with the utmost priority, and have escalated to management for visibility.”
identified
It is a worldwide IBM outage: https://status.aspera.io/
identified
We are continuing to work on a fix for this issue.
identified
We believe the issue has been resolved, SMRT would like to thank you all for your patience.
resolved
This incident has been resolved.
Elevated response times
Début 14 février 2020 à 16:18 UTC · 9m
IssuesIncident mineur
Composants affectés
POS
investigating
We are currently investigating this issue.
resolved
We identified the issue as a temporary networking problem in the AWS data centers. The system is now 100% operational again.
Pusher having dns problems
Début 5 janvier 2020 à 19:26 UTC · 2h 0m
OutageIncident majeur
Composants affectés
POS
identified
The issue is identified, and a fix is being worked on.