Scheduled maintenance
- investigating
We are currently undergoing a routine maintenance, and various services will be impacted. We intend to be back online shortly.
- resolved
Maintenance has been completed
32 Elevio incidents · ноябрь 2014 г. — official updates, affected components, duration and resolution details.
We are currently undergoing a routine maintenance, and various services will be impacted. We intend to be back online shortly.
Maintenance has been completed
On Saturday 27th November there was an incident affecting the Elevio hosted KB which caused it to become intermittently unavailable between 19:10 and 23:30pm UTC. The issue was caused by a resource leak in the KB proxy server, where file descriptors (fd) failed to close after reading a certificate thus rejecting new connections once the fd limit was reached. The issue was caused by a bug in the underlying proxy server software, and was resolved by updating the proxy server software to the latest version. We apologise for this disruption in our service. The issue was particularly difficult to track down since there were no recent updates to the proxy server, and the unfortunate timing of the event (Saturday night / Sunday early morning) meant extra delays in getting the daytime team up to speed
Timeline \(in UTC\):19:10pm: The hosted KB becomes unavailable and service restores without intervention in less than 1 minute. 19:16pm: The hosted KB becomes unavailable and service restores without intervention within 2 minutes. 19:50pm - 20:30pm: Multiple events where the hosted KB becomes unavailable for 1-2 minutes and service restores without intervention. On-call engineer escalates the incident with the backend team. 20:30pm - 22:00pm: The service becomes unavailable and no longer restores. Restarting the servers restores service for short periods of time. 22:15pm: The issue is identified and a temporary fix is deployed. Service resumes as normal 23:30pm: A permanent fix is deployed. During this time, the temporary fix had to be momentarily reverted in order to deploy the new version which caused the service to become unavailable for < 5 mins.
There is currently an issue in the retrieval of articles, we are investigating the issue.
Services have returned to operating as expected. We will continue to monitor.
This incident has been resolved.
On Monday 12 July 2021 at 10AM AEST, we made a change to our main database to improve its stability and performance. We closely monitored the metrics to make sure all systems that depend on it were performing normally, and they were. However, on Tuesday 13 July 2021 at around 12:05AM AEST, the database started degrading, leading to a partial outage. Our failover system kicked in within seconds, and it took 5-10 minutes for most of our services to fully recover and metrics to stabilise. Unfortunately the database gradually degraded again, leading to another partial outage at 4:20AM AEST. It successfully failed over to a backup instance like in the first outage. However, this time around it took our systems 15-20 minutes to fully recover from it. At around 4:50AM AEST everything went back to normal. After investigating the two incidents, we identified the root cause as the change we made on Monday morning. We've decided to roll the database back to its previous configuration and we're cautiously optimistic that the same issue will not happen again.
We are currently investigating this issue.
A fix has been implemented and we are monitoring the results.
This incident has been resolved.
We are currently investigating downtime with a number of AWS services, we will report back when we have further information
The main issues are stemming from major AWS outages, which is bringing down a large portion of the internet at the moment. More information on the current status of AWS can be seen here: https://status.aws.amazon.com/ More details from other major applications being down: https://www.theverge.com/2020/11/25/21719396/amazon-web-services-aws-outage-down-internet
Service is slowly being restored. However, full recovery may take up to a few hours.
A fix has been implemented and we are monitoring the results.
Our main API powering the Assistant has now recovered, however interaction events (article views, searches, module clicks, etc) aren't currently being processed. We continue to investigate the latter.
Our event processing systems are gradually recovering. However, full recovery may take up to a few hours.
All systems are now operating normally.
Hosted knowledge bases experiencing inconsistent downtime.
Related to \(but independant from\) the KB downtime that was experienced on May 14, one of the KB servers experienced an issue whereby excessive caching filled the servers disk capacity causing flow on effects. The issue was confined to only one server, which meant the KB would appear down in some cases, but OK in others after a reloading of the page. To resolve this issue, caching in ares has been removed, as improvements made over time at the API level have rendered the speed improvements received with caching unnecessary and will remove the possibility of this issue from reoccurring in future.
Some customers experienced downtime with their hosted KB
Earlier this morning there was an issue with the knowledge base servers having intermittent troubles displaying. This was due to the servers filling up with log files over time, as well as cached views, and reached a tipping point. We’ve resolved the issues and cleaned up the servers, as well as moved the logging to an external service to prevent this situation from occurring again in the future.
Our CDN provider is currently experiencing issues with our SSL certificate, we're aware of the issue and are working hard to get this resolved ASAP. This effects customers who have the Assistant loaded via Segment, or those still installing V3 (in which case, we recommend upgrading your installation script https://app.elev.io/installation)
This incident has been resolved.
The hosted knowledge base service is currently experiencing issues, we're investigating and will aim to resolve this ASAP.
Hosted knowledge bases appear to be back online, we're currently monitoring the situation.
After monitoring the KB servers for some time, we're confident the issue can be marked as resolved.
We are currently investigating this issue.
Due to a rogue environment variable
We're currently investigating an issue with the API for loading the Assistants home screen.
This incident has been resolved.
We identified an issue affecting one of our analytics database instances which resulted in increased error rates when accessing the home page of the Assistant. We managed to restore the service back to normal after restarting the faulty instance. We are working closely with our cloud provider to identify the root cause to ensure that it does not reoccur.
For a period of approximately 15 minutes, there was an issue with connectivity to our database. This was resolved almost immediately, and steps taken to combat this in future.
We're currently experiencing an issue with the connection between the KB and its API, we're investigating and will update when resolved
This incident has been resolved.
We made some updates to improve the "page not found" messages to be branded, along with other error pages. This update inadvertently set our healthcheck endpoint to throw a 404, which our status page reported as an outage. This is just a note to confirm that during this time, there was actually *no outage*.
We're currently experiencing a server outage with our main API service, we are aware of the issue and working to resolve it. We will have this restored ASAP.
This incident has been resolved.
An update was made to the API early on the weekend (US time) to gather more usage statistics. Things went well initially, but the infrastructure started to buckle under heavy load around 9 AM US EST. As soon as we were notified of the increased error rates by our monitoring systems, we immediately rolled back the change. Unfortunately, during this rollback process, the system subsequently suffered from [thundering herd](https://en.wikipedia.org/wiki/Thundering_herd_problem) due to a huge amount of failed requests being retried concurrently, even with autoscaling on. We decided to double the allowable capacity temporarily to handle the increased load. The system subsequently stabilized a few minutes after that.
There is currently an issue with the sync server, we're looking into the cause. During this period, syncing from third-party content stores may be delayed.
The issue with the sync looks to be resolved, and it is making its way through the backlog now. We'll continue to monitor.
The sync backlog has been cleared, and we'll be keeping an eye on it.
We've had some issues with the article sync server recently. All looks to be OK now, however, there's a large backlog of sync requests to process. Apologies for any inconvenience
The queue has been drained, and syncing should be back to normal now.
We had some downtime with the API, please see the postmortem for more detail.
In the past month, we've had a few unfortunate issues and we'd like to take a moment to document these periods to allay any fears. _In order of oldest to most recent:_ Around a month ago we had some weird behavior on some of the servers. In the end, it was due to the servers running for eight months with no reboot. A quick reboot and all returned to normal. To combat this, we are now doing periodic reboots to rule this out from occurring in future. During the three weeks that followed, we prepped our servers for a new feature release (that wasn’t in public use) which inadvertently caused issues with some customers due to isolated backward compatibility, it only occurred sporadically and was limited to a small set of customers. This was entirely our fault, and with the help of some of those customers that were experiencing the issue, it was promptly resolved. At the same time, we added in additional monitoring to be made aware of similar issues moving foward. To help streamline our reporting infrastructure while we were moving page view counts from one system to another, we modified how we were recording events but failed to load test it enough and it buckled under the load of the full production rollout. We reverted to a previous build until we were able to properly handle the load (and any future load) and this hasn’t been a problem since. Finally, last night, AWS had an outage in one of their subsystems for around 2 hours. We're already running in two data centers, so there was little we could do short of moving our whole infrastructure to a new region. We were in constant contact with AWS during this period to resolve the issue ASAP, but were completely at the mercy of their system. ## In summary In response to the issues that have appeared recently, we’ve put a lot of effort into making our infrastructure more resilient to prevent these and similar issues moving forward. We've also added more granular real-time monitoring to alert us of issues in a more timely fashion. Thanks for your patience and trust during this period. Onward and upward.
In preparation for a coming improvement to access control, as an update is being rolled out to all accounts there will be a short period of downtime as the servers get in sync with each others API. This will resolve itself as the update is rolled out over the coming minutes.
This incident has been resolved.