强壮出局
- monitoring
Robigus主服务器出乎意料地坠毁,影响了一些客户的搜索和索引. 它已经恢复,运作恢复正常。 目前,已经恢复了对确保稳定的观察.
- resolved
一切都恢复正常了,这起事件被认为是解决了的. 抱歉打断你!
自动翻译自官方事件更新。
51 Flying Sphinx incidents · 2015年4月 — official updates, affected components, duration and resolution details.
Robigus主服务器出乎意料地坠毁,影响了一些客户的搜索和索引. 它已经恢复,运作恢复正常。 目前,已经恢复了对确保稳定的观察.
一切都恢复正常了,这起事件被认为是解决了的. 抱歉打断你!
自动翻译自官方事件更新。
The Mors server has unexpectedly crashed, working on bringing up a replacement.
A server is now booted, working on restoring Sphinx daemons.
Sphinx daemons and all other services are now running as expected. Please get in touch if problems are persisting.
No further issues have cropped up. This incident is resolved.
We're seeing a failure from an essential third-party API which is stopping our own requests from working - actions that involve reindexing or restarting Sphinx daemons aren't being completed. Search requests are not impacted by this change, though. The issue has been raised with the third-party vendor, and we're hoping for a swift resolution.
The third-party has implemented a fix, errors on our side have stopped but will continue to monitor the situation.
All functionality has returned to normal. Thanks for your patience, and if you're finding any problems are persisting, please do get in touch.
The primary server for Robigus has stopped responding unexpectedly. Impacted customers have been switched over to the failover server, and should have search queries working again. A new failover will be created shortly.
New failover is now running. If any customers are seeing issues persist, please do get in touch.
All systems are back to normal now.
Both Averna and Strenua servers had some instability in their services, but that has been resolved and everything is back to normal. If any customers continue to be affected, please do get in touch.
The robigus server had stopped responding, but fixes have been applied, things should be working again, but monitoring the situation to confirm.
The incident has been resolved, all services are functioning as expected. If you're seeing any lingering issues, please do get in touch!
Robigus was temporarily overloaded, but is now responding to search requests and API calls. Further investigation is underway to determine the initial cause of the problem.
The primary Robigus server has gotten worse, so the failover has been promoted as the new primary, with all affected customers migrated. If you're still seeing issues, please get in touch. A new failover server is now being prepared.
New failover is standing by, everything should be functioning properly. Please get in touch if that's not the case.
The primary server for Robigus-hosted customers has stopped responding unexpectedly, so the failover has now been promoted (and is actively serving search queries). A new failover will be set up.
The new failover is in place, the new primary is running as expected.
Everything's back to normal - but if any customers are still coming across issues, please get in touch.
The primary that was switched over to yesterday is having some teething issues, so I've shifted affected customers over to the new failover instead while things are fixed up. Everything should already be working again, and I'll get the underlying problem fixed.
The old primary is now in a happy state (and thus can be considered the new failover). The new primary is serving queries and processing API requests without any issues. Keeping an eye on things.
No further issues have cropped up, everything's back to working smoothly, so this issue is considered resolved. If anyone's still seeing problems on their side, please do get in touch.
The primary server for Robigus has failed unexpectedly. Impacted customers have now been switched to the failover server, and search requests should be being served again. A new failover is now being prepared.
Robigus replacement failover is set up, and search queries and API calls should be operating normally for all customers.
Everything's continuing to work as expected, this issue is considered resolved. Please get in touch if you're finding any problems persist.
Our DNS provider is having issues from their European node that impacts Lime (the server for our eu-west customers). Searching should remain unaffected, but API requests for indexing and other daemon actions are currently failing. We've been in touch with the DNS provider, and expect there'll be a resolution soon.
We are continuing to work on a fix for this issue.
Our DNS provider has fixed the issue and API requests are now working again.
There have been no further disruptions, this incident is considered resolved.
The current More instance is not responding, so a replacement server is being fired up to take its place.
The new Mors instance is in place and responding to API calls and search queries.
This incident has been resolved, and all systems are operating as expected. If you're finding any problems persisting, do get in touch.
We've switched over to the secondary server, so customers should have queries and API calls working again. The cause of the initial issue is not yet clear.
Robigus has a new secondary server (with the old secondary working smoothly as the new primary). Everything should be responding as expected now - do get in touch if any issues are still cropping up.
Lime unexpectedly restarted (possibly due to a broader issue in eu-west-1), but came back up within a few minutes with all Sphinx daemons. Everything is back to normal.
Lime has returned to being fully operational. If any customers are seeing issues persist, please do get in touch.
Averna has crashed, currently bringing up a replacement server.
Replacement is up and running, with indexing and search requests responding accordingly.
This issue is now resolved: Averna's returned to operating smoothly. If you're still seeing issues, please get into touch.
Working on getting it restored.
New instance for averna is now up and running.
Everything's operating as expected again. Please get in touch if you're finding that's not the case for you!
Earlier today, Robigus stopped responding to search requests (connections to Sphinx daemons). The part of the Flying Sphinx infrastructure that failed was the Sphinx proxy on your original server - it stopped responding to any TCP requests (though the logs had no suggestion as to why). Clearly, this is a critical part of everything - if the proxy’s down, you can’t connect to your Sphinx daemon at all (and that’s essential in both searching and regenerating). I usually get downtime alerts (with a dedicated phone + SMS messages which should wake me up if it’s the middle of the night), but this wasn’t triggered by the proxy failing. So, I’ve made the following changes: * If the proxy does not respond to health checks, it’s considered a major server failure, and thus I will get SMS alerts. * Also, if it’s not responding to health checks, Monit will restart the proxy process, so resolution should be sorted out within a minute. This is now all in place, but I’ll continue to think through better ways of handling such situations. I’m very sorry for the downtime!
Greatly delayed alert, but Carmenta has unexpectedly restarted itself and crashed halfway through the process. Working on getting a replacement sorted now, expected resolution in 40 minutes.
A new server for Carmenta is running, affected customers should find things are working now. Continuing to monitor to confirm that the issue is resolved.
Everything's now back to normal. Will make sure such fixes don't get delayed so badly again! If anyone's still seeing problems persist, please get in touch.
As part of some DNS changes, SSL requests for flying-sphinx.com were briefly returning an invalid certificate (lasting about 10 minutes). This has been resolved, but if you're seeing the problem persist, please do get in touch.
Marking this issue as resolved given the certificate is now correct again. Will continue to monitor for any affected services.
There was one more glitch from the certificate provider, but things have been smooth since, so marking this as resolved.
While general server maintenance was happening, a bug was introduced for API actions. This has been fixed, and behaviour has returned to normal. Search daemons/requests were unaffected.
Maintenance has finished, behaviour remains as expected.
This issue is now resolved. If you're seeing any persistent problems, please do get in touch.