Główny serwer Robigus rozbił się niespodziewanie, wpływając na wyszukiwanie i indeksowanie dla niektórych klientów. Został odrestaurowany, a operacje wróciły do normy. Obecnie obserwujemy rzeczy, które mają zapewnić stabilność.
resolved
Wszystko wróciło do normy, ten incydent jest uważany za rozwiązany. Przepraszam, że przeszkadzam!
The Mors server has unexpectedly crashed, working on bringing up a replacement.
identified
A server is now booted, working on restoring Sphinx daemons.
monitoring
Sphinx daemons and all other services are now running as expected. Please get in touch if problems are persisting.
resolved
No further issues have cropped up. This incident is resolved.
API requests are failing due to a third-party issue.
Początek 7 listopada 2022 00:33 UTC · 1h 57m
OutageKrytyczny incydent
Dotknięte komponenty
API
identified
We're seeing a failure from an essential third-party API which is stopping our own requests from working - actions that involve reindexing or restarting Sphinx daemons aren't being completed. Search requests are not impacted by this change, though.
The issue has been raised with the third-party vendor, and we're hoping for a swift resolution.
monitoring
The third-party has implemented a fix, errors on our side have stopped but will continue to monitor the situation.
resolved
All functionality has returned to normal. Thanks for your patience, and if you're finding any problems are persisting, please do get in touch.
Robigus primary server has stopped responding
Początek 14 lipca 2021 01:46 UTC · 1h 26m
OutagePoważny incydent
Dotknięte komponenty
Robigus
identified
The primary server for Robigus has stopped responding unexpectedly. Impacted customers have been switched over to the failover server, and should have search queries working again. A new failover will be created shortly.
monitoring
New failover is now running. If any customers are seeing issues persist, please do get in touch.
resolved
All systems are back to normal now.
Instability for Averna, Strenua
Początek 4 maja 2021 06:30 UTC · 0m
Pending
resolved
Both Averna and Strenua servers had some instability in their services, but that has been resolved and everything is back to normal. If any customers continue to be affected, please do get in touch.
Robigus temporarily not responding.
Początek 3 marca 2021 03:39 UTC · 9m
OutagePoważny incydent
Dotknięte komponenty
Robigus
monitoring
The robigus server had stopped responding, but fixes have been applied, things should be working again, but monitoring the situation to confirm.
resolved
The incident has been resolved, all services are functioning as expected. If you're seeing any lingering issues, please do get in touch!
Robigus was temporarily overloaded.
Początek 4 września 2020 21:47 UTC · 1h 9m
OutageKrytyczny incydent
Dotknięte komponenty
Robigus
monitoring
Robigus was temporarily overloaded, but is now responding to search requests and API calls. Further investigation is underway to determine the initial cause of the problem.
monitoring
The primary Robigus server has gotten worse, so the failover has been promoted as the new primary, with all affected customers migrated. If you're still seeing issues, please get in touch. A new failover server is now being prepared.
resolved
New failover is standing by, everything should be functioning properly. Please get in touch if that's not the case.
Unexpected Robigus downtime
Początek 25 sierpnia 2020 22:38 UTC · 53m
OutagePoważny incydent
Dotknięte komponenty
Robigus
identified
The primary server for Robigus-hosted customers has stopped responding unexpectedly, so the failover has now been promoted (and is actively serving search queries). A new failover will be set up.
monitoring
The new failover is in place, the new primary is running as expected.
resolved
Everything's back to normal - but if any customers are still coming across issues, please get in touch.
New Robigus primary is having issues…
Początek 2 lipca 2020 22:00 UTC · 31m
OutagePoważny incydent
Dotknięte komponenty
Robigus
identified
The primary that was switched over to yesterday is having some teething issues, so I've shifted affected customers over to the new failover instead while things are fixed up. Everything should already be working again, and I'll get the underlying problem fixed.
monitoring
The old primary is now in a happy state (and thus can be considered the new failover). The new primary is serving queries and processing API requests without any issues. Keeping an eye on things.
resolved
No further issues have cropped up, everything's back to working smoothly, so this issue is considered resolved. If anyone's still seeing problems on their side, please do get in touch.
Robigus primary has failed.
Początek 1 lipca 2020 19:02 UTC · 35m
OutagePoważny incydent
Dotknięte komponenty
Robigus
identified
The primary server for Robigus has failed unexpectedly. Impacted customers have now been switched to the failover server, and search requests should be being served again. A new failover is now being prepared.
monitoring
Robigus replacement failover is set up, and search queries and API calls should be operating normally for all customers.
resolved
Everything's continuing to work as expected, this issue is considered resolved. Please get in touch if you're finding any problems persist.
DNS resolution issues
Początek 14 kwietnia 2020 09:14 UTC · 2h 0m
OutagePoważny incydent
Dotknięte komponenty
Lime
identified
Our DNS provider is having issues from their European node that impacts Lime (the server for our eu-west customers). Searching should remain unaffected, but API requests for indexing and other daemon actions are currently failing. We've been in touch with the DNS provider, and expect there'll be a resolution soon.
identified
We are continuing to work on a fix for this issue.
monitoring
Our DNS provider has fixed the issue and API requests are now working again.
resolved
There have been no further disruptions, this incident is considered resolved.
Mors is not responding.
Początek 4 listopada 2019 13:29 UTC · 24m
OutageKrytyczny incydent
Dotknięte komponenty
Mors
investigating
The current More instance is not responding, so a replacement server is being fired up to take its place.
monitoring
The new Mors instance is in place and responding to API calls and search queries.
resolved
This incident has been resolved, and all systems are operating as expected. If you're finding any problems persisting, do get in touch.
Robigus Primary Server has stopped responding.
Początek 17 września 2019 22:30 UTC · 3h 4m
OutagePoważny incydent
Dotknięte komponenty
Robigus
identified
We've switched over to the secondary server, so customers should have queries and API calls working again. The cause of the initial issue is not yet clear.
resolved
Robigus has a new secondary server (with the old secondary working smoothly as the new primary). Everything should be responding as expected now - do get in touch if any issues are still cropping up.
Lime Rebooted
Początek 6 czerwca 2019 00:31 UTC · 4m
Pending
Dotknięte komponenty
Lime
monitoring
Lime unexpectedly restarted (possibly due to a broader issue in eu-west-1), but came back up within a few minutes with all Sphinx daemons. Everything is back to normal.
resolved
Lime has returned to being fully operational. If any customers are seeing issues persist, please do get in touch.
Averna has crashed.
Początek 24 stycznia 2019 22:05 UTC · 22m
OutageKrytyczny incydent
Dotknięte komponenty
Averna
identified
Averna has crashed, currently bringing up a replacement server.
monitoring
Replacement is up and running, with indexing and search requests responding accordingly.
resolved
This issue is now resolved: Averna's returned to operating smoothly. If you're still seeing issues, please get into touch.
Averna is not responding
Początek 21 listopada 2018 17:02 UTC · 17m
OutageKrytyczny incydent
Dotknięte komponenty
Averna
investigating
Working on getting it restored.
monitoring
New instance for averna is now up and running.
resolved
Everything's operating as expected again. Please get in touch if you're finding that's not the case for you!
Sphinx daemons not responding on Robigus
Początek 3 października 2018 08:49 UTC · 0m
Pending
Dotknięte komponenty
Robigus
resolved
Earlier today, Robigus stopped responding to search requests (connections to Sphinx daemons).
The part of the Flying Sphinx infrastructure that failed was the Sphinx proxy on your original server - it stopped responding to any TCP requests (though the logs had no suggestion as to why). Clearly, this is a critical part of everything - if the proxy’s down, you can’t connect to your Sphinx daemon at all (and that’s essential in both searching and regenerating).
I usually get downtime alerts (with a dedicated phone + SMS messages which should wake me up if it’s the middle of the night), but this wasn’t triggered by the proxy failing.
So, I’ve made the following changes:
* If the proxy does not respond to health checks, it’s considered a major server failure, and thus I will get SMS alerts.
* Also, if it’s not responding to health checks, Monit will restart the proxy process, so resolution should be sorted out within a minute.
This is now all in place, but I’ll continue to think through better ways of handling such situations. I’m very sorry for the downtime!
Carmenta is down
Początek 21 września 2018 07:52 UTC · 52m
OutageKrytyczny incydent
Dotknięte komponenty
Carmenta
identified
Greatly delayed alert, but Carmenta has unexpectedly restarted itself and crashed halfway through the process. Working on getting a replacement sorted now, expected resolution in 40 minutes.
monitoring
A new server for Carmenta is running, affected customers should find things are working now. Continuing to monitor to confirm that the issue is resolved.
resolved
Everything's now back to normal. Will make sure such fixes don't get delayed so badly again!
If anyone's still seeing problems persist, please get in touch.
Wrong SSL Certificate
Początek 20 kwietnia 2018 04:41 UTC · 1h 40m
Pending
identified
As part of some DNS changes, SSL requests for flying-sphinx.com were briefly returning an invalid certificate (lasting about 10 minutes). This has been resolved, but if you're seeing the problem persist, please do get in touch.
monitoring
Marking this issue as resolved given the certificate is now correct again. Will continue to monitor for any affected services.
resolved
There was one more glitch from the certificate provider, but things have been smooth since, so marking this as resolved.
Issues with requesting actions
Początek 29 grudnia 2017 12:26 UTC · 18m
Pending
identified
While general server maintenance was happening, a bug was introduced for API actions. This has been fixed, and behaviour has returned to normal. Search daemons/requests were unaffected.
monitoring
Maintenance has finished, behaviour remains as expected.
resolved
This issue is now resolved. If you're seeing any persistent problems, please do get in touch.