51 incidente Flying Sphinx înregistrate începând din aprilie 2015, cu actualizări oficiale, componente afectate, durată și informații despre rezolvare.
Serverul principal Robigus s-a prăbuşit pe neaşteptate, afectând căutarea şi indexarea unor clienţi. A fost restaurată, iar operaţiile au revenit la normal. În prezent, trebuie să supraveghem lucrurile pentru a asigura revenirea stabilităţii.
resolved
Totul a revenit la normal, acest incident este considerat rezolvat. Îmi cer scuze pentru întrerupere!
Traducere automată din actualizarea oficială a incidentului.
Unexpected server crash: Mors
A început 13 septembrie 2024 la 16:20 UTC · 4h 32m
OutageIncident critic
Componente afectate
Mors
investigating
The Mors server has unexpectedly crashed, working on bringing up a replacement.
identified
A server is now booted, working on restoring Sphinx daemons.
monitoring
Sphinx daemons and all other services are now running as expected. Please get in touch if problems are persisting.
resolved
No further issues have cropped up. This incident is resolved.
API requests are failing due to a third-party issue.
A început 7 noiembrie 2022 la 00:33 UTC · 1h 57m
OutageIncident critic
Componente afectate
API
identified
We're seeing a failure from an essential third-party API which is stopping our own requests from working - actions that involve reindexing or restarting Sphinx daemons aren't being completed. Search requests are not impacted by this change, though.
The issue has been raised with the third-party vendor, and we're hoping for a swift resolution.
monitoring
The third-party has implemented a fix, errors on our side have stopped but will continue to monitor the situation.
resolved
All functionality has returned to normal. Thanks for your patience, and if you're finding any problems are persisting, please do get in touch.
Robigus primary server has stopped responding
A început 14 iulie 2021 la 01:46 UTC · 1h 26m
OutageIncident major
Componente afectate
Robigus
identified
The primary server for Robigus has stopped responding unexpectedly. Impacted customers have been switched over to the failover server, and should have search queries working again. A new failover will be created shortly.
monitoring
New failover is now running. If any customers are seeing issues persist, please do get in touch.
resolved
All systems are back to normal now.
Instability for Averna, Strenua
A început 4 mai 2021 la 06:30 UTC · 0m
Pending
resolved
Both Averna and Strenua servers had some instability in their services, but that has been resolved and everything is back to normal. If any customers continue to be affected, please do get in touch.
Robigus temporarily not responding.
A început 3 martie 2021 la 03:39 UTC · 9m
OutageIncident major
Componente afectate
Robigus
monitoring
The robigus server had stopped responding, but fixes have been applied, things should be working again, but monitoring the situation to confirm.
resolved
The incident has been resolved, all services are functioning as expected. If you're seeing any lingering issues, please do get in touch!
Robigus was temporarily overloaded.
A început 4 septembrie 2020 la 21:47 UTC · 1h 9m
OutageIncident critic
Componente afectate
Robigus
monitoring
Robigus was temporarily overloaded, but is now responding to search requests and API calls. Further investigation is underway to determine the initial cause of the problem.
monitoring
The primary Robigus server has gotten worse, so the failover has been promoted as the new primary, with all affected customers migrated. If you're still seeing issues, please get in touch. A new failover server is now being prepared.
resolved
New failover is standing by, everything should be functioning properly. Please get in touch if that's not the case.
Unexpected Robigus downtime
A început 25 august 2020 la 22:38 UTC · 53m
OutageIncident major
Componente afectate
Robigus
identified
The primary server for Robigus-hosted customers has stopped responding unexpectedly, so the failover has now been promoted (and is actively serving search queries). A new failover will be set up.
monitoring
The new failover is in place, the new primary is running as expected.
resolved
Everything's back to normal - but if any customers are still coming across issues, please get in touch.
New Robigus primary is having issues…
A început 2 iulie 2020 la 22:00 UTC · 31m
OutageIncident major
Componente afectate
Robigus
identified
The primary that was switched over to yesterday is having some teething issues, so I've shifted affected customers over to the new failover instead while things are fixed up. Everything should already be working again, and I'll get the underlying problem fixed.
monitoring
The old primary is now in a happy state (and thus can be considered the new failover). The new primary is serving queries and processing API requests without any issues. Keeping an eye on things.
resolved
No further issues have cropped up, everything's back to working smoothly, so this issue is considered resolved. If anyone's still seeing problems on their side, please do get in touch.
Robigus primary has failed.
A început 1 iulie 2020 la 19:02 UTC · 35m
OutageIncident major
Componente afectate
Robigus
identified
The primary server for Robigus has failed unexpectedly. Impacted customers have now been switched to the failover server, and search requests should be being served again. A new failover is now being prepared.
monitoring
Robigus replacement failover is set up, and search queries and API calls should be operating normally for all customers.
resolved
Everything's continuing to work as expected, this issue is considered resolved. Please get in touch if you're finding any problems persist.
DNS resolution issues
A început 14 aprilie 2020 la 09:14 UTC · 2h 0m
OutageIncident major
Componente afectate
Lime
identified
Our DNS provider is having issues from their European node that impacts Lime (the server for our eu-west customers). Searching should remain unaffected, but API requests for indexing and other daemon actions are currently failing. We've been in touch with the DNS provider, and expect there'll be a resolution soon.
identified
We are continuing to work on a fix for this issue.
monitoring
Our DNS provider has fixed the issue and API requests are now working again.
resolved
There have been no further disruptions, this incident is considered resolved.
Mors is not responding.
A început 4 noiembrie 2019 la 13:29 UTC · 24m
OutageIncident critic
Componente afectate
Mors
investigating
The current More instance is not responding, so a replacement server is being fired up to take its place.
monitoring
The new Mors instance is in place and responding to API calls and search queries.
resolved
This incident has been resolved, and all systems are operating as expected. If you're finding any problems persisting, do get in touch.
Robigus Primary Server has stopped responding.
A început 17 septembrie 2019 la 22:30 UTC · 3h 4m
OutageIncident major
Componente afectate
Robigus
identified
We've switched over to the secondary server, so customers should have queries and API calls working again. The cause of the initial issue is not yet clear.
resolved
Robigus has a new secondary server (with the old secondary working smoothly as the new primary). Everything should be responding as expected now - do get in touch if any issues are still cropping up.
Lime Rebooted
A început 6 iunie 2019 la 00:31 UTC · 4m
Pending
Componente afectate
Lime
monitoring
Lime unexpectedly restarted (possibly due to a broader issue in eu-west-1), but came back up within a few minutes with all Sphinx daemons. Everything is back to normal.
resolved
Lime has returned to being fully operational. If any customers are seeing issues persist, please do get in touch.
Averna has crashed.
A început 24 ianuarie 2019 la 22:05 UTC · 22m
OutageIncident critic
Componente afectate
Averna
identified
Averna has crashed, currently bringing up a replacement server.
monitoring
Replacement is up and running, with indexing and search requests responding accordingly.
resolved
This issue is now resolved: Averna's returned to operating smoothly. If you're still seeing issues, please get into touch.
Averna is not responding
A început 21 noiembrie 2018 la 17:02 UTC · 17m
OutageIncident critic
Componente afectate
Averna
investigating
Working on getting it restored.
monitoring
New instance for averna is now up and running.
resolved
Everything's operating as expected again. Please get in touch if you're finding that's not the case for you!
Sphinx daemons not responding on Robigus
A început 3 octombrie 2018 la 08:49 UTC · 0m
Pending
Componente afectate
Robigus
resolved
Earlier today, Robigus stopped responding to search requests (connections to Sphinx daemons).
The part of the Flying Sphinx infrastructure that failed was the Sphinx proxy on your original server - it stopped responding to any TCP requests (though the logs had no suggestion as to why). Clearly, this is a critical part of everything - if the proxy’s down, you can’t connect to your Sphinx daemon at all (and that’s essential in both searching and regenerating).
I usually get downtime alerts (with a dedicated phone + SMS messages which should wake me up if it’s the middle of the night), but this wasn’t triggered by the proxy failing.
So, I’ve made the following changes:
* If the proxy does not respond to health checks, it’s considered a major server failure, and thus I will get SMS alerts.
* Also, if it’s not responding to health checks, Monit will restart the proxy process, so resolution should be sorted out within a minute.
This is now all in place, but I’ll continue to think through better ways of handling such situations. I’m very sorry for the downtime!
Carmenta is down
A început 21 septembrie 2018 la 07:52 UTC · 52m
OutageIncident critic
Componente afectate
Carmenta
identified
Greatly delayed alert, but Carmenta has unexpectedly restarted itself and crashed halfway through the process. Working on getting a replacement sorted now, expected resolution in 40 minutes.
monitoring
A new server for Carmenta is running, affected customers should find things are working now. Continuing to monitor to confirm that the issue is resolved.
resolved
Everything's now back to normal. Will make sure such fixes don't get delayed so badly again!
If anyone's still seeing problems persist, please get in touch.
Wrong SSL Certificate
A început 20 aprilie 2018 la 04:41 UTC · 1h 40m
Pending
identified
As part of some DNS changes, SSL requests for flying-sphinx.com were briefly returning an invalid certificate (lasting about 10 minutes). This has been resolved, but if you're seeing the problem persist, please do get in touch.
monitoring
Marking this issue as resolved given the certificate is now correct again. Will continue to monitor for any affected services.
resolved
There was one more glitch from the certificate provider, but things have been smooth since, so marking this as resolved.
Issues with requesting actions
A început 29 decembrie 2017 la 12:26 UTC · 18m
Pending
identified
While general server maintenance was happening, a bug was introduced for API actions. This has been fixed, and behaviour has returned to normal. Search daemons/requests were unaffected.
monitoring
Maintenance has finished, behaviour remains as expected.
resolved
This issue is now resolved. If you're seeing any persistent problems, please do get in touch.