The Mors server has unexpectedly crashed, working on bringing up a replacement.
identified
A server is now booted, working on restoring Sphinx daemons.
monitoring
Sphinx daemons and all other services are now running as expected. Please get in touch if problems are persisting.
resolved
No further issues have cropped up. This incident is resolved.
API requests are failing due to a third-party issue.
์์ 2022๋ 11์ 7์ผ AM 12:33 UTC ยท 1h 57m
Outage์ฌ๊ฐํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
API
identified
We're seeing a failure from an essential third-party API which is stopping our own requests from working - actions that involve reindexing or restarting Sphinx daemons aren't being completed. Search requests are not impacted by this change, though.
The issue has been raised with the third-party vendor, and we're hoping for a swift resolution.
monitoring
The third-party has implemented a fix, errors on our side have stopped but will continue to monitor the situation.
resolved
All functionality has returned to normal. Thanks for your patience, and if you're finding any problems are persisting, please do get in touch.
Robigus primary server has stopped responding
์์ 2021๋ 7์ 14์ผ AM 1:46 UTC ยท 1h 26m
Outage์ค๋ํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Robigus
identified
The primary server for Robigus has stopped responding unexpectedly. Impacted customers have been switched over to the failover server, and should have search queries working again. A new failover will be created shortly.
monitoring
New failover is now running. If any customers are seeing issues persist, please do get in touch.
resolved
All systems are back to normal now.
Instability for Averna, Strenua
์์ 2021๋ 5์ 4์ผ AM 6:30 UTC ยท 0m
Pending
resolved
Both Averna and Strenua servers had some instability in their services, but that has been resolved and everything is back to normal. If any customers continue to be affected, please do get in touch.
Robigus temporarily not responding.
์์ 2021๋ 3์ 3์ผ AM 3:39 UTC ยท 9m
Outage์ค๋ํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Robigus
monitoring
The robigus server had stopped responding, but fixes have been applied, things should be working again, but monitoring the situation to confirm.
resolved
The incident has been resolved, all services are functioning as expected. If you're seeing any lingering issues, please do get in touch!
Robigus was temporarily overloaded.
์์ 2020๋ 9์ 4์ผ PM 9:47 UTC ยท 1h 9m
Outage์ฌ๊ฐํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Robigus
monitoring
Robigus was temporarily overloaded, but is now responding to search requests and API calls. Further investigation is underway to determine the initial cause of the problem.
monitoring
The primary Robigus server has gotten worse, so the failover has been promoted as the new primary, with all affected customers migrated. If you're still seeing issues, please get in touch. A new failover server is now being prepared.
resolved
New failover is standing by, everything should be functioning properly. Please get in touch if that's not the case.
Unexpected Robigus downtime
์์ 2020๋ 8์ 25์ผ PM 10:38 UTC ยท 53m
Outage์ค๋ํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Robigus
identified
The primary server for Robigus-hosted customers has stopped responding unexpectedly, so the failover has now been promoted (and is actively serving search queries). A new failover will be set up.
monitoring
The new failover is in place, the new primary is running as expected.
resolved
Everything's back to normal - but if any customers are still coming across issues, please get in touch.
New Robigus primary is having issuesโฆ
์์ 2020๋ 7์ 2์ผ PM 10:00 UTC ยท 31m
Outage์ค๋ํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Robigus
identified
The primary that was switched over to yesterday is having some teething issues, so I've shifted affected customers over to the new failover instead while things are fixed up. Everything should already be working again, and I'll get the underlying problem fixed.
monitoring
The old primary is now in a happy state (and thus can be considered the new failover). The new primary is serving queries and processing API requests without any issues. Keeping an eye on things.
resolved
No further issues have cropped up, everything's back to working smoothly, so this issue is considered resolved. If anyone's still seeing problems on their side, please do get in touch.
Robigus primary has failed.
์์ 2020๋ 7์ 1์ผ PM 7:02 UTC ยท 35m
Outage์ค๋ํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Robigus
identified
The primary server for Robigus has failed unexpectedly. Impacted customers have now been switched to the failover server, and search requests should be being served again. A new failover is now being prepared.
monitoring
Robigus replacement failover is set up, and search queries and API calls should be operating normally for all customers.
resolved
Everything's continuing to work as expected, this issue is considered resolved. Please get in touch if you're finding any problems persist.
DNS resolution issues
์์ 2020๋ 4์ 14์ผ AM 9:14 UTC ยท 2h 0m
Outage์ค๋ํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Lime
identified
Our DNS provider is having issues from their European node that impacts Lime (the server for our eu-west customers). Searching should remain unaffected, but API requests for indexing and other daemon actions are currently failing. We've been in touch with the DNS provider, and expect there'll be a resolution soon.
identified
We are continuing to work on a fix for this issue.
monitoring
Our DNS provider has fixed the issue and API requests are now working again.
resolved
There have been no further disruptions, this incident is considered resolved.
Mors is not responding.
์์ 2019๋ 11์ 4์ผ PM 1:29 UTC ยท 24m
Outage์ฌ๊ฐํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Mors
investigating
The current More instance is not responding, so a replacement server is being fired up to take its place.
monitoring
The new Mors instance is in place and responding to API calls and search queries.
resolved
This incident has been resolved, and all systems are operating as expected. If you're finding any problems persisting, do get in touch.
Robigus Primary Server has stopped responding.
์์ 2019๋ 9์ 17์ผ PM 10:30 UTC ยท 3h 4m
Outage์ค๋ํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Robigus
identified
We've switched over to the secondary server, so customers should have queries and API calls working again. The cause of the initial issue is not yet clear.
resolved
Robigus has a new secondary server (with the old secondary working smoothly as the new primary). Everything should be responding as expected now - do get in touch if any issues are still cropping up.
Lime Rebooted
์์ 2019๋ 6์ 6์ผ AM 12:31 UTC ยท 4m
Pending
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Lime
monitoring
Lime unexpectedly restarted (possibly due to a broader issue in eu-west-1), but came back up within a few minutes with all Sphinx daemons. Everything is back to normal.
resolved
Lime has returned to being fully operational. If any customers are seeing issues persist, please do get in touch.
Averna has crashed.
์์ 2019๋ 1์ 24์ผ PM 10:05 UTC ยท 22m
Outage์ฌ๊ฐํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Averna
identified
Averna has crashed, currently bringing up a replacement server.
monitoring
Replacement is up and running, with indexing and search requests responding accordingly.
resolved
This issue is now resolved: Averna's returned to operating smoothly. If you're still seeing issues, please get into touch.
Averna is not responding
์์ 2018๋ 11์ 21์ผ PM 5:02 UTC ยท 17m
Outage์ฌ๊ฐํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Averna
investigating
Working on getting it restored.
monitoring
New instance for averna is now up and running.
resolved
Everything's operating as expected again. Please get in touch if you're finding that's not the case for you!
Sphinx daemons not responding on Robigus
์์ 2018๋ 10์ 3์ผ AM 8:49 UTC ยท 0m
Pending
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Robigus
resolved
Earlier today, Robigus stopped responding to search requests (connections to Sphinx daemons).
The part of the Flying Sphinx infrastructure that failed was the Sphinx proxy on your original server - it stopped responding to any TCP requests (though the logs had no suggestion as to why). Clearly, this is a critical part of everything - if the proxyโs down, you canโt connect to your Sphinx daemon at all (and thatโs essential in both searching and regenerating).
I usually get downtime alerts (with a dedicated phone + SMS messages which should wake me up if itโs the middle of the night), but this wasnโt triggered by the proxy failing.
So, Iโve made the following changes:
* If the proxy does not respond to health checks, itโs considered a major server failure, and thus I will get SMS alerts.
* Also, if itโs not responding to health checks, Monit will restart the proxy process, so resolution should be sorted out within a minute.
This is now all in place, but Iโll continue to think through better ways of handling such situations. Iโm very sorry for the downtime!
Carmenta is down
์์ 2018๋ 9์ 21์ผ AM 7:52 UTC ยท 52m
Outage์ฌ๊ฐํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Carmenta
identified
Greatly delayed alert, but Carmenta has unexpectedly restarted itself and crashed halfway through the process. Working on getting a replacement sorted now, expected resolution in 40 minutes.
monitoring
A new server for Carmenta is running, affected customers should find things are working now. Continuing to monitor to confirm that the issue is resolved.
resolved
Everything's now back to normal. Will make sure such fixes don't get delayed so badly again!
If anyone's still seeing problems persist, please get in touch.
Wrong SSL Certificate
์์ 2018๋ 4์ 20์ผ AM 4:41 UTC ยท 1h 40m
Pending
identified
As part of some DNS changes, SSL requests for flying-sphinx.com were briefly returning an invalid certificate (lasting about 10 minutes). This has been resolved, but if you're seeing the problem persist, please do get in touch.
monitoring
Marking this issue as resolved given the certificate is now correct again. Will continue to monitor for any affected services.
resolved
There was one more glitch from the certificate provider, but things have been smooth since, so marking this as resolved.
Issues with requesting actions
์์ 2017๋ 12์ 29์ผ PM 12:26 UTC ยท 18m
Pending
identified
While general server maintenance was happening, a bug was introduced for API actions. This has been fixed, and behaviour has returned to normal. Search daemons/requests were unaffected.
monitoring
Maintenance has finished, behaviour remains as expected.
resolved
This issue is now resolved. If you're seeing any persistent problems, please do get in touch.