We're investigating database performance issues that are causing intermittent errors and long page load times on the loader.io web interface, and delays in processing the load test queue.
monitoring
Test queues have caught up, and web interface and API appear to be functioning as expected now.
resolved
Database performance and test queue processing is back to normal
Web app and test running service unavailable
Начало 2 января 2023 г. в 07:30 UTC · 0m
Pending
resolved
On Jan 2, 2023 the loader.io website, API, and test running services became unavailable due to an expired TLS certificate in a backend service that manages service credentials. The certificate had been renewed, but was not distributed to all servers that needed it.
When the certificate expired, Loader's orchestration system was unable to read credentials for internal connections to databases and other services, and several services failed as a result, including the web interface and the test running & scheduling jobs. Our team was not notified due to a separate failure of alerting systems, and the team was not in the office because of the observance of the New Years Day holiday on the Monday after new years day.
As soon as a team member noticed the outage, service was restored by distributing the renewed certificate to the credential service, and restarting the other failed services.
- Test settings, results, and other account information in the loader.io web interface was inaccessible
- Tests that had been scheduled for Jan 2, 2023 were not run during their scheduled time, and instead would have run after service was restored, early on Jan 3 2023
- Some scheduled tests may have been scheduled twice in error when service was restored, due to retries and delays processing the backlog of tests
We are reviewing our automation and monitoring systems to ensure that critical systems are better automated, and that our team receives alerts promptly!
Main database unreachable
Начало 22 декабря 2021 г. в 12:35 UTC · 16h 35m
OutageКритический инцидент
Затронутые компоненты
Web ConsoleAPITest Queue
investigating
We are investigating an issue connecting to Loader's primary database
monitoring
A replica has been promoted after Loader's primary database failed. Systems are functional and we continue to monitor closely
resolved
This incident has been resolved.
loader.io website down
Начало 15 марта 2021 г. в 09:56 UTC · 13h 0m
Pending
Затронутые компоненты
Web ConsoleAPITest Queue
investigating
We are investigating a problem with our website load balancers.
monitoring
Systems are getting back to normal; we continue to monitor closely
resolved
This incident has been resolved.
Bad Gateway errors
Начало 26 августа 2019 г. в 10:01 UTC · 6h 38m
Pending
Затронутые компоненты
Web ConsoleAPI
monitoring
A recent configuration change caused some of our web servers to stop responding. You may have seen HTTP 502 "Bad Gateway" errors from our web interface and API. The problem has been fixed and we will continue to monitor closely.
resolved
All web traffic is stable, so we are marking this resolved.
test queue delays
Начало 6 февраля 2019 г. в 18:36 UTC · 8h 36m
IssuesНезначительный инцидент
Затронутые компоненты
Test Queue
investigating
We are looking into an issue where some tests are not running
monitoring
The cause of test delays has been identified and fixed. We are continuing to monitor as the test queues catch up.
resolved
This incident has been resolved.
Test delays
Начало 30 ноября 2016 г. в 19:36 UTC · 1h 0m
IssuesНезначительный инцидент
investigating
Some tests are being queued for longer than usual, we are investigating the cause of the slow-down
resolved
Tests should be running normally now
some load tests not running
Начало 27 июля 2015 г. в 15:21 UTC · 25m
IssuesНезначительный инцидент
monitoring
One of our load generation machines stopped responding this morning and has caused a few tests scheduled on it not to run. It has been removed from our fleet of load generators and affected tests should start running. We will be monitoring closely to make sure the issue is resolved.
resolved
Tests are running normally now
tests are being delayed
Начало 29 мая 2015 г. в 11:16 UTC · 59m
IssuesНезначительный инцидент
investigating
We are currently investigating this issue.
investigating
Delayed tests are starting to run, and new tests should run as expected. All systems operational :)
resolved
Tests are now running normally.
https tests not sending correct number of requests
Начало 20 января 2015 г. в 19:36 UTC · 52m
Pending
identified
Some tests against https endpoints are not sending the correct number of requests. We are currently working to resolve this issue.
resolved
We rolled back a recent deploy, https tests should be behaving normally again.
test results not live-updating
Начало 11 декабря 2014 г. в 14:16 UTC · 1h 32m
IssuesНезначительный инцидент
investigating
We are investigating an issue where test results do not update as the test runs. Tests are running and results do appear on page refresh.
resolved
test results are now coming through and updating live as the test runs
Service down
Начало 6 декабря 2014 г. в 03:13 UTC · 50m
OutageКритический инцидент
investigating
We are currently experiencing an unexpected service outage. We are working on resolving the issue.
resolved
We've restored service.
Networking issue
Начало 4 октября 2014 г. в 11:45 UTC · 14m
IssuesНезначительный инцидент
investigating
We are investigating a networking issue that is preventing some tests from running correctly.
resolved
One one our nodes inside of EC2, DNS was resolving EC2 hostnames to public IP addresses instead of internal ones, which prevented some of our internal systems from communicating properly. DNS is resolving properly again.
Database server reboot
Начало 28 сентября 2014 г. в 06:34 UTC · 14m
OutageКритический инцидент
monitoring
The server reboot is going to take longer than anticipated. We should be back around 8:15 AM EDT.
monitoring
We're back up, but our queues are a little backed up, so tests may take a little longer to start for a bit.
resolved
This incident has been resolved.
load generator issues
Начало 1 августа 2014 г. в 17:28 UTC · 54d 6h
IssuesНезначительный инцидент
investigating
We are investigating an issue with our load generators causing a few tests to lose some results
monitoring
load generation is performing normally, but we will keep monitoring to make sure the issue is resolved
resolved
This incident has been resolved.
unplanned maintenance
Начало 4 июня 2014 г. в 20:18 UTC · 12m
OutageКритический инцидент
identified
Our web and API are down right now because of a deploy gone wrong. We're working on getting operational again as soon as possible.
resolved
And we're back. Tests should be verifying and running as usual now.
stalled tests
Начало 23 января 2014 г. в 16:43 UTC · 2h 17m
Pending
investigating
Investigating some stalled tests from the past hour
resolved
Some network issues around 10:20AM EST caused a few tests to stall. Those tests have been aborted and systems operational now.
isolated stalled tests
Начало 19 декабря 2013 г. в 18:51 UTC · 0m
Pending
resolved
A small number of users may have experienced the "preparing screen of death" intermittently over the last 12 hours, where a test shows a preparing message and even the "abort test" button couldn't get you out of it. This was caused by EC2 capacity issues at Amazon, combined with a bug in our handling of that error. A fix for the bug in our code has been deployed, and if you had a test stuck at the preparing screen, we have aborted it for you - instead of the preparing screen, you should now see a message indicating that your test has been aborted. You can run the test again from there.