We're investigating database performance issues that are causing intermittent errors and long page load times on the loader.io web interface, and delays in processing the load test queue.
monitoring
Test queues have caught up, and web interface and API appear to be functioning as expected now.
resolved
Database performance and test queue processing is back to normal
Web app and test running service unavailable
Startede 2. januar 2023 kl. 07.30 UTC · 0m
Pending
resolved
On Jan 2, 2023 the loader.io website, API, and test running services became unavailable due to an expired TLS certificate in a backend service that manages service credentials. The certificate had been renewed, but was not distributed to all servers that needed it.
When the certificate expired, Loader's orchestration system was unable to read credentials for internal connections to databases and other services, and several services failed as a result, including the web interface and the test running & scheduling jobs. Our team was not notified due to a separate failure of alerting systems, and the team was not in the office because of the observance of the New Years Day holiday on the Monday after new years day.
As soon as a team member noticed the outage, service was restored by distributing the renewed certificate to the credential service, and restarting the other failed services.
- Test settings, results, and other account information in the loader.io web interface was inaccessible
- Tests that had been scheduled for Jan 2, 2023 were not run during their scheduled time, and instead would have run after service was restored, early on Jan 3 2023
- Some scheduled tests may have been scheduled twice in error when service was restored, due to retries and delays processing the backlog of tests
We are reviewing our automation and monitoring systems to ensure that critical systems are better automated, and that our team receives alerts promptly!
Main database unreachable
Startede 22. december 2021 kl. 12.35 UTC · 16h 35m
OutageKritisk hændelse
Berørte komponenter
Web ConsoleAPITest Queue
investigating
We are investigating an issue connecting to Loader's primary database
monitoring
A replica has been promoted after Loader's primary database failed. Systems are functional and we continue to monitor closely
resolved
This incident has been resolved.
loader.io website down
Startede 15. marts 2021 kl. 09.56 UTC · 13h 0m
Pending
Berørte komponenter
Web ConsoleAPITest Queue
investigating
We are investigating a problem with our website load balancers.
monitoring
Systems are getting back to normal; we continue to monitor closely
resolved
This incident has been resolved.
Bad Gateway errors
Startede 26. august 2019 kl. 10.01 UTC · 6h 38m
Pending
Berørte komponenter
Web ConsoleAPI
monitoring
A recent configuration change caused some of our web servers to stop responding. You may have seen HTTP 502 "Bad Gateway" errors from our web interface and API. The problem has been fixed and we will continue to monitor closely.
resolved
All web traffic is stable, so we are marking this resolved.
test queue delays
Startede 6. februar 2019 kl. 18.36 UTC · 8h 36m
IssuesMindre hændelse
Berørte komponenter
Test Queue
investigating
We are looking into an issue where some tests are not running
monitoring
The cause of test delays has been identified and fixed. We are continuing to monitor as the test queues catch up.
resolved
This incident has been resolved.
Test delays
Startede 30. november 2016 kl. 19.36 UTC · 1h 0m
IssuesMindre hændelse
investigating
Some tests are being queued for longer than usual, we are investigating the cause of the slow-down
resolved
Tests should be running normally now
some load tests not running
Startede 27. juli 2015 kl. 15.21 UTC · 25m
IssuesMindre hændelse
monitoring
One of our load generation machines stopped responding this morning and has caused a few tests scheduled on it not to run. It has been removed from our fleet of load generators and affected tests should start running. We will be monitoring closely to make sure the issue is resolved.
resolved
Tests are running normally now
tests are being delayed
Startede 29. maj 2015 kl. 11.16 UTC · 59m
IssuesMindre hændelse
investigating
We are currently investigating this issue.
investigating
Delayed tests are starting to run, and new tests should run as expected. All systems operational :)
resolved
Tests are now running normally.
https tests not sending correct number of requests
Startede 20. januar 2015 kl. 19.36 UTC · 52m
Pending
identified
Some tests against https endpoints are not sending the correct number of requests. We are currently working to resolve this issue.
resolved
We rolled back a recent deploy, https tests should be behaving normally again.
test results not live-updating
Startede 11. december 2014 kl. 14.16 UTC · 1h 32m
IssuesMindre hændelse
investigating
We are investigating an issue where test results do not update as the test runs. Tests are running and results do appear on page refresh.
resolved
test results are now coming through and updating live as the test runs
Service down
Startede 6. december 2014 kl. 03.13 UTC · 50m
OutageKritisk hændelse
investigating
We are currently experiencing an unexpected service outage. We are working on resolving the issue.
resolved
We've restored service.
Networking issue
Startede 4. oktober 2014 kl. 11.45 UTC · 14m
IssuesMindre hændelse
investigating
We are investigating a networking issue that is preventing some tests from running correctly.
resolved
One one our nodes inside of EC2, DNS was resolving EC2 hostnames to public IP addresses instead of internal ones, which prevented some of our internal systems from communicating properly. DNS is resolving properly again.
Database server reboot
Startede 28. september 2014 kl. 06.34 UTC · 14m
OutageKritisk hændelse
monitoring
The server reboot is going to take longer than anticipated. We should be back around 8:15 AM EDT.
monitoring
We're back up, but our queues are a little backed up, so tests may take a little longer to start for a bit.
resolved
This incident has been resolved.
load generator issues
Startede 1. august 2014 kl. 17.28 UTC · 54d 6h
IssuesMindre hændelse
investigating
We are investigating an issue with our load generators causing a few tests to lose some results
monitoring
load generation is performing normally, but we will keep monitoring to make sure the issue is resolved
resolved
This incident has been resolved.
unplanned maintenance
Startede 4. juni 2014 kl. 20.18 UTC · 12m
OutageKritisk hændelse
identified
Our web and API are down right now because of a deploy gone wrong. We're working on getting operational again as soon as possible.
resolved
And we're back. Tests should be verifying and running as usual now.
stalled tests
Startede 23. januar 2014 kl. 16.43 UTC · 2h 17m
Pending
investigating
Investigating some stalled tests from the past hour
resolved
Some network issues around 10:20AM EST caused a few tests to stall. Those tests have been aborted and systems operational now.
isolated stalled tests
Startede 19. december 2013 kl. 18.51 UTC · 0m
Pending
resolved
A small number of users may have experienced the "preparing screen of death" intermittently over the last 12 hours, where a test shows a preparing message and even the "abort test" button couldn't get you out of it. This was caused by EC2 capacity issues at Amazon, combined with a bug in our handling of that error. A fix for the bug in our code has been deployed, and if you had a test stuck at the preparing screen, we have aborted it for you - instead of the preparing screen, you should now see a message indicating that your test has been aborted. You can run the test again from there.