Degraded performance affecting availability of newly submitted session results in US-East-1
开始时间 2026年4月17日 UTC 16:06 · 48m
Issues轻微事件
受影响的组件
Availability of session informationLoading and rendering of reports
investigating
As of 16:00 UTC, we are currently experiencing a slowdown in availability of newly submitted submitted session results in the us-east-1 region. There is no data loss, all saves and submissions are being queued as designed, and the assessments stack remains operational.
Delays may currently be experienced using the Data API to fetch session results, and the Reports API to display results. Other workflows that rely on the synchronous use of Data API results may also be affected.
Learnosity Support and Systems Engineering teams are actively investigating the issue, and will follow on with an update and resolution as soon as possible.
identified
As of 16:20 UTC, Learnosity has increased database performance and we are processing queued sessions rapidly.
Learnosity Support and Systems Engineering teams are continuing to actively investigate the issue, and will follow on with an update and resolution as soon as possible.
monitoring
As of 16:25 UTC, we have resolved the slowdown in processing session results and all sessions are being scored normally. We will continue to monitor for a further 30 minutes before closing this incident report.
Learnosity Support and Systems Engineering teams are continuing to actively investigate the issue, and will follow on with an update and resolution as soon as possible.
resolved
As of 16:55 UTC, we are have resolved the degraded performance issue affecting availability of session data in us-east-1.
Learnosity Support and Systems Engineering teams will follow up with a post mortem once we have completed root cause analysis and finalized any next steps or preventative measures required.
Please reach out if you have any questions or concerns.
postmortem
### Affected Systems and Regions
On 2026-04-17, Learnosity experienced a service degradation impacting availability of newly submitted session results for a small subset of customers in US-East-1. The issue began at approximately 13:54 UTC and was resolved at 16:24 UTC. The total duration of customer impact was approximately 2.5 hours.
### Investigation
Service degradation was detected following the accumulation of submitted sessions in the async scoring queue, This was caused by elevated CPU utilization when heavy load coincided with a background data migration task that had been running for several days on the affected EC2 instance. This reduced message processing throughput, leading to queueing of sessions.
### Resolution
Learnosity CloudOps engineers responded by scaling session processing node EC2 instances, and corresponding database proxies, to maximum allowable capacity. The migration task was also stopped, which further reduced CPU pressure and allowed message processing throughput to recover. The session scoring backlog was fully cleared by 16:24 UTC.
### Prevention
Learnosity is implementing the following measures to mitigate:
* Pre-flight health checks and safeguards are being introduced to prevent migration tasks from running during busy periods of production usage.
* Enhanced monitoring is being applied to detect backlog growth and database contention, and apply remedial action more quickly
Degraded performance affecting Live Progress Report in US-East-1
开始时间 2026年4月16日 UTC 14:08 · 1h 16m
Issues轻微事件
受影响的组件
Live Progress (Live Activity by User) report
investigating
As of 13:30 UTC, we are currently experiencing slow downs affecting the Live Progress report in the us-east-1 region.
Atypical use of eventbus may also be affected, such as custom reports/implementations.
Learnosity Support and Systems Engineering teams are actively investigating the issue, and will follow on with an update and resolution as soon as possible.
resolved
As of 14:30 UTC, the Live Progress report event degradation issue in the us-east-1 region has been resolved. Additional capacity processed the event load and these systems are operating normally.
Brief additional latency during this scaling may have been introduced for users of Learnosity's premium Firehose feature, but we've seen no evidence that this affected customers. This, too, was resolved once scaling was complete.
Learnosity Support and Systems Engineering teams will follow up with a post mortem once we have completed root cause analysis and finalized any next steps or preventative measures required.
Please reach out if you have any questions or concerns.
postmortem
### Affected Systems and Regions
On 2026-04-16, Learnosity experienced a service degradation impacting the Live Progress report for a small subset of customers in the us-east-1 region. The issue began at approximately 13:12 UTC and was resolved at 14:40 UTC. The total duration of customer impact was approximately 88 minutes.
### Investigation
The issue was detected following elevated error rates on the load balancer serving the eventbus service. Investigation determined that an unhealthy condition within the EC2 instances led to elevated CPU utilization and memory pressure within the Auto Scaling Group, resulting in application instability and repeated service restarts. Inconsistent recovery caused traffic to concentrate on a subset of hosts, further elevating error rates. A secondary effect of the instability was increased request pressure on downstream dependencies.
### Resolution
Service was restored by stabilizing the affected EC2 instance and restoring consistent application availability across the Auto Scaling Group. Load distribution normalized once all instances returned to a healthy state.
### Prevention
Learnosity is implementing the following measures to mitigate:
* Improve service startup and dependency handling to ensure consistent recovery behavior
* Review resource thresholds to reduce the likelihood of similar instability
* Enhance monitoring to detect and respond more quickly to similar conditions
Issue affecting assessment delivery in VA (us-east-1)
开始时间 2026年2月4日 UTC 18:34 · 40m
Outage重大事件
受影响的组件
Loading and rendering of Item/Activity Edit viewAMER Assessment (Summary)AMER Analytics (Summary)Loading and rendering of Question Edit viewAMER Authoring (Summary)Saving of student responsesLoading and rendering of assessment UILoading and rendering of Items/Questions/FeaturesLoading and rendering of reports
identified
As of 18:30 UTC, we are currently experiencing an issue affecting assessment delivery in VA (us-east-1).
Learnosity Support and Systems Engineering teams are actively investigating the issue, and will follow on with an update and resolution as soon as possible.
monitoring
As of 18:52 UTC, the Learnosity System Engineering team has put a fix in place. We will continue to monitor for some time to ensure a stable recovery, but all services are operational again.
Learnosity Support and Systems Engineering teams will continue to actively monitor the issue, and will follow on with an update and resolution as soon as possible.
resolved
As of 19:12 UTC, traffic has returned to normal and everything remains stable. We have resolved the incident affecting assessments in VA (us-east-1).
Learnosity Support and Systems Engineering teams will follow up with a post mortem once we have completed root cause analysis and finalized any next steps or preventative measures required.
Please reach out if you have any questions or concerns.
postmortem
### **Affected Systems and Regions**
On 2026-02-04, Learnosity experienced an availability incident impacting Questions API access for customers in the AMER region. The issue began at approximately 18:02 UTC and was formally declared resolved at 18:51 UTC. The total duration of customer impact was approximately 49 minutes.
### **Investigation**
The issue was first detected when the Items API attempted to initialize the dependent Questions API. The incident was traced to a misconfiguration within the content delivery network servicing the AMER region, causing a SSL certificate error..
During a certificate management task, the distribution responsible for serving [_questions.learnosity.com_](http://questions.learnosity.com) was configured with an SSL certificate that did not match the _\*.learnosity.com_ domain nomenclature. This caused browsers and client systems to reject connections to the affected subdomain.
### **Resolution**
Once the issue was identified, the response team updated the affected distribution to use the correct SSL certificate. Service was fully restored after the change was completed.
### **Prevention**
Learnosity is strengthening certificate deployment and validation processes to ensure that domain mappings are verified prior to rollout. Additional safeguards have been added to prevent cross-domain certificate assignments and to improve monitoring for sudden traffic drops that may indicate certificate-related failures.
Issue affecting availability of recent session data for a minority of customers in the US-East-1 region.
开始时间 2026年1月12日 UTC 17:06 · 2h 4m
Pending
受影响的组件
Availability of session informationUpdating session response scoresFirehoseLoading and rendering of reports
investigating
As of 4:30 UTC, we are currently investigating an issue that may result in temporary session unavailability for a small number of US-only customers. We can confirm that no session data has been lost, and we are working to restore the affected sessions as quickly as possible.
Learnosity Support and Systems Engineering teams are actively investigating the issue, and will follow on with an update and resolution as soon as possible.
identified
As of 5:30 UTC, we've identified and corrected a load balancing configuration affecting database reads and writes for a very small number of US-only customers with consumers in the US-East-1 region.
- All new sessions will behave as expected.
- Read requests seeking data on sessions created prior to 23:00 UTC yesterday will also now behave correctly.
- We are working on restoring sessions recently affected
While all APIs continue to function normally, this update identifies service areas that may temporarily have limited access to recently saved or submitted sessions.
Learnosity Support and Systems Engineering teams are continuing to actively investigate the issue, and will follow on with an update and resolution as soon as possible.
resolved
As of 7:00 UTC, we've concluded our monitoring and confirmed that all new requests are behaving as expected. We are closing this record and will continue to work with affected customers directly to confirm that all recent sessions are accessible.
Please reach out if you have any questions or concerns.
AWS outage continuing to have some impact on Learnosity services in VA (us-east-1)
开始时间 2025年10月20日 UTC 14:53 · 6h 42m
Outage重大事件
受影响的组件
Branching/Adaptive servicesAMER Assessment (Summary)Creation of report data setsAMER Analytics (Summary)Availability of session informationProcessing and availability of rich student responses (audio/images/files)Scoring endpointSelf-Hosted Adaptive servicesUpdating session response scoresCreation of sessionsAMER Data Centric (Summary)Creation of Item PoolsFirehose
monitoring
Learnosity services in VA are continuing to be partially impacted by the major AWS outage.
AWS first reported the incident at 7:11 UTC As of 14:50 UTC they are continuing to work on its full resolution, but have again set their impact level to "Degraded".
The AWS Health Status dashboard where updates are being posted by the AWS team can be found here: https://health.aws.amazon.com/health/status
Overall, the error rate has fallen and has stayed below what we saw earlier in the day and we have been successful in slowly scaling up the infrastructure, but we're still seeing some issues stemming from the the ongoing AWS outage as we attempt to bring online some new instances, and we have begun to see an uptick in errors.
These have included errors associated with Data API, Feedback Aide, and the delivery of adaptive assessments, delays in scoring and video encoding, as well as other smaller spikes across different systems.
While the overall error rate remains low, Learnosity Support and Systems Engineering teams are actively monitoring the outage and manually intervening where possible to mitigate the issues we're facing.
In spite of our best efforts it is likely that the APIs may not be fully operational in VA until AWS's service is fully restored to normal operation, but we are taking every step within our power to try to minimize the disruption to our services.
investigating
As of 15:25 UTC, we are seeing increasing error rates.
Data API is the most affected service at the moment and we are seeing an increase in scoring delays, but we are continuing to score incoming messages to the best rate our current capacity allows.
AWS is still reporting a severity of "Degraded", but now with an increased count of 81 affected services.
Learnosity Support and Systems Engineering teams are continuing to monitor the issue, and will continue our mitigation efforts wherever possible.
monitoring
As of 19:10 UTC, we are beginning to see errors rates diminish and some Learnosity services are beginning to scale again.
AWS is still reporting a severity of "Degraded" with problems to 104 of their services in us-east-1.
Learnosity Support and Systems Engineering teams are continuing to monitor and manually intervene when possible to mitigate the effects of the outage.
monitoring
As of 20:00 UTC, the error rate has continued to drop, systems remain scaled and the backlogs of scoring messages continues to drop.
AWS is still reporting 109 impacted services and a severity of "Degraded"
Learnosity Support and Systems Engineering teams are continuing to monitor the issue and are manually adjusting the infrastructure where needed to ensure stability.
monitoring
As of 19:30 UTC, the error rate has fallen to 0.1% and continues to fall, we continue to steadily process the backlog of scoring messages.
As of their latest update AWS has downgraded the severity of outage to "Impacted" but still list 109 services as affected.
Learnosity Support and Systems Engineering teams are continuing to actively monitor the issue, and will follow on with an update and resolution as soon as possible.
monitoring
As of 20:40 UTC, the error rate remains low, the infrastructure is scaled to meet demand, and the scoring queues are empty and new submissions are being processed as they come in.
All services are operational and we are processing the remaining backlog of messages from the incident.
AWS still lists 93 affected services and a severity of "Impacted".
Learnosity Support and Systems Engineering teams are continuing to monitor the issue, but at the moment all systems are fully operational.
resolved
As of 21:30 UTC, all systems keep operating normally, so we are resolving the status page about the AWS outage affecting Learnosity Services in the VA (us-east-1) region.
Learnosity Support and Systems Engineering teams will continue to monitor the systems closely while AWS continues to work on the issues on their side.
Please reach out if you have any questions or concerns.
AWS Outage Impacted some Learnosity Services
开始时间 2025年10月20日 UTC 10:11 · 1h 12m
Pending
受影响的组件
Support ticketing systemProcessing and availability of rich student responses (audio/images/files)
monitoring
AWS is going through late-stage recovery of a major outage in US-East-1 bringing down most of their services. Learnosity APIs were stable and unaffected in isolation, but some impact cascaded to Learnosity (such as availability of assets from CloudFront). The full extent of the incident is unknown at this time.
NOTE: Up until the time of this writing (10:05 UTC), this included impact to sign-in services that prevented us from accessing our StatusPage. We immediately posted banner notifications in our Support Portal, and those banners will remain visible until we've been able to verify that service has been restored for sufficient time.
Incident History:
As of 7:00 UTC, AWS experienced a major outage. While Learnosity APIs were fully operational, some Learnosity services use AWS and were affected. These included access to our Support ticketing system (the Learnosity-hosted Help guides were unaffected), and availability of CloudFront-hosted assets, but other systems may have been impacted as well.
As of 10:00 UTC, AWS has restored services and access to Learnosity-adjacent features is becoming available relatively quickly.
Learnosity Support will monitor AWS for a further 30 minutes before calling this issue resolved. We will follow on with an update and resolution at that time, or sooner.
resolved
As of 11:20 UTC, AWS services appear to have been stable for approximately an hour. Learnosity APIs--including services affected by the AWS outage (such as CloudFront)--appear to be fully operational. At this time we are labeling this issue resolved.
Learnosity Support and Systems Engineering teams will continue to investigate impact from this incident and follow up with a post mortem once we hear from AWS.
Please reach out if you have any questions or concerns.
Partner Service Outage May Disrupt Learnosity API Initialization in All Regions
开始时间 2025年10月2日 UTC 21:34 · 1h 9m
Pending
受影响的组件
Loading and rendering of Items/Questions/FeaturesLoading and rendering of Question Edit viewLoading and rendering of Question Edit viewLoading and rendering of Items/Questions/FeaturesLoading and rendering of Items/Questions/FeaturesLoading and rendering of reportsLoading and rendering of Question Edit viewLoading and rendering of reportsLoading and rendering of reports
investigating
As of 21:15 UTC, Learnosity APIs are being impacted by a temporary outage from our Math partner, Desmos. Authoring or Editing content with Desmos questions or features, or initializing assessments that contain Desmos questions or features, can prevent the Author API or Items API from initializing. Specific reports that display question content, such as the Session Detail by Item report, can also fail to initialize
Learnosity APIs are not otherwise impacted, so content that does not include Desmos features are not affected. Until we learn more, we will list APIs that may experience knock-on effects from this temporary Desmos outage. However, as most customers will not experience any issues, we will not list these APIs as down.
Desmos and Learnosity Support and Systems Engineering teams are actively investigating the issue, and will follow on with an update and resolution as soon as possible.
monitoring
As of 22:12 UTC, Desmos services have been restored and Learnosity APIs are initializing without issue.
Desmos and Learnosity Support and Systems Engineering teams will continue to monitor for a further 30 minutes before declaring this issue resolved. We will follow on with an update and resolution as soon as possible.
resolved
As of 22:42 UTC, all Learnosity APIs are continuing to function as expected in all regions, and we are declaring this issue resolved.
Learnosity Support and Systems Engineering teams will follow up with a post mortem once we have completed root cause analysis and finalized any next steps or preventative measures required.
Please reach out if you have any questions or concerns.
Issue affecting availability of recently submitted session analytics in US-East-1
开始时间 2025年8月19日 UTC 16:11 · 5h 21m
Issues轻微事件
受影响的组件
Availability of session information
investigating
As of 3:40 UTC, we are experiencing degraded performance in our analytics stack affecting the availability of session data in US-East-1 for a subset of customers.
This is affecting the Reports API and Data API. Neither authoring nor assessment stacks are affected, and there is no data loss. Submitted sessions are queueing for processing.
Learnosity Support and Systems Engineering teams are actively investigating the issue, and will follow on with an update and resolution as soon as possible.
investigating
As of 4:30 UTC, we are continuing to investigate degraded performance in our Data and Reports APIs.
Only availability of recently submitted session data is affected. Historic session data, as well as all other API stacks, remain unaffected. New submissions are persisting correctly with no data loss and these submission are being queued for scoring.
Learnosity Support and Systems Engineering teams are continuing to actively investigate the issue, and will follow on with an update and resolution as soon as possible.
identified
(Note: We're correcting cited UTC times to the 24 hr format and will include both forms in this update only.)
As of 5:30pm/17:30 UTC, we've identified a possible contributing cause for the degraded performance in our analytics stack.
Learnosity Support and Systems Engineering teams are continuing to actively investigate the issue, and will follow on with an update and resolution as soon as possible.
identified
As of 18:30 UTC, initial remediation efforts have tripled the rate of queued session processing and we are continuing to work toward a full resolution.
Access to recently submitted session results via the Data API and Reports API remains the only affected part of the Learnosity ecosystem. New submissions continue to be safely queued for scoring while the degraded performance remains.
Learnosity Support and Systems Engineering teams are continuing to actively investigate the issue, and will follow on with an update and resolution as soon as possible.
identified
As of 19:30 UTC, we are now processing queued sessions rapidly and more than half of the backlog has already cleared..
Learnosity Support and Systems Engineering teams are continuing to actively investigate the issue, and will follow on with an update and resolution as soon as possible.
identified
As of 20:30 UTC, the scoring queue backlog is almost empty and new sessions will soon be processed without delay.
Learnosity Support and Systems Engineering teams are continuing to actively investigate the issue, and will follow on with an update and resolution as soon as possible.
monitoring
As of 21:00 UTC, all sessions have been cleared from the scoring backlog queue and all submissions are being processed normally.
Learnosity Support and Systems Engineering teams will monitor this situation for a further 30 minutes before calling it resolved.
resolved
As of 21:30 UTC, we are have resolved the issue affecting availability of session data in US-East-1.
Learnosity Support and Systems Engineering teams will follow up with a post mortem once we have completed root cause analysis and finalized any next steps or preventative measures required.
Please reach out if you have any questions or concerns.
postmortem
### Affected Systems and Regions
On 19 August 2025, Learnosity experienced degraded performance in our analytics stack, affecting session data availability in US-East-1 for a subset of customers. This affected the Reports API and Data API, with no other stacks affected, and no data loss.
### Investigation
Monitoring detected a rapid increase in unprocessed and retried messages, along with elevated lock contention in the sessions database. The root cause was traced to a customer implementation issue generating an extraordinarily high number of submissions. This drove excessive retries, magnifying actual traffic volume.
The use of time-ordered v7 UUIDs for session IDs, normally handled without issue, became problematic under this contention. Uniqueness checks on each session ID required more resources and triggered a succession of temporary deadlocks. These deadlocks would usually self-resolve, but the amplified traffic prevented recovery, turning a minor issue into a sustained queue blockage.
### Resolution
Once the issue was identified, Learnosity moved the customer to a dedicated, isolated sync queue, preventing cross‑tenant impact while we investigated. We applied targeted rate limits for the isolated service to protect the database, and drained the backlog. Where safe, long‑running queries were terminated to free locks and allow forward progress.
To support faster diagnosis, Learnosity enabled detailed deadlock logging and expanded metrics around message retries, abandonment, and per‑session activity. Learnosity also worked with the customer to adjust implementation settings, reducing combined saves and submits by two orders of magnitude. Session IDs were also switched to v4 UUIDs which simplified uniqueness checks further preventing deadlocks.
Immediately after these changes were put into use, queues began to rapidly recover, and normal processing resumed. Most sessions for the subset of affected customers saw short delays, while the most significantly delayed session took ~6 hours before final persistence. Throughout, we identified no data loss.
### Prevention
To prevent recurrence, we are:
* Implementing targeted load tests and contention simulations to replicate high-parallelism patterns.
* Reviewing customer identifier schemes for session IDs and auditing usage across all customers \(initial checks confirm none of our Top 50 customers currently use v7 UUIDs\).
* Analyzing adoption of per-tenant fair-use queues \(or equivalent fair-share policies\) to cap burst throughput from a single tenant and protect shared infrastructure.
Possible Issue affecting manual updates to session data in US-East-1 (VA) region
开始时间 2025年8月8日 UTC 13:55 · 24m
Pending
受影响的组件
Updating session response scores
investigating
As of 13:40 UTC, we began investigating a possible issue affecting manual update jobs to session response scores, sessions statuses, and session metadata in the US-East-1 (VA) region. We saw a brief buildup in the queue for these jobs, the queue processed as expected, and is now empty. However, we're investigating further, out of an abundance of caution, to see if there are any unexpected causes for the short queue buildup.
Learnosity Support and Systems Engineering teams are actively investigating the issue, and will follow on with an update and resolution as soon as possible.
investigating
As of 13:55 UTC, all manual update jobs have been executed and the jobs queue is empty. Requests to manually update session response scores, session statuses, and session metadata should execute as expected. We are still investigating possible reasons for the blip.
Learnosity Support and Systems Engineering teams are continuing to actively investigate the issue, and will follow on with an update and resolution as soon as possible.
resolved
As of 14:10 UTC, we've found no negative behavior in the manual update processes and attribute the short burst of jobs to uncharacteristic testing of the new manual grading workflow. The jobs queue functioned as expected, ensuring no loss of requests, and the jobs were processed in a timely fashion.
Learnosity will continue to monitor this activity for a further 60 minutes, but we are calling this resolved at this time. We apologize for the unwarranted alert.
Please reach out if you have any questions or concerns.
postmortem
At 13:40 UTC on 8 August 2025, we noticed a brief backup of asynchronous jobs affecting manual session updates in the US-East-1 \(VA\) region. Possible jobs including manual updates to response scores, session statuses, and session metadata. The queue mechanism worked as expected, and all requests were processed in short order. Out of an abundance of caution, we raised an alert while systems remained operational, but found nothing of concern during our investigation. We hope the unwarranted alert was not a significant distraction.
Issue affecting availability of session information in US-East-1
开始时间 2025年3月3日 UTC 13:48 · 2h 6m
Outage重大事件
受影响的组件
Availability of session informationLoading and rendering of reports
investigating
As of 13:45 UTC, we are experiencing delays in availability of session information in US-East-1
Learnosity Support and Systems Engineering teams are actively investigating the issue, and will follow on with an update and resolution as soon as possible.
investigating
As of 14:00 UTC, we are still investigating an issue affecting the availability of session information in US-East-1. Both Data API and Reports API are experiencing extended delays and intermittent failures in returning session results. Authoring and Assessment APIs appear to be unaffected at this time and all session submissions are being successfully queued.
Learnosity Support and Systems Engineering teams are continuing to actively investigate the issue, and will follow on with an update and resolution as soon as possible.
investigating
As of 15:00 UTC, we are continuing to investigate the issue. Data and Reports API are still impacted.
Assessment and Authoring APIs remain unaffected. Assessment submissions are being queued and will be processed as the system works through any backlog queue. Note that if the Data API or Reports API is used in a preliminary/synchronous assessment delivery workflow, this will likely have a knock-on affect when initializing assessments. Initializing any assessment API directly is unaffected at present.
Learnosity Support and Systems Engineering teams are continuing to actively investigate the issue, and will follow on with an update and resolution as soon as possible.
monitoring
As of 15:15 UTC, all services have been restored. All queued messages have been processed and session information should now be available. Users of the Learnosity Firehose service will still see slight delays as messages are sent out for recently processed events.
Learnosity Support and Systems Engineering teams are continuing to actively investigate the issue, and will follow on with an update and resolution as soon as possible.
resolved
As of 15:45 UTC, after a further 30 minutes of monitoring, we are resolving this issue. All services remain fully operational.
Learnosity Support and Systems Engineering teams will follow up with a post mortem once we have completed root cause analysis and finalised any next steps or preventative measures required.
Please reach out if you have any questions or concerns.
postmortem
### Affected Systems and Regions
On March 3, 2025, we experienced a temporary service slowdown impacting the availability of session information in the US-East-1 region. The issue began at 13:19 UTC and was resolved by 15:15 UTC. The Data API and Reports API experienced delays and intermittent failures when returning session results. Authoring and Assessment APIs were unaffected with no data loss.
### Investigation
We identified a recently introduced inefficient database query was causing unnecessary strain when handling extremely large data sets. This led to delayed responses to the Data API sessions endpoints. This further led to delays or unavailability in select reports via the Reports API. Additionally, some timeouts and safeguards that should have triggered earlier did not activate as expected.
### Resolution
To quickly restore performance:
* We optimized the affected queries, improving efficiency and reducing database load.
* We scaled up the impacted systems to immediately process the backlog.
* We fine-tuned system timeouts and connection limits to prevent similar issues in the future.
Following these fixes, all services resumed normal operations.
### Prevention
To prevent future occurrences:
* We have permanently updated our database handling methods to ensure more efficient query execution.
* We are enhancing automated monitoring to detect and respond to database slowdowns before they impact customers.
We appreciate your patience and remain committed to delivering a seamless experience.
Analytics APIs Experiencing Degraded Performance in US-East-1
开始时间 2025年1月9日 UTC 13:33 · 1h 17m
Issues轻微事件
受影响的组件
Availability of session informationLoading and rendering of reports
investigating
As of 13:30 UTC, We are currently investigating a possible performance degradation affecting Reports API and Data API results in US-East-1.
Learnosity Support and Systems Engineering teams are actively investigating the issue, and will follow on with an update and resolution as soon as possible.
investigating
As of 14:00 UTC, Learnosity is working on identifying the cause of degraded performance in our Analytics APIs, Reports and Data.
Authoring and Assessment stacks remain unaffected
Learnosity Support and Systems Engineering teams are continuing to actively investigate the issue, and will follow on with an update and resolution as soon as possible.
monitoring
As of 14:10 UTC, Learnosity has restored full service for all new requests to all Analytics APIs.
We are monitoring stability and analyzing all queued requests.
Learnosity Support and Systems Engineering teams are continuing to actively investigate the issue, and will follow on with an update and resolution as soon as possible.
resolved
As of 14:50 UTC, after a further 60 minutes of issue-free monitoring, we are resolving this issue affecting availability of session information via our Analytics stack. Reports API and Data API are fully operational.
Learnosity Support and Systems Engineering teams will follow up with a post mortem once we have completed root cause analysis and finalized any next steps or preventative measures required.
Please reach out if you have any questions or concerns.
postmortem
**Affected Systems and Regions**
On 2025-01-09, Learnosity experienced a brief partial outage affecting our analytics stacks, specifically the Reports API and Data API in the AMER region. The issue began at 13:30 UTC and was resolved at 14:06 UTC, lasting 36 minutes. All other API's were unaffected and there was no loss of data.
**Investigation**
We discovered that a large number of atypical, inefficient Data API queries requested by customers were taking too long to complete. This prevented other queries from running in a timely manner, creating a backlog. It was determined that an additional database index would significantly improve the response times of these types of queries.
**Resolution**
Immediately upon discovering this issue, the impacted database instances were successfully scaled up to ease the backlog of requests. The additional index was implemented and all remaining queries were processed quickly. Affected APIs returned to normal operations, and further monitoring ensured the issue was fully resolved.
**Prevention**
Following further testing, the new index is working well and has become a permanent part of the system. We are also adding new automated monitoring and regression testing to ensure similar requests perform as expected.
Investigating issue affecting live progress report / proctoring in VA
开始时间 2024年10月1日 UTC 15:09 · 59m
Issues轻微事件
受影响的组件
AMER Analytics (Summary)Live Progress (Live Activity by User) report
investigating
As of 15:00 UTC, we are currently investigating an issue affecting the event delivery that powers the live progress report (live-activitystatus-by-user) in VA (us-east-1).
Learnosity Support and Systems Engineering teams are actively investigating the issue, and will follow on with an update and resolution as soon as possible.
identified
As of 15:20 UTC, the Systems Engineering team has identified the cause for the slower than normal delivery of events and is proceeding with mitigation steps.
Learnosity Support and Systems Engineering teams are continuing to actively investigate the issue, and will follow on with an update and resolution as soon as possible.
monitoring
As of 15:35 UTC, the Systems Engineering team has deployed a fix, the errors have stopped, and currently all systems are functional.
We will continue to monitor for a time before resolving the incident but we are currently seeing no errors.
We will follow on with an update and resolution as soon as possible.
resolved
As of 16:05 UTC, We are have resolved the issue affecting live progress reports in VA (us-east-1).
As the recovery remains stable, we are now resolving the incident that was affecting live progress reports.
Learnosity Support and Systems Engineering teams will follow up with a post mortem once we have completed root cause analysis and finalised any next steps or preventative measures required.
Please reach out if you have any questions or concerns.
Investigating potential issue affecting live proctoring in us-east-1 (VA)
开始时间 2024年9月13日 UTC 14:35 · 1h 28m
Issues轻微事件
受影响的组件
AMER Assessment (Summary)Sending and receiving Live Activity Status by User report events
investigating
As of 14:30 UTC, We are investigating a potential issue affecting the delivery of some Events API events (used in the live proctoring / live-activitystatus-by-user report) in us-east-1 (VA).
Assessments are otherwise proceeding normally and can be saved and submitted without any issue.
This issue if confirmed would only affect the delivery of Events API events.
Learnosity Support and Systems Engineering teams are actively investigating the issue, and will follow on with an update and resolution as soon as possible.
identified
As of 15:00 UTC, the Learnosity Systems Engineering team has identified the issue preventing the assessment events from being delivered and are working on the implementation of remediation steps.
Learnosity Support and Systems Engineering teams are continuing to work on the issue, and will follow on with an update and resolution as soon as possible.
monitoring
As of 15:30 UTC, events are flowing normally once again and all systems are operational.
We will continue to monitor for some time before we resolve the issue fully to ensure we have a stable recovery.
resolved
As of 16:00 UTC, after monitoring and confirming the stable state of the fixes put into place, we have resolved the Events API events delivery issue affecting proctored assessments in us-east-1 (VA).
Learnosity Support and Systems Engineering teams will follow up with a post mortem once we have completed root cause analysis and finalised any next steps or preventative measures required.
Please reach out if you have any questions or concerns.
Possible issue affecting small number of adaptive sessions not triggering Firehose updates in US-East-1
开始时间 2024年5月22日 UTC 18:27 · 1h 15m
Pending
受影响的组件
Adaptive session information
investigating
As of 18:10 UTC, we are trying to determine if an issue exists causing a very small number of adaptive sessions to not trigger a Firehose event after scoring in US-East-1.
This is just an internal investigation at present, without confirmation, and we have no evidence that any impact outside the described implementation is affected.
Learnosity Support and Systems Engineering teams are actively investigating the issue, and will follow on with an update and resolution as soon as possible.
investigating
As of 18:30 UTC, Learnosity Support and Systems Engineering teams are continuing to investigate this unconfirmed issue and will follow on with an update as soon as possible.
investigating
As of 18:45 UTC, Learnosity Support and Systems Engineering teams are continuing to investigate an unconfirmed issue, examining less than 500 potentially affected sessions out of 4.93 million submitted over the same period. We will follow on with an update as soon as possible.
monitoring
As of 19:00 UTC, we have been unable to find any systemic problems, while the number of potentially impacted sessions remains at one hundredth of one percent ( 0.01%) of total submissions over the period investigated, likely due to a content issue.
Learnosity Support and Systems Engineering teams will continue to monitor for 30 minutes and expect to close this investigation at that time.
resolved
As of 19:40 UTC, we believe we've identified content to be the root cause of this tiny number of sessions not passing through scoring and triggering a Firehose event. We've reached out to customer proactively and will help address this issue directly.
We are now calling this issue resolved and apologize for any distractions. Please reach out if you have any questions or concerns.
Issue affecting authoring and specific kinds of assessments in U.S.-East-1
开始时间 2024年4月30日 UTC 13:29 · 1h 55m
Issues轻微事件
受影响的组件
Availability of session informationCreating and saving of Items/Activities/TagsScoring endpointLoading and rendering of Items/Questions/FeaturesFirehoseLoading and rendering of reports
investigating
We are currently experiencing an issue affecting the availability of newly authored and edited content. This is also having a ripple effect on assessments dynamically retrieving content, such as adaptive and on-the-fly assessments.
We are investigating and will update this record as soon as we learn more.
investigating
As of 13:45 UTC, we are investigating an issue with the availability of authoring and assessment content in the US-East-1 region. This issue is also impacting analytics, with degraded performance of Reports and select Data API endpoints that require content, such as the Session Detail by Item report and Data API item bank and scoring endpoints.
We will continue to update this incident as we learn more.
Learnosity Support and Systems Engineering teams are continuing to actively investigate the issue, and will follow on with an update and resolution as soon as possible.
monitoring
As of 14:30 UTC, we have addressed an issue primarily affecting Data API use. Error rate has reduced to zero, all newly submitted sessions are scoring immediately, and the backlog of sessions awaiting scoring is cleared. The Firehose service queue is clearing quickly dispatching information about recently scored sessions.
Learnosity Support and Systems Engineering teams are continuing to actively investigate the issue, and will follow on with an update and resolution as soon as possible.
resolved
As of 15:00 UTC, we are closing this incident as resolved. The availability and performance of the Data API and downstream systems in US-East-1 has been restored and operating correctly for 30 minutes.
Learnosity Support and Systems Engineering teams will follow up with a post mortem once we have completed root cause analysis and finalized any next steps or preventative measures required.
Please reach out if you have any questions or concerns.
postmortem
On 2024-04-30, Learnosity suffered a service interruption in the AMER \(US East-1\) region that affected a subset of customers. The incident began at 12:52 UTC and was resolved at 14:18 UTC, lasting 1 hour and 26 minutes. We apologize for the impact that this had on our valued customers and learners. We have learned, and will continue to learn, from this incident and have taken steps to ensure this issue doesn’t repeat.
**Details of Incident**
Performance issues were initially detected by our Site Reliability Engineering team, who executed our escalation process in raising Tier1 and subsequently Tier2 in short order.
While assessment and authoring load were typical, we saw a massive spike in Data API activity from a single client. This was limited to the **get\_/itembank** endpoints set typically used for authoring. The single client had never used these Data API endpoints before, went from zero to equal to all other clients combined in a short period of time.
This is a particularly laborious query necessary for authoring use. Calling this at such a high volume and frequency created a hotspot and contention on the cluster of four itembank delivery databases, which handle a subset of customers. When we’re under high volume of typical traffic we operate on about 30% peak database load. The atypical traffic outlined above put the databases at full capacity.
While all of our assessment critical systems auto scale, the itembank delivery databases are over provisioned to have significant headroom, and as such don’t auto scale. Auto scaling of database instances can introduce a risk if anything goes wrong while scaling, and as such when designing the system our team determined manual scaling of these itembank delivery databases was the more reliable option.
This atypical activity led to resource starvation of the Data API, with knock-on effects to the assessments stack. Data API impacts included limiting access to authoring and session data, as well as slowing down session scoring and follow-on dispatching of scoring events via Learnosity Firehose.
During the incident, we experienced degraded performance with up to 30% of requests not succeeding. Assessment impacts included cases where the Data API was used in assessment delivery, and in some cases limited new initializations \(i.e. starting a test\). For the avoidance of doubt, all save and submit calls were successful without data loss.
**Resolution**
Once the atypical traffic was identified, temporary request rate limits were put in place to allow our remediation efforts to quickly scale to support the increased level of activity and process backlog. Once all services were again operating at optimal levels, the temporary request limits were removed.
**Additional Analysis and Prevention**
We are making the following changes as a result of this operational event, to prevent this from happening again.
* We identified that the baseline rate limit configured for the itembank endpoints was not appropriate - and have reviewed and configured this appropriately.
* We are reviewing improvements and resource allocations of the Data API to handle atypical usage patterns and increase resilience.
* We are continuing to perform analysis to determine what other preventative measures are appropriate to implement.
Investigating possible degraded performance affecting some asynchronous jobs in the US-West-2 Region
开始时间 2024年4月25日 UTC 20:43 · 1h 17m
Issues轻微事件
受影响的组件
AMER Data Centric (Summary)Creation of Item Pools
investigating
As of 20:30 UTC, we are investigating a possible issue affecting the processing of item pool and offline package jobs in the US-West-2 region and causing slower than normal processing times.
Learnosity Support and Systems Engineering teams are actively investigating the issue, and will follow on with an update and resolution as soon as possible.
identified
As of 21:15 UTC, the Learnosity Systems Engineering team has identified the issue and is making adjustments to restore normal jobs performance.
Learnosity Support and Systems Engineering teams are continuing to work on the issue, and will follow on with an update and resolution as soon as possible.
monitoring
As of 21:26 UTC, the jobs processing performance has returned to normal and all job queues are empty.
Everything is currently performing as expected, we will continue to monitor for a time to ensure a full recovery before fully resolving the incident.
resolved
As of 21:56 UTC, all jobs continue to be processed at a normal rate and all services are operational.
We are marking the issue affecting the processing of item pool and offline package jobs in US-West-2 as resolved.
Learnosity Support and Systems Engineering teams will follow up with a post mortem once we have completed root cause analysis and finalised any next steps or preventative measures required.
Please reach out if you have any questions or concerns.
Issue affecting Tag Hierarchies in the US-East-1 region
开始时间 2024年1月26日 UTC 15:00 · 2h 14m
Issues轻微事件
受影响的组件
Creating and saving of Tag Hierarchies (author site)Tag-based reports
investigating
As of 14:30 UTC, we are investigating an issue preventing tag hierarchies from being created or updated in the US-East-1 Region.
This only appears to affect new changes to tag hierarchies, so this should only impact the use of select reports that rely on this feature, and only when a newly created or updated hierarchy is required. Routine authoring and assessments are unaffected, as are reports and Data API use that don't rely on newly edited hierarchies.
We're including tag-related reports in this update for visibility, but the Reports API remains operational and only a small subset of reports use tag heirarchies.
Learnosity Support and Systems Engineering teams are actively investigating the issue, and will follow on with an update and resolution as soon as possible.
investigating
As of 16:00 UTC, we are continuing to investigate an issue preventing updates to Tag Hierarchies.
This issue is limited to the creation and editing of tag hierarchies. Only a small subset of reports that make use of this feature may be affected, and only when a change to a tag hierarchy is required. All other APIs, including reports that don't rely on hierarchy updates, remain operational.
Learnosity Support and Systems Engineering teams are continuing to actively investigate the issue, and will follow on with an update and resolution as soon as possible.
resolved
As of 17:00 UTC, we've identified the problem affecting tag hierarchies and it is not an infrastructure/services issue. We've determined that an application-level bug limited to saving a newly edited hierarchy is the root cause and have raised a ticket for processing.
All other API and backend functionality, including the use of existing tag hierarchies, remains operational. We are closing this incident and moving the focus to the software team. If you experience any issues with saving hierarchies please don't hesitate to raise a ticket with the Learnosity Support team and we will provide additional updates. Thank you.
Issues loading assessment content in AMER region
开始时间 2023年12月13日 UTC 14:06 · 5h 36m
Outage重大事件
受影响的组件
Loading and rendering of Item/Activity Edit viewLoading and rendering of Item/Activity Edit viewLoading and rendering of Item/Activity Edit viewLoading and rendering of Items/Questions/FeaturesLoading and rendering of reports
investigating
As of 14:00 UTC, We are seeing intermittent issues with loading demo assessment content and are investigating out of an abundance of caution.
Learnosity Support and Systems Engineering teams are actively investigating the issue, and will follow on with an update and resolution as soon as possible.
identified
As of 14:15 UTC, we've identified that the AWS Auto Scaling Groups failed to create additional EC2 instances causing an elevated number of 502 errors in the assessment stack.
Learnosity Support and Systems Engineering teams are continuing to actively investigate the issue, and will follow on with an update and resolution as soon as possible.
investigating
As of 14:15 UTC, the assessment stack returned to normal operating metrics in the EMEA and APAC regions. As of 14:30 and 14:45 UTC, the NA region error count dropped by approximately 15% each period, and error counts are continuing to reduce as of 15:00 UTC. Intermittent errors continue to occur in the assessment stack in AMER. Authoring and Analytics APIs remain unaffected.
Learnosity Support and Systems Engineering teams are continuing to actively investigate the issue, and will follow on with an update and resolution as soon as possible.
investigating
Correction: The authoring and analytics stacks are also begin impacted intermittently in workflows where the assessment stack is in play. (E.g. editing views where the assessment stack is used for previews.) Errors in the AMER region continue to reduce. 502 errors have cleared and 504 timeout errors are now being investigated. Elevating classifications to include all stacks and partial outage.
Learnosity Support and Systems Engineering teams are continuing to actively investigate the issue, and will follow on with an update and resolution as soon as possible.
identified
We've now identified a possible cause for the current incident and are working to resolve the issue. We will add additional information as soon as it becomes available.
We've also added EMEA authoring and APAC authoring to affected regions for users of the hosted author site, which operates in the AMER region.
Learnosity Support and Systems Engineering teams are continuing to actively investigate the current issue, and will follow on with an update and resolution as soon as possible.
identified
As of 18:20 UTC, we've verified a cause of the current incident and recovery is already partially complete. We have restored part of the assessment stack and are working on full restoration now, with the remaining stacks to follow. We will provide additional updates soon, including a possible ETA as soon as a reasonably accurate estimate has been determined.
Learnosity Support and Systems Engineering teams are continuing to actively investigate the issue, and will follow on with an update and resolution as soon as possible.
monitoring
As of 18:36 UTC, we've now recovered fully and all stacks are operational. We will continue to monitor for a period before resolving the incident.
Learnosity Support and Systems Engineering teams are continuing to actively investigate the issue, and will follow on with an update and resolution as soon as possible.
resolved
As of 19:40 UTC, we've concluded one hour of active monitoring without additional issues. We will continue to keep an eye on everything but are now ready to call this incident resolved.
Learnosity Support and Systems Engineering teams will follow up with a post mortem once we have completed root cause analysis and finalized any next steps or preventative measures required.
postmortem
On 2023-12-13, Learnosity suffered a major outage that began at 13:22 UTC and was resolved at 18:29 UTC, lasting 5 hours and 7 minutes. Its primary impact was on the assessment stack in the AMER region, with partial impact to authoring and select analytics APIs due to the use of assessment functionality in some preview features.
Initially, we thought this incident was a scaling issue because errors were linked to the creation of new virtual servers to meet daily traffic. However, we ultimately discovered that the problem was unrelated to scaling. The root cause was actually a faulty configuration management client, `salt v3006.5`. A [bug in the most recent version of Salt-stack](https://github.com/saltstack/salt/issues/65691), released at 2023-12-12 at 21:38, caused a failure in our `cloudinit` process when spawning new EC2 instances. From 13:05 PM UTC, as new instances were launched, the `salt` bug caused the new machines to fail a health check and fall into a loop of retried launches.
**Resolution**
Immediately after this discovery, we were able to hard code to a prior version of `salt`, removing the regression. New EC2 instances were successfully provisioned and scaled to comfortably handle all traffic. All APIs returned to normal operations soon after.
**Additional Analysis and Prevention**
Upon further investigation, it was discovered that our system became vulnerable to the `salt` regression due to error handling dependencies. To facilitate fast and reliable scale up, our image build process is designed so that dependencies are pre-installed at build time, with only minor applicable config changes required at instantiation. Further investigation discovered that a prior fix for a build dependency issue inadvertently moved the installation of the affected `salt` package from _image build_ time to _instance launch_ time. This meant that the creation of new instances no longer relied on the original image version of `salt`, instead the newest `salt` version was installed during the scaling process.
We are making the following changes as a result of this operational event, to prevent this from happening again.
* We are conducting a full review of the launch process to ensure there are no other unknown launch-time dependencies.
* We are adding additional guard rails in the build process to catch any dependency regression at the imaging phase before production use.
Investigating possible degraded performance affecting Self-Hosted Adaptive in the US-East-1 Region
开始时间 2023年12月5日 UTC 17:29 · 2h 6m
Issues轻微事件
受影响的组件
Self-Hosted Adaptive services
investigating
As of 17:25 UTC, we are investigating a possible issue affecting Self-Hosted Adaptive testing in the US-East-1 region. Due to the very small number of customers affected, we've not yet confirmed if this is something related to Learnosity APIs or the self-hosted algorithm portion of the solution. Most customers using this feature appear to be unaffected, but we're raising this alert out of an abundance of caution.
Learnosity Support and Systems Engineering teams are actively investigating the issue, and will follow on with an update and resolution as soon as possible.
resolved
As of 17:55 UTC, we identified an adaptive Redis cache that was intermittently refusing a small percentage of connections. Self-hosted adaptive makes extensive use of this cache due to the need to store more state information supplied by the self-hosted client. We replaced the cache database, routed all traffic to the new instance, and saw an immediate, dramatic reduction of refused connections. We've monitored for approximately 90 minutes and all activity indicates normal operation of all APIs. We are resolving this issue and will continue to monitor customer comms.
Learnosity Support and Systems Engineering teams will follow up with a post mortem once we have completed root cause analysis and finalized any next steps or preventative measures required.
Please reach out if you have any questions or concerns.