SMS 2FA issues with AWS aggregator
- monitoring
The 2FA issues are down stream. AWS has reached out the aggregator that works with the carriers on this issue.
- resolved
This incident has been resolved.
54 Keeper incidents ยท 2023๋ 2์ โ official updates, affected components, duration and resolution details.
The 2FA issues are down stream. AWS has reached out the aggregator that works with the carriers on this issue.
This incident has been resolved.
EU ์ง์ญ์ Keeper Router ๋ฐ Gateway๋ฅผ ํตํด ์ค๋ฆฝ ๋ ์ฐ๊ฒฐ์ ์คํจํฉ๋๋ค. ์ถ๊ฐ ์ ๋ณด๋ postmortem ๋ณด๊ณ ์์ ๊ฒ์๋ฉ๋๋ค.
# Post-Incident ๋ณด๊ณ ์ ** 7์ 31, 2026 ** ## ์์ฝ 7์ 30์ผ ์คํ 11์ 33๋ถ์์ 2026๋ , EU ์ง์ญ์์ KeeperPAM ์ฐ๊ฒฐ ์ฌ์ ์ ์ค๋ฅ๋ฅผ ๋ฐ์ํ์ต๋๋ค. Router/Gateway ์๋ฌ๋ ์ฝ 12์๊ฐ 45๋ถ ๋์ ์ง์๋๋ฉฐ, 7์ 31์ผ ์คํ 12์ 18๋ถ CT์ ๋ณต๊ตฌ๋ ์ ์ฒด ์๋น์ค์ ๋๋ค. ์ด ์ฐฝ ์ค, ์ฃผ์ Keeper EU ํ๋ซํผ \ (vault Access, ์ธ์ฆ ๋ฐ ๊ธฐํ ๋ชจ๋ Keeper ์๋น์ค \)๋ ์ ๋ฐ์ ์ผ๋ก ์์ ํ ์ด์๋ฉ๋๋ค. ## ์ด๋ค Happened 7 ์ 22 ์ผ, ์ฐ๋ฆฌ์ ์ฐ๊ฒฐ ๋ผ์ฐํ ์๋น์ค์ ๋ํ ์ผ์์ ์ธ ๋ฐฐํฌ \(โKeeper Routerโ\)๋ ์๋น์ค๊ฐ ์์ ๊ตฌ์ฑ์ ๊ฒ์ํ๋ ๋ฐฉ๋ฒ์ ๋ณ๊ฒฝํ๋ ์ ๋ฐ์ดํธ ๋ ์์กด์ฑ์ ํฌํจ. ๋์ ์์ misconfiguration๋ ์ ๋ฝ ์ฐํฉ (EU) ์ง๊ตฌ๊ฐ FIPS ์๋ํฌ์ธํธ๋ฅผ ์ฐพ์ต๋๋ค. ๊ทธ ๊ฒฐ๊ณผ, EU ์ง์ญ์ ์๋น์ค ์ปจํ ์ด๋๊ฐ ๋ฐฐํฌ์ ๋ฐ๋ผ ์ฌ์์ ํ ๋, ๊ทธ๋ค์ ๊ทธ๋ค์ ์์ ๊ตฌ์ฑ์ ๊ฒ์ํ๊ณ ์คํจ ์ํ ์ ๋ ฅ ํ ์ ์์๋ค. ๋ ๊ฐ์ง ์์๋ ์ด๋ฒ ์คํจ๊ฐ ์ผ์ฃผ์ผ ๋์ ๋ฐ๊ฒฌ๋์ง ๋ชปํ์ต๋๋ค. ์ฒซ ๋ฒ์งธ, ๊ธฐ์กด ์ปจํ ์ด๋๋ ๊ณ์ ํธ๋ํฝ์ ์ ๊ณตํ๋ฉด์ ๊ต์ฒด ์ฉ๊ธฐ๊ฐ ์นจ๋ฌต์ ์ผ๋ก ์์๋์ง ์๋๋ก 7 ์ 22 ๋ฐฐํฌ์์ ์ฆ๊ฐ์ ์ธ ๊ณ ๊ฐ ์ํฅ์ด ์์์ต๋๋ค. QA๋ ๋ํ ๋ชจ๋ ์์ฐ ๊ฒ์ฆ ์ํ์ ํต๊ณผํ์ต๋๋ค. ๋ ๋ฒ์งธ, ๊ตฌ์ฑ๋ก๋ ์คํจ๋ ERROR ๋๋ CRITICAL๋ณด๋ค INFO ์์ค์์ ๊ธฐ๋ก๋์์ผ๋ฏ๋ก PagerDuty ๊ฒฝ๊ณ ๊ฐ ์์ฑ๋์ง ์์์ต๋๋ค. 7 ์ 30 ์ผ ๋ฐค์, ๋ง์ง๋ง์ผ๋ก ๊ฑด๊ฐํ ์ฉ๊ธฐ๊ฐ ๋ฐ์ผ๋ก ์ฌ์ดํด๋ง, ์๋น์ค์ ๋ถํ ์์ก์ ์ ๋ก ๊ฑด๊ฐ ํ ๋ชฉํ๋ฅผ ๊ฐ์ง๊ณ , ์๋ ํฌ์ธํธ๋ HTTP 503 ์ค๋ฅ๋ฅผ ๋ฐํ ์์. ์๋ํ๋ ๊ฑด๊ฐ ๊ฒ์ฌ ๊ฐ์๋ ์ด ์์ outage๋ฅผ ๊ฒ์ถํ๊ณ ์ฐ๋ฆฌ์ on-call ํ ํ์ด์ง๋ฅผ ๋ง๋ค์์ต๋๋ค. ## ์ฐ๋ฆฌ๊ฐ ๋ณํํ๋ ๊ฒ 1. ** ์ปจํ ์ด๋๊ฐ ์คํจํ ๋กค์ค๋ฒ์ ๋์ฐฉ. ** CloudWatch๋ ์ปจํ ์ด๋ ์์ ์ด ๋ฐ๋ณต์ ์ผ๋ก ์์๋ ๋ ๋ชจ๋ ์์ญ๊ณผ ํ๊ฒฝ์ ๊ฑธ์ณ CloudWatch ์๋์ ๋ฐฐํฌํ๊ณ , ์กฐ์ฉํ ๋ฐฐํฌ ์คํจ๋ฅผ ๊ฐ์งํฉ๋๋ค. 2. ** ์์ ์คํจ๋ฅผ ์ํ ๋ก๊ณ ์ฌ๊ฐ์ฑ. ** ํ์ํ ์์ ๊ตฌ์ฑ์๋ก๋ํ๋ ๋ชจ๋ ์คํจ๋ ์ด์ ์ ์ ํ ํ๋ ๊ฒฝ๊ณ ๋ฅผ ์ ๋ฐํ๋ INFO ๋์ ERROR/CRITICAL severity์ ๋ก๊ทธ์ธํฉ๋๋ค. ์ฐ๋ฆฌ๋ ํผ๋์ ๋ํ ์ฐ๋ฆฌ์ EU ๊ณ ๊ฐ์๊ฒ ์ฌ๊ณผํฉ๋๋ค.
๊ณต์ ์ธ์๋ํธ ์ ๋ฐ์ดํธ๋ฅผ ์๋ ๋ฒ์ญํ์ต๋๋ค.
Keeper Gateway์ KSM ์ฐ๊ฒฐ ์ค๋ฅ๊ฐ ๋ฐ์ํ์ต๋๋ค. ์ฐ๋ฆฌ๋ ๋ฌธ์ ์ ์ ํ์ธํ๊ณ ๋ช ๋ถ ์์ ํด๊ฒฐ๋ ๊ฒ์ ๋๋ค.
API ์ค๋ฅ๊ฐ ํด๊ฒฐ๋์์ต๋๋ค. Keeper Gateway ๋ฐ KSM ํด๋ผ์ด์ธํธ ๋ฒ์ ์ ๋ ์ด์ ์ค๋ฅ๋ฅผ ๋ฐํํ์ง ์์ต๋๋ค.
๊ณต์ ์ธ์๋ํธ ์ ๋ฐ์ดํธ๋ฅผ ์๋ ๋ฒ์ญํ์ต๋๋ค.
We are aware of an issue with direct record sharing for enterprise users as a result of yesterday's Vault 18.3 update. We will issue a patch before 12PM PST. To work around the issue, you can click the "Can share to users outside the enterprise" option in Role enforcements, or just use the Desktop App.
The issue was resolved at 12:30PM PST.
On March 30 at 8:00 PM PST, Keeper performed a scheduled minor upgrade to our multi-region AWS RDS clusters as part of routine maintenance. The upgrade itself completed in approximately 3 minutes, followed by a standard application restart process that took about 15 minutes. While this process is typically non-disruptive, recent changes to service health checks caused unexpected customer impact during the restart window. This maintenance was not communicated in advance due to an internal miscommunication and the expectation of no impact. We apologize for the disruption and are taking steps to ensure all future maintenanceโregardless of expected impactโis communicated ahead of time, as well as reviewing our health check configurations to prevent similar issues.
We are investigating errors in establishing KeeperPAM connections in the US Data Center.
We have identified the cause of the KeeperPAM connection errors which are related to a networking issue within the AWS ECS environment. We are actively troubleshooting with the AWS support team and will update the ticket as soon as there is a status update.
A fix has been implemented and we are monitoring the KeeperPAM connection stability. We will update this case with additional details soon.
The issue has been resolved. See postmortem page with additional details.
At 9:05 AM PST, alerts were triggered for issues affecting KeeperPAM connections managed through Keeperโs ECS deployments in the US-EAST region. There were no recent changes to the application or environment. Investigation identified a low-level concurrency bug in the Keeper EPM service that caused request failures under high simultaneous load. These failures led to instability in the ECS services supporting KeeperPAM connections. As a temporary mitigation, we blocked the error condition, restoring KeeperPAM connectivity by approximately 1:00 PM PST. The engineering team then developed and deployed an updated Keeper Router version to address the underlying issue and prevent EPM agents from triggering server errors. The fix was fully validated by 3:00 PM PST, at which point all KeeperPAM services were stable and operating normally.
We have identified an issue causing KeeperPAM connection errors in the US Data Center. It will be resolved shortly.
The issue has been resolved. We will update the ticket with additional details shortly.
At 11:08 AM PST, DevOps and customers reported sporadic error messages occurring after login for KeeperPAM customers with connections and tunnels enabled. While vault login itself was not affected, the errors were triggered asynchronously after login due to a communication issue with the KeeperPAM router endpoint. Further investigation determined that a "scheduler" database within the AWS RDS environment - used for certain PAM scheduling operations - was encountering an unexpected error. The operations team resolved the underlying database issue and restarted the affected services. This scheduler service was inadvertently impacting connection establishment. To prevent similar issues in the future, we will be implementing software changes to ensure that scheduling-related errors cannot affect connection functionality. Additional monitoring and alerting have also been implemented for the scheduler databases to help detect and prevent similar issues going forward. Full service for connection and tunneling capabilities was restored by 11:40 AM PST.
As part of our planned GovCloud region maintenance involving database upgrades, restarts of the Redshift cluster were required. As a result, downstream services required configuration updates and all services are now restored.
Maintenance in the GovCloud region was scheduled to begin at 8:00 PM PST with a planned duration of 30 minutes. During the maintenance window, certain services required additional reconfiguration beyond the original scope. Core services were restored within the planned 30-minute window. However, due to the extended reconfiguration activities, some users experienced API errors during login for approximately 15 minutes following service restoration.
During scheduled system updates, we have encountered unexpected NGINX 503 / 502 errors in the EU region. As of 4:50 PM CST, work is still in progress. During this time, users in the EU region will experience brief or intermittent access interruptions when using Keeper.
The issue has been resolved. We will update the postmortem report shortly.
On December 2, our **EU region** experienced an unexpected outage during planned infrastructure updates. The issue began at **1:24 PM PT** and service was fully restored by **4:10 PM PT**. **What Happened** During an update to our server configuration, a region-specific configuration error caused the EU environment to begin throwing API errors. Although the underlying issue was identified quickly, each corrective change required a full autoscaling group rollout \(up to 30 minutes per cycle\) which extended the total recovery time. **Resolution** Our engineering team implemented the necessary fixes and restored service. We have also updated our infrastructure-as-code to ensure this issue cannot recur in future deployments. We apologize for the disruption and appreciate your patience while we resolved the incident.
There are intermittent connectivity issues due to the AWS outage in US-EAST-1 affecting KeeperPAM cloud connections, scheduled rotations and discovery operations.
Clearing this alert as most AWS services are restored.
AWS is reporting errors in the US-EAST region that are affecting multiple services, affecting email and SMS delivery which impacts Keeper device approvals. The AWS Service Health dashboard can be monitored here: https://health.aws.amazon.com/health/status
SMS delivery services have been restored, but AWS Lambda is still down which affects our email delivery of device approvals. We recommend using the "Keeper Push" or TOTP method of device approval.
All Keeper email delivery services in US-EAST-1 were restored around 4AM PST.
We are aware of Auth API errors in EU region and we have identified the isssue. The issue will be resolved shortly.
The issue with Auth API errors in the EU region has been resolved. Additional details will be posted in the postmortem shortly.
On **September 10 at 8:42 AM PDT**, our monitoring detected elevated errors with the **Auth API in the EU region**. The issue was identified promptly and fully resolved by **8:49 AM PDT**. Total impact duration was approximately **7 minutes**. ### Impact * **Region affected:** EU * **Service affected:** Authentication API * **Customer impact:** Customers in the EU experienced failed authentication attempts during the incident window. ### Timeline \(PDT\) * **08:42** โ Incident detected. Elevated Auth API errors in the EU region reported. * **08:43** โ Engineering and DevOps teams engaged. * **08:46** โ Root cause isolated to a Terraform-related configuration change involving a VPC endpoint. * **08:49** โ Configuration change rolled back; Auth API errors resolved. ### Root Cause The outage was caused by a configuration change in our production tenant during an ongoing Terraform migration. As part of the migration, a new VPC endpoint was being added to production environments. During Terraform apply, a subnet association was also being created. A network disruption occurred between the creation of the VPC endpoint and the subnet association \(behavior not seen in prior testing\) leading to temporary Auth API connectivity failures in the EU region. A new Terraform migration plan has been created that eliminates the possibility of disruption.
Resolved - At 9:49 AM PST, we experienced login API request failures due to errors on an RDS database in the US region. The database failover completed successfully, and all systems were fully operational by 10:00 AM PST.
Several customers have been experiencing SSO login errors this morning. The team determined the cause of the issue which is detailed in the postmortem report.
On June 30, 2025, Keeper published an updated SAML Service Provider \(SP\) certificate as part of our standard annual process. This certificate is used for signing SP requests and does not impact SSL termination. During the update, the certificate subject name was unintentionally changed from sso.keepersecurity.com to sso.staging.keepersecurity.com. While most identity providers do not validate the certificate subject name, certain providersโsuch as JumpCloud and Shibbolethโmay enforce strict matching and reject SAML requests when the subject name does not align. On Sunday July 6, 2025, routine backend infrastructure changes caused the cert to be propagated across production systems throughout the morning hours. The issue with the subject name was identified by our team after several troubleshooting calls with customers. The DevOps team then corrected the issue by publishing the certificate with the proper subject name around 12pm PST. After reviewing the case history, this affected a small number of customers. Next year, when the SP cert change is approaching, we will notify all customers with the exact date and time when the cert will be updated, so those customers using strict SP cert matching can be prepared. We apologize for the error and we have taken steps to ensure this does not occur again.
We are investigating vault login errors
We are continuing to investigate this issue.
The issue has been resolved. The outage was caused by a high volume of database locks related to a large BreachWatch data update. The engineering team is tracking down the root cause and will be addressing a preventative solution today.
At 8:15 AM PST, our DevOps team was alerted to elevated API error rates in production. Upon investigation, we identified the following: * The production database was experiencing high table lock contention. * This was triggered by several large-scale BreachWatch data updates. * A recent backend migration had unintentionally doubled the number of parallel update operations, which exacerbated the issue. * A database failover was performed to restore service stability. The issue was fully resolved by 8:40 AM PST. The engineering team has addressed the underlying problem in the update processes and is implementing safeguards to prevent recurrence, including query optimization and improved job orchestration. We apologize for the disruption.
We are investigating login errors in the EU data center.
We have identified the issue causing login errors in the EU data center. Our team is working on a resolution.
Login errors in the EU region have been fully resolved. We will update this incident postmortem with additional information shortly.
At 4:37AM PST, Keeper experienced API errors in the EU region. Service was restored by 5:35AM PST following a database failover. The issue was caused by a bug in the new Device Management API, released the day before. During device registration, the system attempted to delete old devices, which created high database load for accounts with large device counts. The bug was identified, fixed, and a patch was deployed. We apologize for the disruption. To ensure uninterrupted access, we recommend enabling โWork Offlineโ mode for your account.
Some users are receiving errors when downloading file attachments. We are working on resolution.
The issue has been resolved. Additional information is in the postmortem report.
### **Incident Summary** In the US, EU, and AU regions, users experienced vault errors when attempting to download a file attachment immediately after adding it to a record. This issue affected only newly uploaded attachments in these three data center regions. The Gov, JP, and CA regions were not impacted. ### **Root Cause** During our investigation, we identified a discrepancy in how file attachment triggers were implemented across different regions. A recent Terraform change inadvertently removed the triggers responsible for moving file attachments in the US, EU, and AU regions, preventing newly uploaded files from being immediately available for download. ### **Resolution** We corrected the issue by restoring the missing triggers and applying the necessary changes. Additionally, any affected files were successfully moved to their intended locations. ### **Impact & Mitigation** * No data loss occurred. * All impacted files were successfully resolved by our corrections. * We corrected the issue to prevent similar discrepancies in future infrastructure changes. We apologize for any inconvenience and appreciate your patience while we resolved this issue.
Between 5:15PM PST and 5:30PM PST, there were a high number of Vault Login errors.
### **Incident Summary** The Keeper DevOps team was alerted to API errors in the US region. Upon investigation, we identified a high number of database deadlocks affecting performance. To mitigate the immediate impact, we restarted the affected writer instance. ### **Root Cause** Historically, similar issues were linked to specific queries impacting MySQL 8.x updates on the RDS Aurora platform. In this case, we discovered that a utility designed to monitor deadlocks and capture diagnostic data was inadvertently contributing to the problem. A known bug in the MySQL minor version used in production caused our data-gathering queryโintended to analyze deadlocksโto trigger additional locking, exacerbating the issue. ### **Resolution & Mitigation** * We promptly disabled the problematic monitoring utility and applied database parameter optimizations in collaboration with the AWS team to enhance performance under peak load. * These changes were fully implemented within 24 hours, and no further excessive row locks have been observed. * Our DBA team is actively working with backend engineers to further optimize write operations for improved performance. * A backend code update with additional optimizations is scheduled for release the week of March 17. We appreciate your patience and will continue to monitor and enhance system performance to prevent similar issues in the future.
We are investigating login errors and our team is currently working on the issue.
We are continuing to investigate this issue.
The issue has been resolved
We are still working on the issue which is causing login errors in the US region. Please use Keeper's "work offline" while we are working to resolve the issue.
The issue has been resolved and root cause determined. We will update the postmortem on this issue shortly.
We identified a backend API and resulting database query which triggered a high database load and affected production services. We are publishing a permanent fix on Monday, Feb 24.