JumpCloud is currently experiencing delays in assigning Users and Systems to Dynamic Groups. Administrators may see delays in these Group membership updates, and thus delays in provisioning to connected resources.
identified
The issue has been identified and a fix is being implemented.
identified
We are continuing to work on a fix for this issue.
monitoring
A fix has been implemented and we are monitoring the results.
monitoring
We are continuing to monitor for any further issues.
We are currently investigating delays or timeouts loading the user console. We are investigating the cause of the issues currently, and will provide an update within one hour.
identified
We have identified that a third party outage is impacting our service. More information can be found at https://www.cloudflarestatus.com/incidents/v3yl7jqmqj51
resolved
The third party issue impacting our service has been resolved. Our internal monitoring is showing service states returned to normal.
postmortem

**Date**: Jun 25, 2026
**Date of Incident:** Jun 22, 2026
**Description**: RCA for Third-Party CDN Network Disruption — Cloudflare / Zayo Fiber Cut
#### **Summary:**
On June 22, 2026, beginning at approximately 13:25 UTC, some JumpCloud customers experienced degraded access to web-facing services including the Admin Console, User Portal, API endpoints, and SSO authentication flows. Customers connecting through North America or accessing services routed through North American infrastructure were most affected.
The root cause was a fiber cut in Eastern North America, a physical break in underground or undersea cables carrying internet traffic. Cloudflare attributed the disruption to Zayo, a network transit provider, experiencing an outage on some of its network routes, which caused reachability issues for services routing through those paths. Because JumpCloud's traffic transits the Cloudflare network before reaching JumpCloud's origin infrastructure, degradation at the Cloudflare/Zayo layer directly impacted our customers' ability to reach JumpCloud services.
By late morning, most affected platforms had stabilized, with Cloudflare indicating its fix was being actively rolled out. JumpCloud's backend services and data plane remained fully operational throughout the event. The issue was isolated to the traffic path between end users and JumpCloud's origin infrastructure.
#### **What Happened:**
JumpCloud utilizes Cloudflare as the security and networking layer in front of all its global infrastructure. All HTTP/HTTPS requests to JumpCloud's web services traverse the Cloudflare network before reaching JumpCloud's origin infrastructure.
#### **Provider Root Cause:**
Cloudflare engineers first detected elevated error rates and latency at approximately 13:25 UTC on June 22, 2026. By 14:37 UTC, they had traced the root cause to a fiber cut in Eastern North America. Cloudflare confirmed that Zayo, a network provider, was experiencing an outage on some of its network routes, causing sites and services routing through those paths to become unreachable.
Cloudflare's status page indicated that customers connecting through North America or accessing services in Europe may have seen increased latencies and timeouts as Cloudflare engineers worked to mitigate the issue. Traffic engineering efforts successfully mitigated the majority of congestion and packet drops, with services reported as largely stable with only minor residual impact remaining.
Cloudflare's scheduled Newark \(EWR\) datacenter maintenance overlapped in timing but was a separate event - it was not the root cause of the fiber cut outage.
#### **Corrective Actions \(Target End-of-July 2026\)**:
While this incident was caused by a third-party provider we recognize the impact it had on our customers. JumpCloud is committed to reducing our exposure to single-provider failures and improving our resilience posture. The following corrective actions are actively underway:
* **Multi-CDN Redundancy:** Cloudflare remains part of our architecture. We are adding a parallel, independently operated CDN path with weighted routing and security parity across both, so traffic can shift automatically if one path degrades. When complete, JumpCloud's edge will be capable of routing around a complete outage of any single CDN provider without customer action or service interruption. This work is in active delivery within our Platform Engineering organization.
* **Resilient DNS Management:** We are modernizing how we manage DNS with phased, automated changes and rollback capabilities, reducing the risk that DNS operations themselves become a source of disruption.
* **Validated Failover:** We are running controlled failover exercises including a production-scale test before relying on this redundancy in production.
* **Enhanced Third-Party Monitoring & Automated Alerting:** We are expanding our monitoring to include automated detection of third-party provider degradation, enabling faster incident declaration and more rapid customer communication. This includes synthetic monitoring from multiple geographic regions that can differentiate between JumpCloud-origin issues and upstream provider issues
#### **Provider Remediation:**
Cloudflare has communicated the following remediation actions:
* **Transit Path Restoration:** Cloudflare is coordinating with Zayo and fiber providers to repair the affected Eastern North American routes and restore full transit capacity.
* **Traffic Engineering Response:** Cloudflare's traffic engineering teams implemented manual rerouting to redistribute load across unaffected network paths during the incident window.
* **Ongoing Investigation:** Cloudflare has indicated a full post-incident review is underway to assess automated failover capabilities for large-scale transit provider failures.
JumpCloud takes the reliability and availability of our platform seriously. We understand that our customers depend on JumpCloud for critical identity and device management operations, and any disruption, regardless of its origin, impacts their business.
We are actively investing in infrastructure resilience to ensure that our dependency on any single third-party provider does not create an unacceptable risk to service availability.
Degraded Console Service - Devices Page slowness
Started May 8, 2026 at 6:41 AM UTC · 1h 47m
IssuesMinor incident
Affected components
Admin Console - US Region
investigating
We are currently investigating delays or timeouts loading the devices page in the admin console. We are investigating the cause of the issues currently, and will provide an update within one hour.
identified
The issue has been identified and the team is working on the fix.
monitoring
A fix has been implemented. We are currently monitoring the results.
resolved
This incident has been resolved.
Degraded Agent Service on MacOS, Windows and Linux
Started April 28, 2026 at 8:01 AM UTC · 1h 54m
IssuesMinor incident
Affected components
Agent - US Region
investigating
We are currently aware of reports with agent Installation failing, this is affecting MacOS, Windows and Linux. We are investigating the cause of the issues currently, and will provide an update within one hour.
identified
The issue has been identified and a fix will be implemented soon.
monitoring
A fix has been implemented, and Agent installations should now be functioning as expected. Our team is actively monitoring the situation to ensure continued stability.
resolved
The incident has been resolved.
LDAP Directory Processing Delay
Started April 2, 2026 at 5:08 PM UTC · 47m
IssuesMinor incident
Affected components
LDAP - US Region
investigating
We are currently investigating a delay in LDAP directory synchronizations. While updates are processing, changes to user attributes and passwords may take longer than expected to propagate.
The team is actively working to increase processing capacity and accelerate synchronizations.
identified
The issue has been identified and a fix is being implemented.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
Directory Dispatch Delays
Started March 31, 2026 at 12:32 AM UTC · 5h 52m
IssuesMinor incident
Affected components
Admin Console - US Region
identified
JumpCloud is currently experiencing dispatch delays for core Directory services. Administrators may experience delays updating associations for users, groups, policies, directories, and commands. We have identified the cause of the issue and are actively implementing a fix.
identified
We have made progress in reducing the backlog affecting Directory services and are continuing to work through the remaining queue. Administrators may still experience delays when updating associations for users, groups, policies, directories, and commands.
The team is actively implementing additional changes to increase processing capacity and accelerate resolution. We will provide further updates in one hour.
identified
The backlog affecting Directory services has been significantly reduced and continues to decrease. Administrators may still experience delays when updating associations for users, groups, policies, directories, and commands.
The team continues to work on additional changes to accelerate processing. We will provide further updates in an hour.
monitoring
The backlog affecting Directory services continues to decrease and we are approaching resolution. Administrators may still experience some delays when updating associations.
The team is continuing to monitor, and our next update will be to confirm full resolution.
resolved
This incident has been resolved.
postmortem

**Date**: Apr 7, 2026
**Date of Incident:** Mar 30, 2026
**Description**: RCA for Directory Association Processing Delays
**Summary:**
Starting March 30th at approximately 15:40 MDT, JumpCloud customers experienced significant delays in directory-related updates. This included latency in password changes, user-to-group associations, and outbound provisioning reflecting in downstream systems. The root cause was identified as a specific code deployment in our Devices service that inadvertently flooded a background processing queue with unpartitioned messages, causing a bottleneck that prevented updates from processing in real-time. The issue was fully resolved by 00:25 MDT on March 31, 2026.
**What Happened:**
The incident was caused by a change in how the JumpCloud agent retrieves software application configurations.
1. **Traffic Spike:** The new code shifted the "source of truth" for these configurations to a new database. If a device polled the system and did not find its record in the new database, the code automatically enqueued a "track collect" request to sync the data.
2. **Unexpected Volume:** We anticipated a "lazy backfill" \(where records are created over time\), but underestimated the number of devices that had no existing software bindings. This resulted in an immediate, massive spike of nearly 280,000 messages.
3. **The Bottleneck \(Partitioning\):** Crucially, these specific messages were enqueued without a "Partition ID." In our high-scale FIFO \(First-In-First-Out\) queue architecture, messages without a partition ID are processed one-by-one rather than in parallel. This effectively "serialized" the queue, preventing us from scaling up workers to process the backlog faster and causing the observed latency.
**Resolution and Recovery**:
Once the offending code was rolled back, the "tap" was turned off, and no further unpartitioned messages were added to the queue.
Because the bottleneck was caused by the lack of partitioning, simply scaling horizontally could not speed up the processing of the existing backlog. The team monitored the queue throughput and determined that the safest and fastest path to recovery was allowing the worker to process the existing messages sequentially rather than risking further disruption by attempting to manually manipulate the production queue.
**Corrective Actions**:
To ensure this type of bottleneck does not occur again, we have committed to the following:
* Improving pre-production testing to better simulate the scale and conditions that can occur in production queue processing
* Reviewing other areas of the platform where similar patterns could produce unexpected request spikes
* Enhancing monitoring and alerting thresholds to enable faster detection and response when queue backlogs begin to form
* Strengthening our deployment validation process to more thoroughly account for background data migrations before releasing dependent code changes
Increased error rates with JumpCloud Agent backend.
Started March 12, 2026 at 4:54 PM UTC · 4h 54m
OutageMajor incident
Affected components
Agent - US Region
investigating
We are currently investigating an issue with delays in syncing user, commands, and policy information with the JumpCloud Devices Agent backend service. We are investigating the cause of the issues currently, and will provide an update within one hour. Agent / Device logins are operating as expected.
investigating
We are continuing to investigate an issue with delays in syncing user, commands, and policy information with the JumpCloud Devices Agent backend service. Local Agent / Device logins without MFA or leveraging TOTP are operating as expected.
identified
The issue has been identified and we have implemented a fix. We are starting to see some recovery with many agents checking in. We will update as our agent traffic normalizes.
monitoring
Agent traffic is reaching normal levels, and the majority of customers are seeing full service restoration. Our engineering teams continue to actively monitor backend stability and traffic patterns. We expect to move to 'Resolved' status shortly as final systems stabilize.
resolved
This incident has been resolved.
postmortem

**Date**: Mar 17, 2026
**Date of Incident:** Mar 12, 2026
**Description**: RCA for Agent Backend \(HAProxy\) System Degradation
**Summary:**
On March 12, 2026, from 10:05 AM to 2:45 PM MDT, JumpCloud experienced a significant service degradation affecting Agent-related activities. During this window, agent updates, including syncing users, passwords, policies and other agent data, as well as new agent installations were unavailable.
This was caused by a "thundering herd" event triggered by a backend traffic-shaping change. We have since identified the root causes and implemented infrastructure changes to prevent a recurrence.
**What Happened?**
At 10:00 AM MDT, our engineering team enabled a feature flag \(a "circuit breaker"\) designed to protect our System Insights API from high load by returning `503 Service Unavailable` responses for certain non-critical requests.
While the flag performed its intended function, it had an unforeseen secondary effect on the JumpCloud Agent’s connection logic. Because the agents could not reuse existing connections for these specific failed requests, hundreds of thousands of agents in our main production environment attempted to establish new mTLS \(mutual TLS\) connections simultaneously.
This created a "Thundering Herd" event that saturated our HAProxy ingress layer, exhausting CPU resources and causing a cascade of connection failures.
**Root Cause:**
The prolonged nature of this incident was the result of three distinct, overlapping bottlenecks that our team had to isolate and resolve one by one:
1. **CPU-Intensive SSL Handshaking:** Establishing an mTLS connection is a CPU-intensive process. The sheer volume of simultaneous connection attempts pushed our HAProxy pods to their resource limits. This caused the pods to become unresponsive, leading to "Out of Memory" \(OOM\) kills and failed health probes.
2. **Health Check Death Spiral:** Our internal health checks initially relied on a Layer 7 SSL validation. Because the CPU was 100% occupied with agent reconnections, the pods couldn't respond to their own health checks in time. This caused the system to erroneously mark healthy pods as "down”, removing them from the rotation and further overwhelming the remaining pods.
3. **Load Balancer Handshake Saturation:** As we attempted to scale our infrastructure, the Application Load Balancer \(ALB\) encountered a throughput bottleneck specifically related to the rate of new connection establishments. The surge of agents attempting to negotiate new SSL handshakes at the same time exceeded the ALB's burst capacity, temporarily preventing even healthy backend pods from receiving and processing traffic.
**Why It Took Time to Resolve:**
While reverting the flag was the correct first step, the agents were already in an aggressive retry loop that continued even after the 503 errors stopped. We had to experiment with several configurations \(adjusting health check intervals and timeout windows\) to find a balance that allowed pods to stay "alive" long enough to process the backlog. Stability was achieved only once we implemented Concurrency Control. By lowering the maximum allowed concurrent connections per pod, we stopped the CPU from over-committing to handshakes, allowing the system to reliably process a controlled flow of traffic until the global queue cleared.
**Corrective Actions / Risk Mitigation:**
**1.\) Edge Infrastructure Hardening**
We are standardized on a new high-availability configuration for our HAProxy ingress layer.
* **Concurrency Governance**: We have implemented a strict maxconn limit per pod. This acts as a "pressure valve," ensuring that the CPU remains available to process existing requests rather than becoming saturated by new connection attempts.
* **Dynamic Capacity Management via Autoscaling**: We are implementing Horizontal Pod Autoscaling \(HPA\) for our HAProxy ingress layer, calibrated to trigger based on both CPU utilization and active connection counts. This ensures we can absorb sudden traffic fluctuations and also maintain a controlled flow of requests to our backend services.
**2.\) Agent Connectivity Optimization**
We are updating the JumpCloud Agent’s communication layer to be more "network-aware" during degraded states:
* **Enhanced Connection Pooling**: We are reconfiguring the agent's HTTP transport logic to maximize the reuse of existing idle connections. This significantly reduces the "Connection Tax" on our backend during high-traffic events.
* **Streamlined Resource Handling**: We are implementing stricter protocols for draining and closing HTTP response bodies, ensuring that pooled connections are returned to the rotation immediately and reliably.
**3.\) Adaptive Retry Logic \(Jitter\)**
To further break up "synchronized" traffic spikes:
* **Introduction of Jitte**r: While our agents currently use exponential backoff for poll requests, we are adding randomized "jitter" to our retry intervals. This spreads reconnection attempts across a wider window, preventing large blocks of agents from hitting the service at the exact same millisecond.
* **Standardizing Resilient Retry Logic:** We are transitioning the Agent’s default HTTP client to a unified **exponential backoff** model for all request types.
* **Controlled Rollou**t: This update will be managed via a staged rollout to monitor for any unforeseen side effects on fleet-wide connectivity patterns.
Console slowness due to CloudFlare Network issue in Asia region
Started November 21, 2025 at 12:34 PM UTC · 1h 57m
IssuesMinor incident
Affected components
Cloudflare CDN/Cache
investigating
We had reports of delays or timeouts when accessing the Admin and User Portal consoles but now recovering.
This issue has been traced to an ongoing network problem within the Asia region reported by Cloudflare.
We are closely monitoring the situation and will provide an update as soon as new information becomes available.
monitoring
Per the latest update from CloudFlare, a fix has been implemented and results are being monitored.
https://www.cloudflarestatus.com/incidents/5p8c3l54b5cl
resolved
This incident has been resolved.
Intermittent issues accessing the JumpCloud Admin portal
Started November 19, 2025 at 6:46 AM UTC · 31m
OutageMajor incident
Affected components
Admin Console - US Region
investigating
We're currently investigating reports of intermittent access issues with JumpCloud's Console Service for administrator at https://console.jumpcloud.com/. We are investigating the cause of the issues currently, and will update the status event once resolved.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
Issue is resolved and closed.
postmortem

**Date**: Nov 21, 2025
**Date of Incident:** Nov 19, 2025
**Description**: RCA for Admin Portal Login Errors
**Summary:**
On November 19, 2025, starting at approximately 04:30 UTC, between 1-5% of requests experienced intermittent failures to successfully authenticate to the Admin Console, lasting until roughly 06:30 UTC. Users attempting to authenticate received an “unexpected” error message during this window, but subsequent retries may have been successful.
**Root Cause:**
This issue was triggered during a standard infrastructure update and traffic shift intended to move services to a new, updated cluster. The core issue was a combination of an infrastructure configuration mismatch and gaps in our detection and validation processes.
1. Configuration Drift: The new infrastructure cluster \(Green Cluster\), intended to host the service, was missing a single but essential configuration value used by the control plane’s service mesh. This value had been recently applied to the existing cluster \(old cluster\) but was inadvertently excluded when the new cluster's baseline configuration was created and branched. When production traffic began routing to the new cluster, the missing configuration caused some access components to fail, leading to the login errors.
2. Detection Gaps: The application logged the configuration failure as a _Warning_ message, rather than a critical error. This meant our automated monitoring system did not trigger an immediate alert or rollback when the issue first occurred.
The team quickly isolated the issue to the new Green Cluster, and an emergency process was initiated to immediately revert all production traffic back to the stable old cluster.
**Corrective Actions / Risk Mitigation:**
1. Automated Configuration Diff Check - Implementing an automated process to continuously compare and ensure 100% configuration parity between old and new production clusters during all transition phases.
2. Clear Rule Enforcement - Reinforcing and automating the process to ensure all configuration changes are applied consistently across all active and future clusters.
3. Multi-Layer Error Monitoring - Implementing error rate monitoring at every layer of the network and application stack to ensure no failure goes undetected.
Remote-Assist Intermittent issues
Started November 11, 2025 at 8:17 AM UTC · 1h 23m
IssuesMinor incident
Affected components
Remote Assist - US Region
investigating
We are currently investigating reports of intermittent issues with launching Remote Assist tool. We are investigating the cause of the issue currently, and will provide an update within one hour.
identified
The issue has been identified and a fix is being implemented.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
SSO-OIDC Authentication Issue
Started November 6, 2025 at 12:28 PM UTC · 1h 19m
IssuesMinor incident
Affected components
SSO OIDC - US Region
investigating
We are currently investigating an issue with the authentication on SSO OIDC applications. We have already identified the cause and will provide an update within one hour.
identified
The issue has been identified and a fix is being implemented.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
postmortem

**Date**: Nov 13, 2025
**Date of Incident:** Nov 6, 2025
**Description**: RCA for SSO/OIDC Service Degradation
**Summary:**
On November 6, 2025, starting at approximately 12:00 UTC, customers experienced failures to launch any application relying on JumpCloud's OIDC-based Single Sign-On \(SSO\), lasting for roughly one hour.
**Root Cause:**
The outage was caused by a combination of two errors during a scheduled compliance procedure:
1. Faulty Password Generation: Our automated system for rotating database passwords created a new credential that contained unsafe special characters.
2. Missing Special Character Logic: The entrypoint script for our core SSO service was missing logic to handle these special characters before using the password to construct a database connection string.
When the SSO service attempted to restart and use the newly rotated password, the presence of the unsafe characters caused the connection string to be misinterpreted as invalid, leading to a parsing failure and service degradation.
This issue stemmed from a latent configuration bug that was masked by prior rotation processes. Previously, database passwords were rotated manually using an older system \(`random_password` IAC resource\) which was explicitly configured to only generate alphanumeric characters. These characters are inherently safe in a URL context, so the underlying bug in the SSO service's connection logic was never exposed. When the credential management was successfully migrated to the new, more robust rotation process, the new function began generating highly complex passwords, including special characters, for the first time. This immediately triggered the latent parsing flaw in the SSO service’s entrypoint script.
**Corrective Actions / Risk Mitigation:**
1. Hardening code logic - All services that construct database connection strings will be audited and updated to explicitly encode the password component eliminating character misinterpretation.
2. Enhanced rotation alerting - New monitoring and alerting dashboards are in place to track the health and success schedule of all automated credential rotation jobs, providing an immediate alert if a rotation creates an invalid credential.
3. Update password generation logic - The automated credential rotation function has been updated to explicitly generate passwords that are safe, avoiding complex, reserved characters.
User Console - US RegionAdmin Console - US RegionRADIUS - US RegionLDAP - US RegionSSO - US RegionTOTP / MFA / JumpCloud Protect - US Region
investigating
We are seeing issues with SSO authentication. We are investigating this currently and will update within 1 hour
investigating
We are continuing to investigate this issue.
identified
We have identified an issue that is causing intermittent login issues with the JumpCloud User Portal and Admin Portal. We have also identified issues with JumpCloud MFA, LDAP, RADIUS, and authentication with SSO. We are working on implementing a fix and will provide another update as soon as possible.
identified
We continue to see intermittent issues with accessing the JumpCloud User and Admin Portal, MFA, LDAP, RADIUS, and SSO. During this time access attempts to LDAP and RADIUS are only impacted if a user is authenticating with MFA.
Our team is working on implementing a fix and will provide another update as quickly as possible.
monitoring
We have implemented a fix and users should now be able to access the User and Admin Console, MFA, LDAP, RADIUS, and SSO without issue. We will continue to monitor the results of the fix.
resolved
Services have been fully restored and this incident has been resolved. We will provide a formal postmortem as a follow up.
postmortem

**Date**: Nov 7, 2025
**Date of Incident:** Nov 4, 2025
**Description**: RCA for Auth Database Degradation
**Summary:**
On November 4, 2025, a number of customers experienced intermittent failures, timeouts and increased latency when attempting to authenticate to multiple JumpCloud Services, including consoles, LDAP, RADIUS and SAML, or use Multi-Factor Authentication.
**Root Cause:**
The incident was triggered by an issue in the deployment process involving a database schema change and a subsequent application code release.
During this deployment, a planned database change unintentionally removed several database indexes required by the existing application code.
The sequence of failure was as follows:
1. Deployment Order Error: The database schema change \(which removed necessary indexes\) was applied to the production database before the new application code \(which did not require those indexes\) was deployed.
2. Performance Collapse: The existing, high-volume authentication code \(used for functions like TOTP and push authentication\) was forced to run against the now-inefficient database structure. Queries that normally took milliseconds suddenly took several seconds.
3. Connection Exhaustion: These slow queries held database connections open for extended periods, quickly overwhelming the database server's available connection pool.
4. Full Outage: With no available connections, the main authentication API could not communicate with the database, leading to 100% CPU utilization on the database server and triggering the intermittent timeouts and failures experienced by our customers.
**Why Testing Did Not Catch This:**
The issue was not identified during testing in our Development or Staging environments due to insufficient Load Simulation. The resource consumption issues and connection exhaustion only manifest under the extreme pressure of peak production traffic volume. The simulated load profiles in our lower environments were not sufficient to expose this specific failure mode.
**Corrective Actions / Risk Mitigation:**
1. Mandatory schema change review - All database schema changes must now undergo an additional level of review to explicitly assess index dependencies and impact.
2. New deployment phasing - We are implementing new tools and checks to enforce that application code dependent on a schema change is deployed before a database change is executed.
3. Enhance alerting - We are implementing new monitors and alerts specifically for the Auth-API's database connection pool health and CPU utilization.
4. Enhanced load testing - We are revisiting the load profiles used in our staging environments looking for opportunities to more accurately simulate peak production traffic.
3rd Party Provider Operational Issues
Started October 20, 2025 at 3:44 PM UTC · 6h 2m
IssuesMinor incident
Affected components
Premium Chat SupportPrivileged Access Management (PAM) - US Region
investigating
Due to an issue with our Cloud Provider, JumpCloud is experiencing intermittent issues with our PAM service and some Support tools which may affect case creation. We are investigating this with our provider and will update this incident as we know more.
identified
We are monitoring our services and continue to work with our provider to restore full functionality.
monitoring
We've seen an increase in recovery and are monitoring
resolved
This incident has been resolved.
postmortem

**Date**: Oct 27, 2025
**Date of Incident:** Oct 20, 2025
**Description**: RCA for Service Degradation Linked to AWS US-EAST-1 Regional Disruption
**Summary:**
On October 20, 2025, between approximately 07:00 UTC and 10:00 UTC, the JumpCloud platform experienced significant performance degradation. This primarily affected the responsiveness of our core APIs, administrative console access, user console access, single sign-on \(SSO\) functions for customers, failed trial creation, user updates, and the potential loss of some Directory Insights events. Residual performance degradation, affecting services like our Privileged Access Management \(PAM\) feature and specific Support Portal APIs, persisted until approximately 17:00 UTC, at which point all services were fully restored.
The degradation was not caused by a failure within the JumpCloud platform code or infrastructure configuration, but was a direct consequence of a severe, cascading failure within the Amazon Web Services \(AWS\) US-EAST-1 region. Although many of our services are deployed across multiple Availability Zones \(AZs\) within the region for resilience, the nature of the AWS issue - impacting fundamental regional services - compromised inter-AZ communication preventing some of our standard failover mechanisms from operating successfully. Services returned to full operational status after AWS reported stability with the impacted foundational services.
**Root Cause:**
Based on the [post-incident analysis of AWS,](https://aws.amazon.com/message/101925/) the service disruption was not a single event but a sequence of cascading failures across three fundamental AWS services.
1. DynamoDB Failure Due to Latent DNS Race Condition \(Initial Trigger\)
2. EC2 Launch Congestion and Network Propagation Delays \(Sustained Impact\)
3. Network Load Balancer \(NLB\) Health Check Instability \(Final Phase\)
All affected service teams were paged by our monitoring and alerting systems, and our incident management team was engaged to coordinate efforts and assess areas where we could throttle traffic to stabilize remaining capacity.
**Corrective Actions / Risk Mitigation:**
1. While our current architecture mitigates some single-Availability Zone \(AZ\) failures, our primary focus is to eliminate single-region dependency entirely. JumpCloud is actively engaged in strategic engineering initiatives to further strengthen the platform's foundation. We are conducting a thorough review of cross-region dependencies and replication strategies to enhance our service's resilience against widespread environmental disruptions, always striving to meet the highest standards of availability.
3rd Party Provider Operational Issues
Started October 20, 2025 at 7:52 AM UTC · 2h 22m
IssuesMinor incident
Affected components
User Console - US RegionAdmin Console - US RegionAdmin Console - EU RegionUser Console - EU Region
investigating
Due to an issue with our Cloud Provider, JumpCloud is experiencing intermittent issues with multiple services. We are investigating this with our provider and will update this incident as we know more.
identified
We are monitoring our services, and starting to see some recovery.
monitoring
We are continuing to monitor as we see further recovery with all services.
resolved
This incident has been resolved.
postmortem

**Date**: Oct 27, 2025
**Date of Incident:** Oct 20, 2025
**Description**: RCA for Service Degradation Linked to AWS US-EAST-1 Regional Disruption
**Summary:**
On October 20, 2025, between approximately 07:00 UTC and 10:00 UTC, the JumpCloud platform experienced significant performance degradation. This primarily affected the responsiveness of our core APIs, administrative console access, user console access, single sign-on \(SSO\) functions for customers, failed trial creation, user updates, and the potential loss of some Directory Insights events. Residual performance degradation, affecting services like our Privileged Access Management \(PAM\) feature and specific Support Portal APIs, persisted until approximately 17:00 UTC, at which point all services were fully restored.
The degradation was not caused by a failure within the JumpCloud platform code or infrastructure configuration, but was a direct consequence of a severe, cascading failure within the Amazon Web Services \(AWS\) US-EAST-1 region. Although many of our services are deployed across multiple Availability Zones \(AZs\) within the region for resilience, the nature of the AWS issue - impacting fundamental regional services - compromised inter-AZ communication preventing some of our standard failover mechanisms from operating successfully. Services returned to full operational status after AWS reported stability with the impacted foundational services.
**Root Cause:**
Based on the [post-incident analysis of AWS,](https://aws.amazon.com/message/101925/) the service disruption was not a single event but a sequence of cascading failures across three fundamental AWS services.
1. DynamoDB Failure Due to Latent DNS Race Condition \(Initial Trigger\)
2. EC2 Launch Congestion and Network Propagation Delays \(Sustained Impact\)
3. Network Load Balancer \(NLB\) Health Check Instability \(Final Phase\)
All affected service teams were paged by our monitoring and alerting systems, and our incident management team was engaged to coordinate efforts and assess areas where we could throttle traffic to stabilize remaining capacity.
**Corrective Actions / Risk Mitigation:**
1. While our current architecture mitigates some single-Availability Zone \(AZ\) failures, our primary focus is to eliminate single-region dependency entirely. JumpCloud is actively engaged in strategic engineering initiatives to further strengthen the platform's foundation. We are conducting a thorough review of cross-region dependencies and replication strategies to enhance our service's resilience against widespread environmental disruptions, always striving to meet the highest standards of availability.
Navigation affected at the JumpCloud website
Started October 14, 2025 at 8:52 AM UTC · 5h 26m
IssuesMinor incident
investigating
We are currently investigating an issue, where navigating links (specifically redirects) at the JumpCloud website are affected.
We are suspecting an ongoing "Bot Management Cookie" incident declared by CloudFlare to be at the root.
More details can be found at https://www.cloudflarestatus.com
We will update as soon as CloudFlare has fixed the issue.
identified
CloudFlare has identified the issue and is implementing a fix. We are monitoring https://www.cloudflarestatus.com for updates.
identified
CloudFlare continues to work on the issue and we are monitoring their status page at https://www.cloudflarestatus.com
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
JumpCloud Remote Assist Unavailable
Started October 9, 2025 at 2:20 PM UTC · 20m
Pending
Affected components
Remote Assist - EU RegionRemote Assist - US Region
investigating
We are currently investigating an issue where the Remote Assist and Background tools actions are unavailable in the Admin Console.
investigating
We are continuing to investigate this issue.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
Admin Portal Performance
Started October 7, 2025 at 9:31 AM UTC · 1h 52m
IssuesMinor incident
Affected components
Admin Console - US Region
investigating
We are investigating slow performance with the Admin Portal. We will update within an hour.
monitoring
Our investigation uncovered intermittent latencies at several end-points which caused slow response times at the Admin portal for a few customers.
Out of an abundance of caution, we have rolled back a recent update and will continue monitoring.
resolved
The issue is now marked as resolved.
LDAP Degredation
Started October 5, 2025 at 10:00 AM UTC · 6h 21m
Pending
Affected components
LDAP - US Region
investigating
Currently investigating LDAP performance degradation.
investigating
We are currently investigating failures to authenticate using LDAP. We are investigating the cause of the issues currently, and will provide an update within one hour.
identified
We are continuing to investigate failures to authenticate using LDAP. We have found the cause of the issue and are working to resolve the situation. We will provide an update within one hour.
identified
We are currently seeing recovery in the LDAP service. We will continue to monitor this situation.
resolved
This incident has been resolved.
postmortem

**Date**: Oct 8, 2025
**Date of Incident:** Oct 5, 2025
**Description**: RCA for LDAP authentication failures
**Summary:**
On October 5, 2025, a number of customers experienced intermittent failures when attempting to authenticate to LDAP. Users and services attempting to authenticate received an error message indicating a failure to successfully establish a connection.
**Root Cause:**
The incident was caused by a failure in our automated certificate renewal process.
1. Certificate Expiration: A critical internal Transport Layer Security certificate, used for secure communication within our infrastructure, expired.
2. Automation Failure: The automated system responsible for proactively renewing this certificate failed to run its scheduled update.
3. Cascading Affect: Because the core certificate was not renewed, dependent LDAP services could not renew their own certificates, leading to connection failures with our core database and security vault.
The team manually executed the renewal script to update and deploy the expired certificate across all necessary servers, and restarted services on systems that did not pick up the new certificates immediately, restoring normal operation to the LDAP services.
**Corrective Actions / Risk Mitigation:**
1. Immediately execute the renewal script and restart services - DONE
2. Implementing dedicated, proactive alerting on the expiration dates of these infrastructure certificates - IN PROGRESS.
3. Add automation checks that verifies the successful execution of the certificate renewal process - IN PROGRESS
Software Inventory Delay
Started September 18, 2025 at 11:57 AM UTC · 2h 38m
IssuesMinor incident
Affected components
System Insights - US Region
investigating
We are aware of a delay in software inventory data updating. We are currently working on this and aim to update within 1 hour.
identified
We have identified the issue and are applying a fix to resolve this issue. We will update within an hour.
identified
We are still applying the fix. We will update as soon as the delay reduces.
monitoring
The fix has been applied. Software Inventory data is now within the acceptable time range. We will continue to monitor.
resolved
The data is now remaining within the acceptable range. We consider this to be resolved.
Admin Portal Access Issues
Started September 15, 2025 at 7:55 AM UTC · 0m
OutageMajor incident
resolved
There were issues with the following services; Admin Portal & Mobile Admin App Login & Direct API calls. This is resolved now.
postmortem

# **Incident Report**
**Date**: Sep 15, 2025
**Date of Incident:** Sep 16, 2025
**Description**: RCA for Admin Portal Access Failures
**Summary:**
On September 15, 2025, between 08:59 UTC - 09:11 UTC customers experienced failures when attempting to access the Admin Portal. Users and services attempting to authenticate received a loading error message during that window.
**Root Cause:**
The incident originated during an update to support a new API for our SaaS Management Service. The update was designed to allow a new console subdomain to handle traffic, in addition to the existing SaaS subdomain.
While this change successfully passed all testing in our pre-production environment, our automated deployment system failed to apply the full configuration in production. It applied the new console subdomain but, due to a misconfiguration, it omitted a critical path prefix. This resulted in an issue where traffic intended for a different service was routed incorrectly, causing a brief service outage.
As soon as the errors appeared, the team immediately rolled back the change, and service was restored. We have identified the root cause in our automated deployment process and are implementing a fix to prevent this from happening again.