We're currently seeing heightened alerts related to connectivity between zones, we're engaging our infrastructure team to investigate further. We'll update here as soon as we know more, in the meantime please raise a ticket with us if you're experiencing any degradation to service.
resolved
The issues we saw should now be resolved; this was limited to hosts in the SJC zone. We will send out a full RCA in due course, please let us know if you experience any further issues.
We are currently investigating Zabbix alerts/reports of outages affecting portal/calls. We will send out an update as ASAP.
investigating
This appears to a problem with Level3 currently; our infrastructure team is investigating implementing routing changes to workaround the issue.
investigating
Cloudflare appears to have stopped advertising our 8.X addresses causing this outage; we have now re-enabled them. We're working with CloudFlare to investigate the root cause but are seeing issues resolve from our side.
resolved
This incident has been resolved.
Degraded performance in ORD zone due to Cloudflare issues
Początek 27 stycznia 2026 20:26 UTC · 20h 57m
IssuesDrobny incydent
Dotknięte komponenty
Telephony Services
identified
We are seeing degraded performance and intermittent issues in the ORD zone due to a Cloudflare issue: https://www.cloudflarestatus.com/
We are monitoring the situation. Users who are experiencing issues can route away from the ORD zone by cutting over to the east or west zones.
monitoring
Cloudflare has reported a fix has been implemented. We are monitoring for issues.
We are investigating an issue with a partial outage due to loss of connectivity to one of our zones.
investigating
We are experiencing uptime issues with our San Jose datacenter and are working to resolve it. Phones that are properly configured to register to SRV/NAPTR records should fail over to the other zones.
investigating
We are seeing a fiber related issue in our San Jose datacenter and are working to resolve it.
monitoring
We have resolved the issue and believe we are operational at this time. We are monitoring the incident.
resolved
This incident has been resolved.
ZSwitch SMS Issues
Początek 2 września 2025 17:11 UTC · 2h 6m
OutagePoważny incydent
Dotknięte komponenty
Telephony Services
investigating
We've observed issues with Inbound/Outbound SMS and MMS and have confirmed this appears to be consistently impacting all of ZSwitch.
resolved
This incident has been resolved.
Zswitch US outage.
Początek 20 sierpnia 2025 10:50 UTC · 11h 12m
OutageKrytyczny incydent
Dotknięte komponenty
Telephony Services
investigating
We are currently investigating an issue with Zswitch US' database services. We have engaged our engineering and operations team to work this as highest priority. We will update this page as soon as we know more.
investigating
This issue has now been resolved and the database services are operating as expected. We are working on a full RCA and will update this statuspage with it in due course. If you have any questions in the meantime please contact support using the normal methods.
monitoring
This issue has now been resolved and the database services are operating as expected. We are working on a full RCA and will update this statuspage with it in due course. If you have any questions in the meantime please contact support using the normal methods.
We will continue to monitor throughout the day.
monitoring
We've stabilized services and are continuing to monitor the situation.
monitoring
We are seeing further degradation in the system and are investigating
monitoring
We have identified a database system outage and are working with our engineers to restore the system.
monitoring
We have brought services back online and are continuing to monitor for stability.
resolved
This incident has been resolved. We will continue to monitor for stability.
Possible issues
Początek 30 września 2024 16:03 UTC · 3d 23h
IssuesDrobny incydent
Dotknięte komponenty
Telephony Services
investigating
Please be advised that we are currently hearing of widespread ISP issues across the country. Major media outlets are reporting Verizon-related outages, but we suspect the impact is at least somewhat wider. No 2600Hz services are directly impacted at this time. However, resellers can expect to see increased reports from clients, especially Verizon clients
monitoring
Please be advised that we are currently hearing of widespread ISP issues across the country. Major media outlets are reporting Verizon-related outages, but we suspect the impact is at least somewhat wider. No 2600Hz services are directly impacted at this time. However, resellers can expect to see increased reports from clients, especially Verizon clients
resolved
This incident has been resolved.
Investigating call center issues
Początek 7 maja 2024 21:11 UTC · 2h 12m
Pending
Dotknięte komponenty
Management Portal
investigating
We are currently investigating reports of issues with loading Call Center in our UI and comm.land.
monitoring
We have identified that we were receiving alerts from two of our apps servers having issues connecting to each other. We've restarted Kazoo Apps on both servers. We've received confirmation from some customers that this is resolved, we've received no new reports, and we've been unable to replicate since the restarts. We're monitoring to ensure this is resolved.
monitoring
We are continuing to monitor for any further issues.
resolved
This incident has been resolved.
Investigating issues on EWR
Początek 9 kwietnia 2024 21:08 UTC · 2h 34m
Pending
Dotknięte komponenty
Telephony Services
investigating
We are investigating reports of degraded call services in our EWR zone
monitoring
We have mitigated the issue and our engineers are working to identify the root cause
resolved
This incident has been resolved.
Reports Of Call Completion Issues in EWR
Początek 27 lutego 2024 21:06 UTC · 2h 59m
IssuesDrobny incydent
Dotknięte komponenty
Telephony Services
investigating
We are receiving reports of call completion issues on EWR. We are investigating.
monitoring
We have mitigated the issue and are monitoring the system for any further issues.
resolved
This incident has been resolved.
Unanswered semi-attended transfers not hitting voice mail
Początek 16 stycznia 2024 21:34 UTC · 3h 55m
OutagePoważny incydent
Dotknięte komponenty
Telephony Services
investigating
We are seeing increasing reports of semi-attended transfers not reaching the recipient's voice mail if the recipient does not answer.
This is impacting our Hosted Platform only and we are currently investigating the issue.
identified
We believe we have identified the issue. We have deployed a patch to the SJC zone and are currently testing before we deploy to ORD and EWR.
monitoring
Testing has been successful so far and we've deployed the patch to all remaining Hosted Platform servers. We believe this to be fixed now, but will continue monitoring for any further issues.
resolved
This incident has been resolved.
Increased Reports Of Call Completion Issues using WSS
Początek 5 stycznia 2024 14:50 UTC · 7h 9m
IssuesDrobny incydent
Dotknięte komponenty
Telephony Services
investigating
We are receiving reports of call completion issues over web sockets registered devices including comm.land phones
monitoring
At this time we believe this issue affected clients that were using comm.land AND immediately updated their comm.land software this morning. This issue should now be resolved, we are monitoring the situation to validate.
resolved
We've had multiple customer reports that this issue is now resolved
Single Kamailio Server - No TCP
Początek 14 sierpnia 2023 22:49 UTC · 1d 3h
IssuesDrobny incydent
Dotknięte komponenty
Telephony Services
investigating
Presently, a concern has arisen wherein one out of the three Kamailio servers located in EWR is encountering difficulty in processing SIP registrations. It's important to note that this issue is not expected to affect our services, as phones that are appropriately configured will automatically attempt registration through alternate SBCs and/or zones. If you're currently facing challenges with phone registrations, we encourage you to reach out to our support team. This way, we can assist you in enhancing your phone's configuration to proactively avoid similar issues moving forward.
monitoring
We would like to inform you that several clients have reported successful phone registrations after we implemented a TCP traffic block to and from this specific SBC. This leads us to believe that the issue is currently fully mitigated. In light of this, we are intentionally leaving this SBC in a degraded state to facilitate thorough investigation by our engineering teams as required. Our intention is to conduct a server reboot tomorrow evening. We will keep this incident active until that time to facilitate full transparency into the issue.
resolved
This incident has been resolved.
Increased reports of call failures in EWR
Początek 22 czerwca 2023 12:00 UTC · 0m
Pending
resolved
We are currently investigating this issue.
postmortem
After investigation Kamailio was restarted which resolved the issue - Please let us know if you encounter any further disruption. We are working on investigating the root cause.
Increased Reports Of One Way Audio Issues in EWR datacenter
Początek 15 czerwca 2023 15:57 UTC · 1h 57m
Pending
Dotknięte komponenty
Telephony Services
investigating
We are receiving increased Reports Of One Way Audio Issues in EWR datacenter
investigating
At this time we do believe this is isolated to clients on a single ISP. We are continuing to investigate
identified
We do believe this issue is with a particular peer. We are advertising new BGP routes to attempt to work around this issue
resolved
We have multiple confirmations that our workaround has resolved the issue
Increased Reports Of Call Completion Issues in EWR
Początek 2 marca 2023 17:19 UTC · 5m
IssuesDrobny incydent
Dotknięte komponenty
Telephony Services
investigating
We are receiving reports of call completion issues in EWR
monitoring
We have paused the EWR zone and are monitoring for any further issues
resolved
This incident has been resolved.
US-West and US-East connectivity
Początek 1 marca 2023 21:30 UTC · 0m
IssuesDrobny incydent
resolved
At 13:13:56 PT we received alerts that multiple BGP sessions had dropped in our US-West datacenter. This included our backhaul between US-East and US-West, as well as one of our outbound ISPs. The backhaul traffic failed over to the secondary by 13:14:03 PT, and the session with the outbound ISP recovered by 13:25:21 PT. We believe there was no impact on inbound, or outbound traffic to/from the 2600Hz network during this incident.
We have engaged our providers to get more information on what occurred, but all seems to have recovered by 13:25:21 PT.
Thank you,
2600Hz Infrastructure
Increased Reports Of Call Completion Issues in EWR
Początek 16 lutego 2023 21:12 UTC · 1h 2m
IssuesDrobny incydent
Dotknięte komponenty
Telephony Services
investigating
We are receiving reports of call completion issues in EWR.
monitoring
We have paused the EWR zone in full and are redirecting traffic arriving in EWR to ORD in Kamillio. Call completion rates seem to be returning to normal levels
monitoring
We have had no further reports of issues and all alerts cleared. Please write into support if you have any example calls after 13:35 PT
resolved
Multiple clients are reporting no further issues
postmortem
On Thursday the 16th of February we experienced a Zswitch outage between 19:25-19:45 GMT.
As soon as we noticed alerts appearing for disconnects a 911 bridge was started with Support/Operations/Engineering/ the CTO.
This appeared to be localised to EWR; we therefore paused all EWR servers as a precaution which remediated the issues with calls.
From here we troubleshooted to determine what could have caused the outage; it was realised compaction was started on the bc003.ewr server only a few minutes beforehand; we could also see that the load was unreasonably high on this server. The compaction was stopped which brought the load down to a stable level. We also noticed that this BigCouch server was delivering DB responses in the 5-6 second range, rather than the below 50ms we would expect. This was due to the server not behaving in the expected manner when compaction was running.
Further tests were completed which helped us to determine this was an issue with the server itself and put a plan in to migrate it over to new hardware.
After the migration was complete, we further stress tested the new server to confirm compaction did not cause any issues. It’s clear that this server was major component in the 4 most recent outages.
We have added new alerting which will help us to have higher visibility of response times from all servers to Bigcouch as this will allow us to notice any similar issues before they start to impact. We are also further improving this alerting to exclude crossbar \(API\) calls which will help keep out any false positives, as well as adding similar active testing across all HA Layers in Kazoo.
During the holiday weekend we staged a “All hands on deck” meeting to go over the outages we’ve seen and the steps we can put in place to improve. We want to convey that the seriousness of the downtime we have seen is understood and we’re continuing to work to prevent any recurrences. Following this call we have come out with a list of action items that we’ll be undertaking with the highest priority; including the aforementioned monitoring changes, and reprioritising engineering load to speed up the migration over to CouchDB 3 from BigCouch.
If anyone has any further questions, or you’d like deeper information on what steps we’re taking to prevent further outages please don’t hesitate to get in touch.
Cluster Wide Outage
Początek 2 lutego 2023 20:30 UTC · 3h 18m
OutageKrytyczny incydent
Dotknięte komponenty
Management PortalTelephony Services
investigating
We are receiving and have verified reports that there are call failures happening in all zones on ZSwitch
investigating
We are continuing to investigate this issue.
investigating
We have identified this as a database issue. We are still investigating the root cause
identified
We are having success improving the DB response rate. Though we are still seeing a small percentage of intermittent failures
resolved
We believe the issue is now resolved. We are continuing to monitor services
postmortem
On Thursday 2nd Feb we encountered an issue on Zswitch causing call failures and issues accessing the UI. We were first alerted to issues by our monitoring system; telling us there were errors on FS within EWR. As a precautionary measure we decided to pause FS on the alerting nodes. The issue instantly spread to the rest of the FS servers; at this point we involved the engineering team and jumped on an all hands "911" call.
Our engineering, operations and support teams further investigated the issues which uncovered that BigCouch seemed to be in a bad state. \(We could see when trying to pull information from the DB, the apps were occasionally getting timeouts or at least seeing very long response times.\) This again pointed to the databases being overloaded. All of the DB nodes were checked, all compactions stopped, and a slow restart of all the DB nodes was carried out.
After the DB's were back up and working we found the following culmination of tasks caused the issues. \(We would like to stress that any individual, or even two of these tasks wouldn't usually cause issues; this is a very edge case scenario.\) -
1. The "DB of DB's" \(A file called dbs.couch which contains all the locations of all shards within the database\) was much larger than it should be. This would cause DB tasks not to be completed as quickly and left BigCouch in a more fragile state.\) To shrink this a separate compaction task needs to be carried out. Separate to the usual compaction carried out on a regular basis.
2. A new feature was added into the latest version to be able to bill for ephemeral tokens executed an inefficient implementation of a monthly roll up which put heavy load on the already struggling DB.
3. Two compaction tasks were running on separate nodes throughout the cluster to address disk space alerts. \(This is a normal and routine task, although the culmination of the two other issues caused heavier load on the DB\).
To fully resolve the issue the DB of DB's was subsequently compacted on all of the databases followed by a restart of BigCouch.
To prevent this issue from reoccurring we are taking the following steps in the short term -
1. A full audit of the DB compaction procedure to make sure we are compacting everything which needs to be and there's no way BigCouch can reach the same fragile state it did again.
2. Our engineering team are fixing the ephemeral report generator so it's not as aggressive with the DB.
In the long term, we have a plan to upgrade all servers onto CouchDB 3, which will compact automatically in a much more proficient manner.
We again apologise for the interruption to service, please be assured we working hard to improve our service. We strive to provide the best telecommunications platform available. This was as mentioned previously a very edge case scenario, a perfect storm of DB tasks which resulted in instability; any two of the tasks outlined we are confident wouldn't have caused an issue. If anyone has any further questions, please do send a ticket into support. We'll be more than happy to address any concerns and go into further detail with the steps we're taking to improve.