Intermittent temps d'arrêt / difficulté d'accès ChargeOver
Début 24 août 2026 à 21:18 UTC · 37m
OutageIncident majeur
Composants affectés
Main Application
investigating
Nous enquêtons actuellement sur cette question.
monitoring
La plate-forme ChargeOver se stabilise et nous suivons la situation de près. Plus de mises à jour à suivre.
resolved
Cet incident a été résolu.
Un post mortem suivra.
Traduit automatiquement depuis la mise à jour officielle de l'incident.
ChargeOver est inaccessible
Début 1 août 2026 à 06:24 UTC · 8h 34m
OutageIncident critique
Composants affectés
Main Application
identified
ChargeOver est actuellement inaccessible.
Nous travaillons sur une résolution dès que possible.
monitoring
ChargeOver service a été restauré.
La production automatisée de factures et le traitement automatisé des paiements sont retardés, mais ils se rattraperont rapidement et ne nécessiteront aucune intervention manuelle de la part des commerçants. La synchronisation des données avec d'autres plates-formes (QuickBooks, Xero, Salesforce) sera retardée, mais se rattrapera rapidement, et ne nécessitera aucun re-sync manuel.
D'autres mises à jour et un postmortem seront affichés plus tard.
monitoring
Nous restons à l'affût de tout autre problème.
monitoring
Nous continuons de surveiller toute autre question.
Les processus automatisés (créer des factures; effectuer des paiements; synchroniser avec QuickBooks/Xero/Salesforce) continuent de rattraper leur retard.
resolved
Cet incident a été résolu et tous les services fonctionnent normalement maintenant.
Un post mortem suivra.
Traduit automatiquement depuis la mise à jour officielle de l'incident.
Network provider outage
Début 14 avril 2026 à 04:46 UTC · 7h 58m
OutageIncident critique
Composants affectés
Main Application
investigating
ChargeOver is currently unreachable. Our upstream network provider is experiencing an outage affecting connectivity to our data center. Our systems are online and operational; this is a network issue outside of our infrastructure.
Our provider is aware of the issue and is actively working toward a resolution. We do not have an ETA at this time.
We will post an update every 60 minutes until this is resolved.
investigating
No change in status. Flexential has confirmed that carrier circuits at their Chaska, MN facility remain down. Their network engineers are engaged with the carrier vendor, who is actively investigating. No restoration ETA has been provided.
Our systems remain online and operational. This continues to be a network connectivity issue upstream of our infrastructure.
Next update in 60 minutes or sooner if status changes.
investigating
Progress has been made in identifying the cause. Flexential's carrier vendor has isolated a faulty optical component on the affected circuit and a technician is en route to replace it. No restoration ETA has been provided yet.
This continues to be a network connectivity issue upstream of our infrastructure.
Next update in 60 minutes or sooner if status changes.
investigating
No change in status. This continues to be a network connectivity issue upstream of our infrastructure.
Next update in 60 minutes or sooner if status changes.
investigating
ChargeOver is accessible again. Flexential has restored network connectivity to our data center and we are actively monitoring performance and stability. We have not yet received an official all-clear from Flexential.
We will post a final update once we have confirmed full resolution.
monitoring
We are continuing to monitor connectivity, and await an update from our primary data center (Flexential).
ChargeOver is operating normally at this time.
Some payments may need to be re-attempted. Some syncs to external platforms (QuickBooks, Xero, Salesforce, etc.) may be delayed or need re-sync.
resolved
Connectivity issues have been resolved.
Some payments may need to be re-attempted. Some syncs to external platforms (QuickBooks, Xero, Salesforce, etc.) may need re-sync.
ChargeOver unavailable for some users
Début 26 mars 2026 à 16:57 UTC · 21m
OutageIncident majeur
Composants affectés
Payment ProcessingMain Application
investigating
We are investigating the issue.
We are expecting to have this resolved within the next 30 minutes.
resolved
This has been resolved.
We are investigating what occurred here still. A postmortem will follow.
Problem accessing our help documentation
Début 26 décembre 2025 à 22:57 UTC · 32m
Pending
Composants affectés
Developer Docs
investigating
We are currently investigating this issue.
investigating
We are continuing to investigate this issue.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
Emails failing to send with a "Invalid handle provided" error message
Début 11 novembre 2025 à 18:21 UTC · 1h 51m
IssuesIncident mineur
Composants affectés
Email Sending
investigating
We are currently investigating this issue.
Some emails are not sending, instead triggering a "Invalid handle provided" message.
Payment processing and integrations are unaffected.
identified
We have identified the problem, and are expecting a resolution within the next 30 minutes.
We will provide an update here again shortly.
monitoring
This issue should be resolved now. We are monitoring this for any further action.
A postmortem will follow.
monitoring
We are continuing to monitor for any further issues.
resolved
This incident has been resolved.
A postmortem will follow.
Delayed sync of data to QuickBooks Online, Xero, and Salesforce for approximately 30% of accounts
Début 30 octobre 2025 à 13:41 UTC · 5h 53m
IssuesIncident mineur
Composants affectés
Integrations
identified
We have identified a component failure which is delaying the sync of data to some systems for approximately 30% of ChargeOver accounts.
The affected integrations are:
* QuickBooks Online
* Xero
* Salesforce
Syncing will catch up automatically, but may be delayed.
We are still investigating root cause and will follow up with a postmortem.
monitoring
Syncing has completely caught up for all but approximately 10% of ChargeOver accounts.
The remaining ChargeOver accounts continue to catch up quickly, but are still delayed.
We are continuing to monitor the situation, and a postmortem will follow.
resolved
Syncing for all affected integrations is complete now.
A postmortem will follow.
ChargeOver is down for some users.
We are investigating.
monitoring
ChargeOver is accessible again.
We are investigating root cause and investigating the outage.
resolved
ChargeOver was inaccessible for most users from approximately 9:44am CT to 9:58am CT, due to a database connection issue.
No data loss occurred. Scheduled subscription invoice creation and payment processing were not affected - all scheduled payments processed normally.
We are still investigating root cause. A postmortem will follow.
Login outage
Début 18 juillet 2025 à 19:11 UTC · 1h 3m
OutageIncident critique
Composants affectés
Main Application
investigating
We are currently investigating this issue.
identified
We believe we have identified an issue with third party analytics service causing the login process to take longer than expected to give users access to their instance.
monitoring
Our third party analytics service was not being responsive and we've implemented a fix to bypass this issue. Logins should resume as expected. We're monitoring and will update with information as it becomes available.
resolved
ChargeOver had longer than expected wait times from 14:11 CDT to 15:07 CDT.
We're monitoring the solution to the root cause. A postmortem will follow.
postmortem
# Incident details
We received messages from internal & external users that they were not able to log into their instances. After investigation, we confirmed that the login sequence was taking an abnormal amount of time. This behavior was present on July 18th from `14:00 CST` to `15:14 CST`.
The impact of this incident was limited to only the login process for ChargeOver. This did not affect payment processing, hosted pages, or other automated processes.
# Root cause
ChargeOver uses a third party metrics service, PostHog, to track user input to help improve the platform.
Requests to PostHog were determined to be the cause of the slow login sequence. We found a URL endpoint was no longer being serviced, updated the pointing of our load balancer, and rebooted any systems in order for the change to take affect.
# Incident timeline
* 14:00 CDT - Received notice of logins not working.
* 14:11 CDT - Assembled a team to start investigating root causes and update our status page.
* 14:44 CDT - Identified the issue with the third party metrics service. Determined that logins were working, but they were taking significantly longer than expected.
* 15:07 CDT - Implemented fix for third party metrics service and deployed to user instances. Began monitoring for any further issues.
* 15:14 CDT - Verified that logins were working as expected and began drawing up plans for future remediation.
# Remediation plan
We understand that it is frustrating to lose access to your ChargeOver instance and we are very sorry that this has happened.
In the future, we will handle these requests so that they do not get in the way of normal use by separating the logic independently of the rest of ChargeOver. We will also be implementing an SOP to make sure that any services that may cause disruptions won't affect the app in this way.
ChargeOver inaccessible / network outage
Début 5 juillet 2025 à 08:30 UTC · 0m
Pending
resolved
ChargeOver was inaccessible from approx 3:27 CDT, to 4:57 CDT.
This was a networking-system related outage.
We are still investigating root cause. A postmortem will follow.
postmortem
# Incident details
On July 5th, a firewall-related issue caused ChargeOver to be inaccessible for approximately 90 minutes.
Customers were unable to log in or access the ChargeOver application during this time frame.
# Root cause
At approximately 3:27pm CT, one of the high-availability firewalls experienced an error which caused it to stop passing traffic into/out-of ChargeOver’s internal network. Although the firewall was configured in a way which should trigger automatic fail-over \(we use industry standard Netgate firewalls, which use CARP to share IP and network status information and monitor for fail-over\) to a secondary firewall, the automatic fail-over did not occur.
ChargeOver staff were immediately and automatically notified, and after troubleshooting were able to resolve the outage by power-cycling the affected firewall.
# Incident timeline
* 3:27pm CT - Firewall experiences an error, which stops some inbound/outbound traffic. ChargeOver staff are immediately alerted.
* 3:28pm CT - ChargeOver staff respond and begin troubleshooting.
* 4:57pm CT - Problem resolved after power-cycling the affected firewall.
# Remediation plan
We are still investigating what caused the error which took the firewall offline, as well as why the firewall did not fail over correctly to a secondary firewall when the error occurred. As we investigate further, we’ll have a better idea of whether this requires replacement or reconfiguration of the firewall.
We are improving documentation, to enable our engineering team to troubleshoot quicker/isolate root cause quicker for future incidents.
We are also working with data center staff to establish some better monitoring tools to help us troubleshoot quicker/isolate root cause quicker.
ChargeOver inaccessible
Début 8 juin 2025 à 18:54 UTC · 4h 29m
OutageIncident critique
Composants affectés
Main Application
identified
We are working with our primary data center to resolve the issue.
At this time, we believe this to be a carrier networking issue at the data center itself, outside of ChargeOver's infrastructure.
resolved
This incident has been resolved.
A postmortem will follow.
postmortem
We have posted a postmortem for this incident here:
* [https://status.chargeover.com/incidents/3zfsdz4msf3s](https://status.chargeover.com/incidents/3zfsdz4msf3s)
We have identified a networking issue, and are working on a resolution.
monitoring
A networking fix has been applied, and services are coming back up.
We are monitoring to ensure everything is accessible.
monitoring
We are continuing to monitor for any further issues.
resolved
This incident has been resolved.
A postmortem will follow.
postmortem
## Incident details
ChargeOver experienced an outage on June 7th, and an outage again on June 8th, as a result of a planned data center power outage, and an unplanned networking stack -related failure.
On Saturday, June 7th, ChargeOver was inaccessible from approximately 13:12 CT to 16:53 CT.
On Sunday, June 8th, ChargeOver was inaccessible from approximately 13:32 CT to 17:05 CT.
## Root cause
ChargeOver’s primary data center \(Flexential in Chaska, MN\) performed a planned maintenance task which de-energized the four uninterruptible power supply \(UPS\) systems serving entire data center, one UPS system at a time.
ChargeOver’s network infrastructure is designed to be redundant, and so ChargeOver has two separate power feeds coming from the four data center UPS systems. ChargeOver’s stacked network switches are plugged into either the first power feed, or the second power feed, to provide redundancy.
When one of ChargeOver’s redundant network switches lost power due to the data center UPS maintenance, the expectation was that the switch would automatically fail over to the second redundant switch. Either the switches did not fail over correctly, or the switch lost it’s configuration on fail-over, which resulted in ChargeOver’s systems becoming unavailable to the outside world.
## Incident timeline
* May 29 - Flexential informed ChargeOver of anticipated maintenance. Due to a miscommunication, it was unclear to ChargeOver staff that the maintenance would impact our environment.
* June 7 - 13:12 CT - Flexential de-energizes and then re-energizes the first power feed, causing a network switch to go offline. The network switch either fails to fail-over correctly, or loses it’s configuration as the power feed comes back on. ChargeOver staff are immediately notified.
* June 7 - 13:19 CT - Attempts to reset the network switch fail to restore service.
* June 7 - 16:51 CT - One network switch is removed, and the configuration on the second switch is re-loaded.
* June 7 - 16:51 CT - Network availability is restored.
* June 7 - 13:32 CT - Flexential de-energizes and then re-energizes the second power feed, causing a network switch to go offline. The network switch either fails to properly boot, or re-loses it’s configuration as the power feed comes back on. ChargeOver staff are immediately notified.
* June 7 - 14:41 CT - Attempts to reset the network switch fail to restore service.
* June 7 - 17:03 CT - The configuration on the second switch is re-loaded.
* June 7 - 17:05 CT - Network availability is restored.
## Remediation plan
We deeply apologize for the downtime, and are already working towards many improvements towards our infrastructure.
* We have replaced a suspected-faulty network switch.
* We are working internally to establish some better procedures for coordinating planned maintenance intervals in the future.
* We are investigating network switch options and fail-over configuration, to identify possible improvements we may be able to make to our network stack to avoid future service interruptions of this type.
* We are investigating possible UPS upgrades or replacements we may be able to make to provide extra power redundancy beyond what is already provided by the Flexential data center.
Some emails are stuck in a SENDING... state
Début 11 mars 2025 à 00:00 UTC · 0m
IssuesIncident mineur
resolved
We have identified an issue where some emails sent the evening of March 10 / the morning of March 11 appear to be stuck in a SENDING state, instead of being delivered.
We have identified and resolved an issue where an internal network attached storage device became disconnected, causing some emails to get stuck in a SENDING... state, instead of actually being delivered.
We are still investigating root cause further.
Delayed syncs to QuickBooks Online, Xero, Salesforce
Début 8 janvier 2025 à 12:52 UTC · 11h 35m
IssuesIncident mineur
Composants affectés
Integrations
identified
We have identified an issue causing delayed syncs of data to QuickBooks Online, Xero, and Salesforce for a small subset of accounts.
Data will still be synced to QuickBooks Online, Xero, and Salesforce automatically, but the sync to those systems may be delayed by several hours. Manual syncing of data will not be required.
We are working on a resolution.
monitoring
Data syncs to QuickBooks Online, Xero, and Salesforce are catching up quickly, but there is a significant number of items to sync before everything has fully caught up.
Data will still be synced to QuickBooks Online, Xero, and Salesforce automatically, but the sync to those systems may be delayed by several hours. Manual syncing of data will not be required.
We are continuing to monitor the situation.
monitoring
Data syncs to QuickBooks Online, Xero, and Salesforce are catching up quickly, and are almost fully caught up.
Data will still be synced to QuickBooks Online, Xero, and Salesforce automatically, but the sync to those systems may be delayed by a few more hours. Manual syncing of data will not be required.
We anticipate syncing being fully caught up within the next ~8 hours.
resolved
This issue has been resolved.
All data syncs to Xero, QuickBooks Online, and Salesforce have caught up and fully synced.
Investigating issues with ChargeOver
Début 18 novembre 2024 à 21:44 UTC · 23m
OutageIncident majeur
Composants affectés
Public WebsiteEmail SendingIntegrationsSearchPayment ProcessingMain ApplicationDeveloper Docs
investigating
We are aware of ongoing issues with access to our platform. We are currently investigating and will release details here as they become available.
identified
We've identified the problem and are implementing a fix.
resolved
This incident has been resolved.
A post-mortem containing further details to come.
postmortem
## Incident details
On November 18th, ChargeOver suffered a partial service outage.
Many customers were unable to access most parts of the ChargeOver application.
## Root cause
ChargeOver has a cluster of redundant ingress servers which accept all application traffic, and then route that traffic internally to different backend services.
This cluster of ingress servers \(HAProxy servers\) is backed by a distributed network file store \(GlusterFS\), which stores data \(encrypted SSL/TLS certificates\) needed by HAProxy to serve secure connections for incoming HTTP requests.
Sometime between November 7th and November 17th, the connection between two of our ingress load balancers and our distributed network file store was disconnected. We are still investigating the cause of the disconnection.
On November 18th, we deployed a routine update to our ingress load balancers. The deployment failed on 2 of the 3 redundant ingress load balancers due to the lost connection to the distributed network file store. This caused approximately 2/3 of traffic to ChargeOver to be dropped. The third ingress load balancer stayed active and continued to serve traffic.
Automatic DNS fail-over occurred, automatically removing the 2 affected ingress load balancers from serving traffic within 5 minutes of the outage.
ChargeOver staff were notified immediately. We restarted the two affected ingress load balancers to re-establish the distributed network file store connection, and re-deployed the ingress load balancer update.
This restored access to all services.
## Incident timeline
* Sometime between November 7th and November 17th - connection between load balancers and distributed network file store is lost
* November 18 - 3:29pm CT - Our team starts a routine deployment to our ingress load balancers
* November 18 - 3:33pm CT - The deployment completes, and we are immediately notified of a problem with 2 of the 3 load balancers
* November 18 - 3:38pm CT - Automatic DNS fail-over removes 2 of our 3 ingress load balancers from DNS, as expected
* November 18 - 3:44pm CT - We post an initial update to [https://status.chargeover.com/](https://status.chargeover.com/) of the partial outage
* November 18 - 3:55pm CT - After assessing the situation, we reboot the two problematic load balancers
* November 18 - 3:59pm CT - Re-deploy of the ingress update for load balancers succeeds across all 3 servers
* November 18 - 4:02pm CT - Our team confirms everything looks good, and DNS fail-over automatically adds the 2 failed servers back to DNS
* November 18 - 4:07pm CT - Update posted to [https://status.chargeover.com](https://status.chargeover.com)/ indicating the issue has been resolved
## Remediation plan
We’ve identified a number of take-aways from this incident.
* We are implementing additional monitoring to proactively monitor the connection between the ingress load balancers and the distributed network file store.
* We are discussing possible solutions to reduce/remove the dependency on the distributed file store from the load balancers.
* There were several internal DNS names which pointed at specific ingress load balancers, rather than using the round-robin, more fault-tolerant DNS which our public services use. This hampered our ability to recover from this situation more quickly, so we are moving those internal DNS names to use round-robin DNS.
Some customers are having trouble accessing ChargeOver
Début 12 novembre 2024 à 17:05 UTC · 1h 6m
Pending
Composants affectés
Main Application
investigating
We are receiving reports that some customers are having trouble accessing the ChargeOver app.
We are investigating further.
All ChargeOver services appear to be operating normally, but it looks like some web traffic may be having problems accessing ChargeOver services.
investigating
All ChargeOver services are operating normally.
Some customers are having trouble accessing ChargeOver. At this time, we believe this is due to a carrier network routing issue outside of ChargeOver.
We are still investigating and working with our data center providers to get more information.
monitoring
All ChargeOver services are operating normally, and are accessible to everyone now.
A networking or connectivity incident occurred with our data center provider (Flexential) which prevented some customers from accessing ChargeOver. We are still waiting for full details, and then will post an update.
We continue to monitor connectivity. We believe any impact is now resolved, and everyone should be able to access ChargeOver at this time.
An update will follow with more details.
resolved
All connectivity issues are resolved now.
Our primary data center (Flexential - Chaska) experienced a networking issue which impacted the ability to connect to ChargeOver services for some customers.
postmortem
## Incident details
On November 12th, some customers had trouble accessing ChargeOver.
This was a networking issue at the primary data center ChargeOver is hosted within.
## Root cause
On 12 November at 10:26 CT, ChargeOver began receiving automated reports of connectivity issues.
The issue was tracked down to a networking issue at the data center ChargeOver’s primary location is within \(Flexential in Chaska, MN\).
The problem was immediately escalated to Flexential engineers who identified an internet circuit at Flexential’s Chicago POP as the source of the issue.
Flexential opened a ticket with the circuit provider who confirmed there was an issue with packets over a certain size traversing the device due to an MTU misconfiguration. The circuit provider confirmed that they fixed the issue, however Flexential kept production traffic off the link while more investigation is performed.
Total impact to service was around 41 minutes.
## Incident timeline
* Nov 12 - 10:26am CT - ChargeOver’s automated uptime monitoring identifies and alerts our engineering team of an issue
* Nov 12 - 10:53am CT - ChargeOver identifies the issue as a bandwidth/networking issue outside of the ChargeOver network, and escalates the issue to Flexential engineering
* Nov 12 - 11:05am CT - Flexential resolves the issue
## Remediation plan
ChargeOver is working with our data center provider Flexential to improve monitoring and connectivity.
Delayed delivery of some emails
Début 28 juillet 2024 à 03:09 UTC · 2d 12h
IssuesIncident mineur
Composants affectés
Email Sending
investigating
Some customers are experiencing a delay in delivery of outgoing emails from ChargeOver. Our team is working with our email provider to investigate the delay.
At this time, we expect all emails to be sent successfully, but some delivery delays may occur.
identified
We have identified the problem, and are working with our email provider to resolve it.
At this time, we expect all emails to be sent successfully, but some delivery delays may occur.
identified
Emails of the following types are being delivered without any delays:
* invites sent to new admin users
* admin password resets
* scheduled reports
* any sort of admin notifications (e.g. notifications when a quote is accepted, when custom domains are configured, etc.)
We are aware that many of our customers are seeing significant delivery delays for transactional emails (e.g. invoice due emails, payment receipt emails, etc.).
We are working with our email provider to get the delays resolved as soon as possible.
If your account is using your own SMTP server or your own SendGrid, Mailgun, or Mandrill account, your account is unaffected by the delivery delays.
identified
We are diverting some email delivery to a secondary email account.
Some delivery delays are still occurring. While we are working to resolve this, some emails sent may not be sent with correctly aligned DKIM policies/signatures.
identified
We have diverted all email delivery to a secondary email account.
Some delivery delays are still occurring. While we are working to resolve this, some emails sent may not be sent with correctly aligned DKIM policies/signatures.
We continue to work with our email provider to resolve this.
monitoring
The email delays have been resolved. We are monitoring delivery of emails.
Some users may still experience some delays and/or higher-than-normal bounce rates for a short amount of time.
Further updates and a postmortem will follow.
resolved
Email delivery is operating normally.
A postmortem will follow.
postmortem
## Incident details
From July 25th 14:00 PST to July 29 20:41 PST, some emails sent via ChargeOver were delayed.
Only emails sent via ChargeOver’s email provider were affected \(emails sent via ChargeOver’s custom SMTP, custom SendGrid, Mailgun, and Mandrill integrations were _not_ impacted.\)
## Root cause
On two separate days in July, malicious users logged into two separate ChargeOver accounts, and used them to send a large number of spam/scam emails.
The ChargeOver accounts were _customers of ChargeOver_ - no
Both ChargeOver accounts used extremely easy to guess passwords, used those passwords across many other applications beyond ChargeOver, and did not have 2FA/MFA enabled within ChargeOver. Malicious actors were able to guess the ChargeOver users' passwords to log in to ChargeOver. _ChargeOver itself was not hacked and did not suffer any sort of security breach._
This led to ChargeOver’s primary email provider \(SendGrid\) temporarily placing a hold on some outgoing email from ChargeOver.
ChargeOver worked closely with SendGrid to restore normal email delivery.
## Incident timeline
* Early July - two malicious actors guess ChargeOver user passwords \(or re-use passwords from other applications\), and send many malicious emails via ChargeOver
* July 12 - 06:22 CT - ChargeOver’s monitoring automatically alerts our team of unexpected activity on two ChargeOver accounts, and we notify the affected customers and force password changes on the affected accounts
* July 13 - 08:25 CT - ChargeOver begins work towards future mitigation strategies
* \(July 13 - July 29\) - ChargeOver pushed out 19 separate updates through this 2-week period to mitigate impact, protect against future attacks, and migrate some critical pieces of infrastructure away from the affected SendGrid account\)
* July 25 - 16:00 CT - SendGrid places a suspension/hold on one of ChargeOver’s email accounts, but does not notify us
* July 26 - 11:55 CT - ChargeOver staff reach out to SendGrid because we are seeing some email being delayed
* July 29 - 22:41 CT - SendGrid removes the hold/suspension on the account, and normal email delivery returns
## Remediation plan
We’ve identified a number of items to be addressed to protect against future attacks, and mitigate impact.
**The most important item is to encourage all ChargeOver customers to** [**enable 2FA/MFA on their ChargeOver account**](https://help.chargeover.com/docs/getting-started/configuration/two-factor-authentication-2fa/)**.** Enabling 2FA/MFA is the quick, easy, free, and the single most important thing anyone can to do protect against any sort of unauthorized access to any application you use online. ChargeOver customers will see a push for 2FA/MFA adoption across all ChargeOver accounts.
Items already accomplished during the July 13 - July 29th period:
* Improved / more proactive monitoring of spam reports
* Added rate limiting for sending email through the ChargeOver admin panel
* Move critical email infrastructure to separate, independent SendGrid accounts
Future plans:
* Enforce rate limiting for email sending through the REST API
* Enforce rate limiting for email sending through the ChargeOver.js API
* Add various rate limits into a number of in-app endpoints
Delayed processing of search / sync to 3rd-party platforms / email sending
Début 4 décembre 2023 à 17:05 UTC · 7h 53m
IssuesIncident mineur
Composants affectés
Email SendingIntegrationsSearch
investigating
ChargeOver is experiencing delays when syncing data to 3rd-party platforms (Salesforce, QuickBooks, Xero, etc.) and some delays with search and email sending.
We are investigating the delays.
identified
We have identified the cause of the degraded performance, and are working towards a resolution.
Search and email sending is performing normally now. Sync to 3rd-party integrations is catching up, but still delayed.
monitoring
3rd-party integrations sync is continuing to catch up, and now impacting only a small percentage of users.
We are continuing to monitor the sync.
monitoring
We are continuing to monitor to ensure that all 3rd-party integration syncs catch up.
resolved
This issue has been resolved. A postmortem will follow.
We are aware of the problem, and are investigating.
We will post further updates as we have them.
identified
We have identified the problem.
ETA to resolution is less than 30 minutes.
identified
We are continuing to work on a fix for this issue.
monitoring
All systems have been restored.
We are monitoring the fix. A postmortem will be posted.
resolved
The incident has been resolved.
A postmortem will be provided.
postmortem
## Incident details
The ChargeOver team follows an agile software development lifecycle for rolling out new features and updates. We routinely roll out new features and updates to the platform multiple times per week.
Our typical continuous integration/continuous deployment roll-outs look something like this:
* Developers make changes
* Code review by other developers
* Automated security, linting, and security testing is performed
* Senior-level developer code review before deploying features to production
* Automated deploy of new updates to production environment
* If an error occurs, roll back the changes to a previously known-good configuration
ChargeOver uses `Docker` containers and a redundant set of `Docker Swarm` nodes for production deployments.
On `July 19th at 10:24am CT` we reviewed and deployed an update which, although passing all automated acceptance tests, errantly deployed an update which caused the deployment of `Docker` container images to silently fail. This caused the application to become unavailable, and a `503 Service Unavailable` error message was shown to all users. The deployment appeared to be successful to automated systems, but due to a syntax error actually only removed existing application servers rather than replacing them with the new software version. No automated roll back occurred, because the deployment appeared successful but had actually failed silently.
## Root cause
A single extra space \(a single errant spacebar press!\) was accidentally added to the very beginning of a `docker-compose` `YAML` file, which caused the `docker-compose` file to be invalid `YAML` syntax.
The single-space change was subtle enough to be missed when reviewing the code change. All automated tests passed, because the automated tests do not use the production `docker-compose` deployment configuration file.
When deploying the service to `Docker Swarm`, `Docker Swarm` interpreted the invalid syntax in the `YAML` file as an empty set of services to deploy, rather than a set of valid application services to be deployed. This caused the deployment to look successful \(it successfully deployed, removing all existing application servers, and replacing them with nothing\) and thus automated roll-back to a known-good set of services did not happen.
## Incident timeline
* 10:21am CT - Change was reviewed and merged from a staging branch, to our production branch.
* 10:24am CT - Change was deployed to production, immediately causing an outage.
* 10:29am CT - Our team posted a status update here, notifying affected customers.
* 10:36am CT - Our team identified the related errant change, and started to revert to a known-good set of services.
* 11:06am CT - All services became operational again after deploying to last known good configuration.
At this time, all services were restored and operational.
* 11:09am CT - Our team identified exactly what was wrong - an accidentally added single space character at the beginning of a configuration file, causing the file to be invalid `YAML` syntax.
* 11:56am CT - Our team made a change to validate the syntax of the `YAML` configuration files, to ensure this failure scenario cannot happen again.
## Remediation plan
There are several things that our team has identified as part of a remediation plan:
* We have already deployed multiple checks to ensure that invalid `YAML` syntax and/or configuration errors cannot pass automated tests, and thus cannot reach testing/UAT or production environments.
* Our team will work to improve the very generic `503 Service Unavailable` message that customers received, directing affected customers to [https://status.chargeover.com](https://status.chargeover.com) where they can see real-time updates regarding any system outages.
* Customers logging in via [https://app.chargeover.com](https://app.chargeover.com) received generic `The credentials that you provided are not correct.` messages, instead of a notification of the outage. This will be improved.
* Our team will do a review of our deployment pipelines, to see if we can identify any other similar potential failure points.
In-app pages are not loading correctly; trouble logging in
Début 19 mai 2023 à 01:40 UTC · 18m
OutageIncident majeur
Composants affectés
SearchMain Application
identified
We have identified an issue causing trouble logging in, and causing some pages to load incorrectly.
ETA to fix is less than an hour.
monitoring
A fix has been implemented, and we are monitoring to make sure all services have recovered.
resolved
This issue has been resolved.
Please contact us if you continue to have any trouble.