Partial outage for the SES Service Broker
- resolved
De 9h20 à 9h55 le mercredi 5 août 2026, le courtier en services Cloud a connu une panne partielle. Pendant ce temps, le courtier en services Cloud n'était pas disponible, ce qui signifie que les clients ne pouvaient pas créer de nouveaux services, mettre à jour les services existants ou supprimer les services existants. Toutefois, les services existants créés par le courtier ont continué de fonctionner normalement. À partir de 9h55, la panne a été résolue et le courtier en services Cloud était pleinement opérationnel. Afin d'identifier les causes de cet incident et de l'empêcher de se reproduire, l'équipe de Cloud.gov effectuera un incident post mortem et partagera nos conclusions dans les prochains jours. Merci d'être un client Cloud.gov.
- postmortem
# SES service broker outage on August 5, 2026 ## Summary From approximately 9:13 AM ET to 9:45 AM ET on Wednesday August 5, 2026,[ ](http://cloud.gov)the [SES service broker](https://docs.cloud.gov/platform/services/aws-ses/) was unavailable. During this time, customers could not create, update, or delete brokered SES services. The outage affected only the management of brokered SES services. Existing service instances and email delivery were not interrupted. ## Impact During the incident, customers were unable to: * Create new brokered SES services * Update existing brokered SES services * Delete existing brokered SES services No action was required from customers after service was restored. ## Cause An automated cleanup process deleted a container image that was still being used by the deployed SES service broker. When the broker restarted during a production deployment, it could not retrieve the image and failed to start. ## Timeline All times are Eastern Time. * August 4, 4:24 AM - An automated pipeline created a new container image for the SES service broker * August 4, 3:03 PM - An automated process deleted the container image used by the deployed broker * August 5, 7:17 AM - A [Cloud.gov](http://cloud.gov) production deployment began. * August 5, approximately 9:13 AM - The broker attempted to restart but could not retrieve its container image. * August 5, approximately 9:28 AM - The [Cloud.gov](http://cloud.gov) team received an alert and began investigating. * August 5, approximately 9:45 AM - Engineers manually redeployed the broker with an available container image and restored service. **Resolution** [Cloud.gov](http://cloud.gov) manually redeployed the SES service broker using an available container image. The team then reran the automated deployment pipeline and confirmed that the deployment and acceptance tests completed successfully. **Follow-up Actions** [Cloud.gov](http://cloud.gov) is taking the following actions to reduce the risk of a similar incident and improve recovery: * Improve monitoring and alerting for SES service broker outages. * Update the container image cleanup process so it does not delete images that are still in use. * Update the deployment process so the SES service broker references the intended current container image. Thank you for your patience while we resolved this issue. Questions may be sent to us at [[email protected]](mailto:[email protected]).
Traduit automatiquement depuis la mise à jour officielle de l'incident.