--- title: "Deployment Rollback Template" description: "Template for communicating deployment failures and rollback procedures during service disruptions." author: "openstatus" publishedAt: "2026-01-19" category: "template" howto: totalTime: "PT1H" steps: - name: "Detect and acknowledge the deployment issue" text: "Identify that a recent deployment is causing problems through monitoring alerts or user reports. Immediately acknowledge the issue and announce rollback decision to users." url: "#investigating" - name: "Execute the rollback" text: "Roll back to the previous stable version while providing status updates. Monitor error rates and system health during the rollback process." url: "#identified" - name: "Monitor post-rollback stability" text: "Verify that the rollback completed successfully and all services have returned to normal operation. Continue monitoring for any lingering issues." url: "#monitoring" - name: "Confirm resolution and communicate lessons" text: "Send final confirmation that services are restored. Schedule post-mortem to identify root cause and document preventative measures for future deployments." url: "#resolved" faq: - question: "How quickly should I decide to roll back a deployment?" answer: "If error rates exceed 5% or critical functionality is broken, initiate rollback immediately. Don't wait to debug in production. Roll back first, then investigate the issue in a safe environment. Time is critical - every minute of impact affects user trust." - question: "Should I complete the rollback before communicating, or communicate during?" answer: "Communicate the rollback decision immediately, then provide updates during the process. Users appreciate knowing you're taking action. Example: 'We've initiated a rollback and expect completion in 10 minutes' is better than waiting until it's done." - question: "What if the rollback doesn't fix the issue?" answer: "This suggests the problem isn't with the recent deployment. Pivot your communication to general incident response, investigate the root cause, and consider whether you need to roll back further or take other remediation steps. Update users with revised information." - question: "Do I need a post-mortem after every rollback?" answer: "Yes. Even quick rollbacks deserve analysis. Document what went wrong, why it wasn't caught in testing, and what process changes will prevent recurrence. Share key findings with users if appropriate to show continuous improvement." --- Use this template when a deployment causes issues and requires rolling back to a previous version. Communicates the issue, mitigation, and resolution clearly. ## When to Use This Template - Deployment causes service errors - New release breaking existing functionality - Configuration changes causing issues - Need to roll back to previous version ## Template Messages ### Investigating We encountered an issue during a scheduled deployment and have initiated a rollback to restore service stability. We apologize for the disruption and are working to complete the rollback as quickly as possible. ### Identified The deployment issue has been identified. We are completing the rollback to the previous stable version. ### Monitoring The rollback has been completed successfully. We are monitoring system stability to ensure all services are functioning normally. ### Resolved All services have been restored to normal operation. The deployment issue has been resolved and we will investigate the root cause. ## Real-World Examples ### Vercel: "Elevated error rates across new deployments" **Context**: Build system changes affecting deployments **Duration**: ~32 minutes **Impact**: New deployments and builds failing **Communication approach**: - Clear about what was affected ("new deployments") - Specific about the problem ("elevated error rates") - Quick resolution time communicated ### GitHub: "Configuration error during a model update" **Context**: Copilot model deployment issue **Duration**: ~1.5 hours **Impact**: 18% error rate on average, spiking to 100% **What worked**: - Transparent about the cause ("configuration error") - Quantified impact (error rates) - Showed progression: investigation → rollback → recovery monitoring - Committed to improvement **Sample update progression**: 1. "We're investigating reports of Copilot failures" 2. "Configuration error identified during model update. Rolling back." 3. "Rollback complete. Monitoring for full recovery." 4. "Fully resolved. Error rates back to normal. RCA in progress." ## Best Practices ### During Rollback **Do**: - ✅ Announce the rollback immediately - ✅ Set expectations on timeline if known - ✅ Acknowledge the user impact - ✅ Monitor closely during rollback **Don't**: - ❌ Wait until rollback is complete to communicate - ❌ Over-promise on fix timing - ❌ Skip the post-incident analysis ### After Resolution Always include: 1. Confirmation that rollback succeeded 2. Current system status 3. Next steps (investigation, new deployment plan) 4. Preventative measures (if ready to share) ## Communication Patterns ### For User-Facing Services ``` We've rolled back a deployment that was causing issues with [feature]. Service is now restored. We're investigating to prevent this in the future. ``` ### For API Services ``` A recent deployment caused elevated error rates in our API. We've rolled back to the previous version and error rates have returned to normal. All endpoints are functioning correctly. ``` ### For Internal Tools ``` Deployment to production caused issues with [system]. Rolled back to previous version. Services restored. Post-mortem scheduled for tomorrow to identify root cause and improve deployment process. ``` ## Preventative Context After resolution, consider adding context about prevention: > "We're improving our deployment process to catch issues like this > before they reach production. This includes enhanced pre-deployment > testing and gradual rollout procedures." ## Timing Guidance - **0-5 minutes**: Initial notification of issue and rollback decision - **5-20 minutes**: Rollback in progress updates - **20-30 minutes**: Rollback complete, monitoring - **30-60 minutes**: Confirmed resolution - **24 hours later**: Root cause analysis summary (optional) ## Related Templates - Use **API Service Disruption** if rollback affects external integrations - Use **Database Performance** if rollback impacts database operations - Use **Security Incident** if deployment exposed security issues