DevOps8 min read

Zero-Downtime Deployments with Terraform and GitOps: CI/CD at Scale

Zero-downtime deployments on AWS using Terraform, Docker, and GitOps CI/CD pipelines that eliminate deployment downtime and configuration drift.

DM

Deep Mehta

Founder & Cloud Engineer

Deploying software updates on a Friday afternoon should not be terrifying. Yet in companies without automated GitOps pipelines, releases require maintenance windows, user lockouts, manual database rollbacks, and frantic troubleshooting.

Real engineering velocity requires zero-downtime deployments where code moves from pull request to production automatically, reliably, and safely. Here is how 3 Dices Technology builds those pipelines.

Git as the single source of truth

GitOps extends Git version control to infrastructure and deployment configuration. Every change to infrastructure (Terraform) or container image versions is tracked via commits. If something breaks, rolling back is as simple as reverting the commit, which automatically triggers a declarative rollback in production.

Blue-green and canary strategies

To guarantee zero downtime, we never overwrite running containers in place:

  • Blue-green deployments spin up a complete new green environment alongside the active blue one. Health checks run against green, and once verified, traffic switches at the load balancer with zero dropped connections.
  • Canary deployments route 5% of production traffic to the new release, monitor error rates in Datadog or CloudWatch for ten minutes, and then progressively shift 25%, 50%, and 100%.

Database migrations without locks

The biggest obstacle to zero-downtime deployment is schema change. We mandate the expand/contract pattern:

  1. Expand: add new columns or tables as backward-compatible additions, and deploy code that reads from both old and new columns.
  2. Backfill: migrate historical data in background asynchronous batches.
  3. Contract: deploy code that points only to the new columns, then drop legacy tables cleanly without locking active transactions.

The pipeline stages

Our automated multi-stage workflow:

  1. PR validation: GitHub Actions, ESLint, and Jest run linting, unit tests, and coverage verification.
  2. Security and compliance: Trivy, SonarQube, and TFSec run container vulnerability scanning, SAST, and IaC checks.
  3. Build and artifacts: Docker and Amazon ECR produce multi-stage builds with immutable image tags keyed to the git SHA.
  4. Staging deploy: Terraform Cloud or ArgoCD deploys automatically to an isolated staging environment.
  5. Production rollout: AWS ALB and Route 53 weighted routing drive an automated canary or blue-green traffic shift with auto-rollback.

Ship with confidence

Ship faster with complete confidence. Our DevOps and CI/CD automation work designs and implements GitOps infrastructure end to end, built on sane Terraform state foundations.

#DevOps#Terraform#CI/CD
DM

About the author

Deep Mehta

Deep is the founder of 3 Dices Technology, a cloud engineering studio shipping AWS architecture, DevOps automation, and production AI systems for startups and SMBs.

Connect on LinkedIn

Frequently Asked Questions

How do you deploy database schema changes without downtime?
The expand/contract pattern: add backward-compatible columns first, backfill data asynchronously, then switch the code to the new columns and drop the old ones, all without locking active transactions.
Blue-green or canary?
Blue-green swaps all traffic to a verified new environment at once. Canary shifts a small percentage first and watches error rates before progressing. Both avoid overwriting running containers in place.
Why GitOps?
With Git as the single source of truth for infrastructure and image versions, rolling back is a git revert that triggers a declarative rollback in production.

Have a Question This Didn't Answer?

Ask us directly, we're happy to share what we know about your specific situation.