Voiceware: Operations & Lessons · deep-dive
Rollback in GitOps Is Simple Only When Releases Are Identifiable
Connecting Git reverts, ArgoCD reconciliation, and immutable artifacts into one recovery model.
The day-two operations question is connecting Git reverts, ArgoCD reconciliation, and immutable artifacts into one recovery model.
Day-two reality
I had GitOps reconciliation working while several services still used mutable latest tags, which made rollback behavior worth treating separately.
Getting the first successful deployment is only the beginning. Operations starts when the system has to survive repeated releases, dependency failures, load changes, partial outages, and engineers who were not present during the original build.
Signals
For this workload class I care about change diff, deployment status, product behavior, recovery time, repeatability from clean state. I want the signal to map to the actual failure mode, not just to whatever metric is easiest to collect.
change -> review -> deploy -> observe -> diagnose -> recover -> document
Failure scenarios
The primary risk is: Reverting YAML is not sufficient if the image reference does not uniquely identify the previous binary.
I also plan for partial failure. A controller can reconcile while the product is broken. A process can run while a dependency is slow. A queue can accept jobs while completion latency grows. An internal Service can resolve while no useful endpoint answers.
Operational response
I traced this as change -> review -> deploy -> observe -> diagnose -> recover -> document and checked change diff, deployment status, product behavior, recovery time, repeatability from clean state at each handoff. That kept the debugging path concrete.
I make the smallest change that addresses the layer with a concrete signal. If an emergency manual change is required, I treat it as temporary state and reconcile the intended fix back into Git afterward.
Safer defaults
The problem came down to this: Reverting YAML is not sufficient if the image reference does not uniquely identify the previous binary. What I carried forward was this: Git history and artifact history must line up for rollback to be deterministic.
Runbook notes
- Define rollback before the risky change.
- Keep secrets out of plaintext Git.
- Document what is implemented versus what remains a hardening next step.
- Change one layer at a time during migration.
- Review infrastructure by blast radius, not line count.
The operational principle is Git history and artifact history must line up for rollback to be deterministic. That lesson has been more reusable for me than any particular YAML pattern.