← blog · August 19, 2026

How to Do GitOps, and What It Actually Buys You

Why the pull model differs from push-based deployment, where self-healing turns into a fight, why environments belong in directories rather than branches, how to handle secrets and image tags, how rollback really behaves, and when GitOps is overkill for a small team.

If someone asks what is actually running in a cluster and the only way to answer is to read kubectl get output, your deployment system is not telling you the truth. A pipeline applied a manifest months ago, someone bumped a replica count by hand at midnight, someone else patched a ConfigMap. None of it is written down anywhere. The problem GitOps sets out to solve is not automating deployment, that was solved long ago. It is binding the state inside the cluster to a single auditable source.

What pull actually changes

In a conventional pipeline a CI job pushes toward the cluster with kubectl apply or helm upgrade. That job needs cluster credentials, which usually live in an environment variable on a shared runner. In the pull model a controller runs inside the cluster, watches the repository and applies changes itself.

The difference shows up in three places.

The direction of trust flips. With push, an external system holds broad rights over your cluster. With pull, the cluster grants nothing outward and only reaches out with a credential that can read a repository. For clusters whose API server sits on a private network, push means a VPN or a bastion; pull sidesteps that entirely.

The second difference is continuity. A push is an event: the job finishes and nothing looks again. A pull is a loop. The gap between what the repository says and what the cluster holds is remeasured on every reconciliation.

The third is scope. In a push world there is no complete list of what belongs in the cluster, only the residue of jobs that once ran. An object that should have been deleted survives for years unless somebody wrote the job that deletes it.

Drift and self-healing

Drift is the divergence between the manifest and the object in the cluster: a manual patch, a field an operator added, an old object nobody removed. Self-healing pulls it back.

On the Argo CD side it comes down to two flags:

spec:
  syncPolicy:
    automated:
      prune: true
      selfHeal: true

prune is off by default, for good reason: with it on, a file accidentally deleted from the repository deletes an object from the cluster. A second guard refuses to wipe everything when the repository turns up empty, and you disable it explicitly with allowEmpty: true. Flux takes the other route and makes .spec.prune a required field, so you cannot create the resource without answering the question. I prefer that: an explicit answer beats a quiet default.

Here is the trap. Turn on self-healing while a horizontal autoscaler changes replica counts and the two will fight: as long as the repository says replicas: 2, every scaling event gets reverted. Drop that field from the manifest, or tell the controller to ignore it. The same goes for fields that admission webhooks and operators graft onto objects after creation; that is usually why an application nobody touched keeps reporting itself out of sync.

Reconciliation is not infinitely frequent. Argo CD polls repositories under a three minute timeout by default (timeout.reconciliation in argocd-cm), while Flux carries the interval on each resource's own interval field. Both let you add a webhook and cut the delay to seconds, but do not turn polling off: when a webhook goes missing, the loop is your only remaining guarantee.

Argo CD and Flux grab the same problem from different ends

Argo CD is application-centric: a strong UI, one control plane for many clusters, its own authorization and single sign-on layer. If you want a dashboard that developers who are not fluent in Kubernetes can read, that alone can settle the choice.

Flux is cluster-native. Instead of a UI you get custom resources and a CLI, and Helm, Kustomize, OCI repositories and SOPS decryption are built in rather than bolted on.

My rule: pick Argo CD when humans will watch a dashboard and you want central, multi-tenant control; pick Flux when you want the platform itself declarative, everything from installation to secrets in the repository. Running both in one cluster is possible, but never let two reconcilers own the same object; nobody can predict which wins.

Repository layout: directories yes, branches no

Keep application code and manifests in separate repositories. Put them together and the automated commits that bump image tags trigger the code pipeline, an easy way to build a loop that feeds itself. People patch this with a skip marker in the commit message, but that is tape over a design problem.

Splitting environments by branch is the more common mistake. In order:

Promotion degenerates into cherry-picking. Moving a fix from staging to production means resolving conflicts tangled up with the genuine differences between environments. That is exactly how a staging-only replica count eventually leaks into production.

Branches drift apart. Six months in, the question "what differs between staging and production" stops being a diff and becomes archaeology.

And it is unnecessary. Kustomize overlays and Helm value files already exist to express environment differences; using branches for it puts version control in place of a configuration layer.

On a single branch the layout looks like this:

apps/
  payments/
    base/
    overlays/
      staging/
      production/
infra/
  base/
  overlays/

Promotion is now a one line diff: the image tag from the staging overlay moves to the production overlay, and the reviewer sees exactly what changed.

Secrets

Three reasonable options.

Sealed secrets, encrypted with the public key of a controller in the cluster and committed to the repository, carry the fewest dependencies. Only that cluster's private key opens the value, and the default scope binds the object's name together with its namespace, so you cannot copy the ciphertext into another namespace and decrypt it there. The cost: backing up that private key is now part of your disaster recovery plan, and losing it turns every secret in the repository into garbage.

SOPS encrypts at file level and keeps the key in a cloud key service or an age key. It is built into Flux, where .spec.decryption.provider: sops on a Kustomization is enough; Argo CD needs a plugin. Unchanged values produce unchanged ciphertext, so diffs stay readable, and that small detail matters a lot during review.

With an external vault and an operator that pulls values into the cluster, the repository holds only references. That is the cleanest separation, but the repository is no longer the single source of truth and your deployment gains a runtime dependency.

For small and mid-sized teams I recommend SOPS with age: an afternoon to set up, no vault to operate. If you already run a vault, the external operator is the right answer. Committing a plain Secret never is; base64 is an encoding, not encryption.

Who updates the image tag

A moving tag quietly breaks GitOps. When the tag stays fixed and the image behind it changes, the link between the commit and the binary in the cluster is severed and rollback loses its meaning. Tags should be versions or digests, and updating them should be automated.

Flux does this with three resources: ImageRepository scans the registry, ImagePolicy decides which tag wins, and ImageUpdateAutomation commits the choice back to the repository. A comment marks the field to update:

spec:
  template:
    spec:
      containers:
        - name: api
          image: registry.example.com/api:1.4.2 # {"$imagepolicy": "flux-system:api"}

On the Argo CD side a separate component does the same job through annotations. Watch the write-back method: choose the one that commits to the repository. Writing straight to the application object is faster, but then the repository no longer describes reality, and the day you re-apply it the tag snaps back.

A sensible split: tags update automatically in staging, promotion to production goes through a pull request. Automation does the boring part, a human makes the decision.

How rollback really works

Rollback means reverting a commit and waiting for reconciliation. On paper, that is all. In practice it trips over three things.

Database migrations do not roll back. Reverting a manifest does not revert a schema, so the old version is stuck reading the new one. The cure is writing migrations to be backward compatible: add the column first, ship the version that uses it second, drop the old one only once you are certain you will not go back.

Persistent volumes and stateful workloads do not roll back. Restoring a StatefulSet manifest does not restore what is on disk.

The rollback button in the UI moves the cluster to an older commit but leaves the repository alone, so with self-healing on the next reconciliation undoes it. That button buys time during an incident; it is not the fix. The real rollback happens in the repository.

When GitOps gets in the way

It is the middle of the night, you need to change one field, you have no appetite for a pull request and a review, and the controller reverts your manual edit within minutes. Write this scenario down before you live it.

Document the suspend path. flux suspend kustomization <name> stops reconciliation for a resource and flux resume restarts it; in Argo CD, temporarily disabling automated sync does the same. If these are not in the runbook, nobody recalls them at 3am.

Make suspension raise an alert. A resource left suspended is a silent bomb: weeks later nobody can work out why deployments stopped landing.

Shorten the process rather than removing it. A single-approval, auto-merging emergency pull request path beats bypassing review altogether.

Before you resume, write your manual change into the repository. If you do not, reconciliation erases it and you get to live the incident a second time.

When you do not need it

Three people, one cluster, a few deployments a day. In that picture GitOps hands you a controller, a repository layout, an encryption chain and a fresh failure surface. Most of what you get back is available far more cheaply: keep manifests in version control and apply them from the pipeline. Drift correction and audit trails start earning that price only once more than one pair of hands touches the cluster.

Adopt it when two of these three hold: more than one environment or cluster; more than three people touching the cluster; an obligation to show who changed what and when. Otherwise start with the cheap half. Putting manifests under version control is nearly free; the continuous reconciliation layer, in a small team, usually comes back as maintenance.