What Is GitOps and Should You Adopt It? How It Works, Benefits, Challenges, Tools & Best Practices

A hotfix goes out at two in the morning through a direct kubectl edit, the incident closes, and nobody records the change anywhere. Three weeks later a routine deployment overwrites it, the original failure returns, and the team spends hours working out why production no longer matches what the repository says it should be. GitOps exists to remove that class of problem. It treats a version-controlled description of the system as the only legitimate way to change it, and it uses software agents that keep checking whether reality still matches that description. This guide explains what GitOps means under the vendor-neutral OpenGitOps definition, how the workflow and reconciliation loop operate, what the model costs in complexity, which tools belong on a shortlist, and how to decide whether your organisation should adopt it now, pilot it, or wait.
TL;DR
GitOps is not "YAML in Git." OpenGitOps requires desired state that is declarative, versioned and immutable, pulled automatically by software agents, and continuously reconciled against the live system.
The pull-and-reconcile loop separates GitOps from a CI pipeline that merely pushes manifests. It lets an agent detect and correct drift, not only apply new commits.
The main benefits are auditability, reproducible environments, reviewable change and controlled recovery. Each depends on disciplined repositories, secrets handling and access control.
Adoption fits teams with Kubernetes, several environments or clusters, and strong Git review habits. Small, simple or heavily imperative estates may not repay the extra machinery.
Argo CD and Flux are the two best-known Kubernetes options, and both come from CNCF graduated projects. The right choice depends on interface needs, tenancy model and team skills, so run a short pilot before committing.
What is GitOps?
GitOps is an operating model in which a system's desired state is declared in version-controlled storage, usually Git, and software agents continuously pull that state and reconcile the live environment toward it. Changes happen through reviewed commits rather than manual commands, which makes deployments auditable, repeatable and self-correcting when drift occurs.
Table of Contents
What Is GitOps?
GitOps is an operating model for managing software systems in which the desired state is described declaratively, stored in a versioned and immutable store, and applied by software agents that pull that state and continuously reconcile the running system toward it. OpenGitOps, the community standard overseen by the GitOps Working Group, publishes the definition as four principles in version 1.0.0. It is most often used by platform, SRE and DevOps teams running containerised workloads.
Four parts make the model work. The desired state is the configuration sufficient to recreate the system, and the OpenGitOps glossary notes that it generally excludes persistent application data such as database contents. The state store keeps immutable versions of that state. Git is the canonical example, and the glossary allows any system that provides access control and auditing. The agent, also called a controller, runs in or beside the target environment, retrieves the desired state and compares it with the actual state. Reconciliation is the process of making the two match, triggered whenever they diverge, whether through a new commit or through unintended drift.
That is why "Git plus YAML" is not enough. A repository of manifests that a pipeline applies once and then forgets gives you version history but no continuous convergence. Without an agent that keeps observing the live system, nothing notices when someone edits a resource by hand, when a cluster is rebuilt, or when an applied change fails halfway. The loop, not the repository, is the distinguishing feature.
It also helps to separate three things that teams often blur. Application source code lives in source repositories. Deployment configuration, which says which version runs where and with which settings, is the desired state that GitOps manages. Built artifacts such as container images belong in an artifact registry and are referenced from configuration by tag or digest, not stored in Git.
The modern GitOps ecosystem centres on Kubernetes for a practical reason. Kubernetes describes workloads declaratively and relies on controllers that, in the project's own words, watch cluster state and try to move the current state closer to the desired state, as the Kubernetes documentation on controllers explains. GitOps tools extend that pattern outward to a repository. The principles themselves do not mention Kubernetes, so the model can apply to any system with a declarative interface that an agent can observe and correct. In practice, mature tooling outside Kubernetes is far thinner.
GitOps in one sentence: declare what should run, version and review that declaration, and let an agent continuously make reality match it.
How GitOps Works
A representative GitOps workflow has nine steps, shown here for a containerised application on Kubernetes.
A developer or operator proposes a change, such as a new application version or a resource limit, by editing declarative configuration on a branch.
The change is submitted as a pull request or merge request.
Reviewers and automated checks examine it: schema validation, linting, policy checks and any required approvals.
After approval, the change is merged into the branch or path that represents desired state for the target environment.
The GitOps controller detects the new revision, either by polling on an interval or after a webhook prompts an early fetch, and pulls it.
The controller compares the desired state with the live state reported by the target API.
It applies the changes needed to close the difference, in the order and with the safeguards the configuration defines.
It keeps observing after the apply and reports health and sync status.
If the live state later drifts, the controller reports or corrects the difference according to the configured policy.
Several components take part. The application repository holds source code. Continuous integration builds, tests, scans and packages that code and publishes an image to the artifact registry. The configuration repository holds the declarative description of what should run. The GitOps controller runs in or near the cluster and does the pulling and applying. The target API, for Kubernetes the API server, is where changes land. Observability tooling shows whether reconciliation is healthy.
CI and GitOps delivery complement each other. CI answers whether an artifact is good, while GitOps delivery answers whether an environment matches what was approved. The Argo CD documentation on automated sync notes that with automated sync a pipeline no longer needs direct access to the Argo CD API server to deploy. It commits to the Git repository instead.
A concise example shows the flow. A team merges a pull request in the application repository. CI builds the container image, scans it and pushes it to the registry with an immutable digest. An automated step, or a person, opens a pull request against the configuration repository that changes the image reference in the staging overlay. Reviewers approve it, the merge lands, the controller pulls the new revision and rolls out the Deployment, and a dashboard shows the application as synced and healthy. Promotion to production repeats the same pull request pattern against the production path, which keeps cloud deployment changes reviewable at every stage.
The Four Core GitOps Principles
OpenGitOps v1.0.0 defines four principles. Each is short, and each carries operational consequences that are easy to miss.
Declarative
OpenGitOps states that a system managed by GitOps must have its desired state expressed declaratively. A declarative description says what the system should look like, not the commands that produce it. That matters because a controller can compare two descriptions of state, which it cannot do with an ordered script. A Kubernetes Deployment manifest that specifies three replicas of a named image is a typical example. A common misunderstanding is that declarative means hand-written. Helm charts and Kustomize overlays are acceptable when the result is a declarative description with pinned versions. Some values may also be deliberately left out of Git: the Argo CD best-practices guide notes that if a Horizontal Pod Autoscaler owns the replica count, the manifest should not track it.
Versioned and Immutable
Desired state must be stored in a way that enforces immutability, versioning and a complete history. This is what makes audit, diff and revert possible. Git provides commit hashes, but immutability still depends on protections such as blocked force-pushes and reviewed merges. Templating can quietly break it. The Argo CD guidance warns that a Kustomize base pointing at an upstream repository's HEAD can change meaning without any commit in your own repository, and recommends pinning to a tag or commit. Flux's security guidance makes a similar point about remote bases.
Pulled Automatically
Software agents automatically pull the desired state from its source. The OpenGitOps glossary explains why pull is specified: agents must be able to reach the desired state at any time, not only when a change triggers a push event, which is a prerequisite for continuous reconciliation. Pull also keeps cluster credentials inside the cluster boundary, since the controller reaches out to the repository and CI does not need broad deployment rights. Webhooks do not turn pull into push. They only tell the agent to fetch sooner. Some products marketed as GitOps have a pipeline push manifests after a Git merge. That can be a useful pattern, but it falls outside the strict OpenGitOps definition.
Continuously Reconciled
Agents continuously observe actual system state and attempt to apply the desired state. The glossary clarifies that "continuous" means reconciliation keeps happening, not that it is instantaneous. The Flux core concepts documentation illustrates the behaviour: changes made to a managed cluster with kubectl are promptly reverted unless reconciliation is suspended or the change is pushed to the repository. The misunderstanding to avoid is that reconciliation guarantees success. Attempts can fail, wait on dependencies or take several cycles, and the controller reports that status rather than hiding it.
GitOps Architecture and Reconciliation
Architecturally, GitOps is a closed control loop. The OpenGitOps glossary describes it in control-theory terms: feedback shows how previous attempts to apply the desired state affected the actual state, and the agent responds by retrying, rolling back or alerting a human.
The loop observes, compares, acts and reports, then repeats:
Pull request -> review and checks -> merge
|
v
State store (Git repository or OCI registry)
^
| pull
|
GitOps controller --- compare ---> Live state (cluster API)
| ^
+------- apply and report ----------+Desired state is what the repository declares. Live state is what the cluster actually runs. Convergence means the two are equal. Drift is the glossary's term for a system whose actual state has moved, or is moving, away from the desired state.
Timing matters. Controllers check on intervals and sometimes on events. Argo CD's automated sync interval is governed by a reconciliation timeout that defaults to 120 seconds with up to 60 seconds of jitter, according to its documentation, and webhooks can shorten detection. Flux sources are checked on a defined interval and can also be prompted by receivers. The result is eventual consistency. Between a merge and convergence there is a window in which the repository and the cluster disagree, and operations teams should plan for it.
Sync state and health state are different questions. In Argo CD, an application whose live state deviates from Git is reported as OutOfSync, while health status analyses whether the resources are actually working. An application can be synced but unhealthy, such as a Deployment that matches Git while its pods crash. It can also be healthy but out of sync, such as a deployment someone scaled by hand.
Manual changes follow the tool's policy. Argo CD does not self-heal by default: according to its documentation, changes made to the live cluster do not trigger an automated sync unless self-heal is enabled, and automated sync does not prune deleted resources unless pruning is enabled. Flux Kustomizations revert out-of-band edits on the next reconciliation unless suspended.
The state store need not be Git. Flux documents a "Gitless GitOps" model in which controllers rely on container registries, with OCI artifacts holding the desired state, while Git remains the interface where people propose and review changes. This still satisfies the pull and reconcile principles, and it removes the Git server as a production dependency.
GitOps vs DevOps vs CI/CD vs Infrastructure as Code
These four ideas overlap, and most organisations use several at once. They answer different questions.
Concept | What it is | Source of truth | Execution model | Typical tools |
|---|---|---|---|---|
GitOps | Operating model for declarative, pulled, reconciled delivery | Versioned desired state in a state store | Agent pulls and reconciles continuously | Argo CD, Flux, Fleet |
DevOps | Culture and practices that join development and operations | Not applicable | Not applicable | Varies by team |
Continuous integration | Automated build, test and scan on each change | Source repository | Pipeline runs on triggers | Jenkins, GitHub Actions, GitLab CI |
Continuous delivery | Automated path to a releasable, deployable state | Pipeline and release definitions | Often pipeline-driven and push-based | Pipelines, release tools |
Infrastructure as Code | Provisioning infrastructure from declarative code | IaC files and state | Plan and apply when run | Terraform, Pulumi, CloudFormation |
DevOps is a way of working, not a mechanism, so GitOps is better described as one concrete way to practise parts of it. Teams that already follow DevOps principles often adopt GitOps to make deployment consistent and auditable.
Traditional CI/CD pipelines are usually triggered by events and push changes to the target environment. As the OpenGitOps glossary puts it, automation in traditional CI/CD is generally driven by pre-set triggers, whereas in GitOps reconciliation is triggered whenever there is a divergence. The two are complementary: CI builds and verifies artifacts, and GitOps delivery brings environments into line with approved configuration.
Infrastructure as Code describes and provisions infrastructure, often through a plan and apply cycle run on demand. A Git-triggered Terraform pipeline is not automatically GitOps. It is declarative and versioned, but nothing pulls and reconciles the environment between runs. Controller-based approaches to Terraform exist, and any of them should be evaluated for current maintenance and fit rather than assumed equivalent. Many teams use IaC to create clusters and cloud resources, and GitOps to manage what runs inside them. Configuration management tools sit nearby and may also coexist.
Benefits of GitOps
The benefits below are real but conditional. Each states what it depends on.
Auditable change history
Every change is a commit with an author, a timestamp and a review trail, which gives investigators a precise record of what changed and why. Git history is not automatic compliance, though. It supports audit trail needs only when branch protections, identity and review requirements are enforced.
Reproducible, consistent environments
When environments are generated from the same declarative sources, differences between staging and production become visible diffs rather than surprises. Reproducibility depends on pinned versions and on treating the repository as the only route to change.
Safer, reviewable change
Pull requests put automated validation and human review ahead of the production path. The benefit depends on the quality of the checks and the willingness of reviewers to read configuration diffs carefully.
Drift detection and self-healing
Because agents keep comparing, manual edits and failed applies are noticed. Whether they are corrected automatically is a policy choice, and automatic correction needs safeguards, covered later in this guide.
Controlled recovery
Configuration can be reverted by reverting a commit, and a replacement cluster can be rebuilt from the repository. This restores declarative configuration only. Persistent data, external services and schema changes need their own recovery plans.
Reduced direct production access
In a pull model, engineers change the repository rather than the cluster, and CI need not hold cluster credentials. That shrinks the set of people and systems with write access to production, provided the controller's own permissions are scoped carefully.
Collaboration and platform self-service
Teams already know pull requests. A platform team can publish standard patterns, and application teams can request environments or deployments through the same workflow, which supports cloud automation without handing out cluster administration rights.
Multi-cluster consistency
One repository structure can drive many clusters, so a policy change or a new add-on rolls out through the same reviewed path everywhere. This is where GitOps tends to repay its setup cost most clearly.
Challenges and Limitations of GitOps
The model moves complexity rather than deleting it. Teams should weigh the costs with the same seriousness as the benefits.
Operational and organisational hurdles
There is a learning curve: engineers must understand the controller's resources, sync semantics and failure modes. Process change is real too, since emergency fixes through direct cluster access must now follow a path through the repository, with a defined break-glass exception. Repository sprawl and configuration duplication grow quickly across many applications and environments. Migrating legacy deployment models is gradual work, and a big-bang cutover is rarely wise.
Technical hurdles
Debugging reconciliation requires reading controller status, events and logs, and a failure may sit in rendering, authentication, ordering or the target API. Dependency ordering is another difficulty, such as installing custom resource definitions before the resources that use them. Eventual consistency means the live system briefly lags the repository. Git or registry outages can delay changes, and controller availability becomes part of your delivery path. Conflicting automation, such as an autoscaler and a controller both managing replicas, produces fights that need explicit design. Imperative operations, stateful changes and database migrations do not fit a pure declarative loop and need separate handling.
Security hurdles
The repository, its credentials, the controller and the automation paths become high-value targets. A compromised controller with broad permissions can change everything it manages. Multi-tenant clusters add questions about who may reference whose sources and which identities each reconciliation uses. Observability demands grow as well, because someone must watch reconciliation health, not only application health.
Why GitOps can amplify weak platform design
GitOps automates whatever is in the repository. If the platform has unclear ownership, tangled environments, untested configuration or unmanaged secrets, a reconciling controller applies those flaws faster and more consistently. It is a delivery model, not a substitute for sound architecture, a point worth remembering when assessing cloud-native architecture decisions that sit upstream of the repository.
When GitOps Is a Good Fit—and When It Isn't
GitOps pays off most where change is frequent, declarative and spread across many targets. The table summarises the pattern.
Factor | Stronger fit | Weaker fit |
|---|---|---|
Platform | Kubernetes or other declarative, controller-driven systems | Mostly imperative systems with no declarative interface |
Scale | Multiple environments or clusters | One small environment with rare changes |
Change pattern | Frequent, repeatable configuration changes | One-off or manual infrastructure work |
Governance | Audit, approval and separation-of-duties needs | Little need for change evidence |
Team practice | Strong Git hygiene and pull request review | Inconsistent review or branching habits |
Drift | Drift is common or costly | Environments are short-lived or rebuilt often |
Operating model | Platform team offering shared services | No owner for tooling or platform |
Kubernetes estates with several environments are the clearest fit, especially where a platform team serves many application teams and where cloud-native infrastructure changes often. Audit and compliance pressure strengthens the case, because pull requests and commit history give a natural change record, which pairs well with cloud compliance programmes.
GitOps may be premature for very small, simple environments where a plain pipeline is easier to run and explain. It is harder to justify for workloads dominated by imperative external actions, such as long-running manual operations on legacy systems, and for teams that cannot yet maintain Git hygiene and review discipline. Not being ready is not a failure. It usually means a foundation, such as declarative configuration or a secrets strategy, should come first.
Should You Adopt GitOps? A Practical Decision Framework
Score each factor below from 0 to 2 and add the results. The scale is a conversation tool, not a benchmark, so adjust the weighting to your context.
Factor | 0 points | 1 point | 2 points |
|---|---|---|---|
Kubernetes adoption | None planned | Partial or planned | Primary platform |
Environments and clusters | One | Two or three | Many |
Change frequency | Rare | Weekly | Daily or more |
Compliance and audit needs | None | Informal | Formal evidence required |
Git and review maturity | Ad hoc | Some pull request practice | Enforced reviews and branch protection |
Declarative or IaC maturity | Mostly manual | Partly declared | Most configuration declared |
Platform capability | No owner | Part-time owner | Dedicated platform or SRE team |
Observed drift | Rare | Occasional | Frequent or costly |
Secrets strategy | Plaintext or none | Partial | Encryption or external store in place |
Debugging capability | Limited | Some | Comfortable with controller troubleshooting |
A total of 15 to 20 marks a strong candidate. Proceed, but still start with a bounded pilot. A total of 8 to 14 suggests piloting first: choose one low-risk application, one non-production environment and a time-boxed evaluation, and close the lowest-scoring gaps in parallel. A total of 0 to 7 suggests delaying adoption and investing in prerequisites such as declarative configuration, review habits and secrets handling. Two overrides apply regardless of total. A zero for Git and review maturity or for secrets strategy should pause production adoption until it is fixed, because GitOps makes both more consequential. The balanced conclusion for many organisations is a small pilot, honest measurement and a decision based on evidence from their own estate.
GitOps Tools Compared
The landscape was checked against current documentation for this guide. The table covers tools with a clear, documented role. It does not rank them, and it omits pricing, which changes frequently.
Tool | Approach | Interface | Best fit | Key trade-off |
|---|---|---|---|---|
Argo CD | Application-centric controller for Kubernetes | Web UI, CLI and API | Teams wanting central visibility and a developer-facing UI | One control plane to secure, scale and operate |
Flux | Toolkit of composable controllers | CLI and Kubernetes resources, with ecosystem UIs | Platform teams building standardised, API-driven workflows | More assembly and a steeper mental model |
Rancher Fleet | GitOps at scale across clusters, preinstalled in Rancher | Rancher UI | Rancher-managed fleets | Natural inside Rancher estates, less so outside |
GitLab with Flux | Flux integrated through the GitLab agent for Kubernetes | GitLab UI plus Flux | GitLab-centred organisations | Couples delivery to the GitLab platform |
Azure GitOps extension | Flux v2 as a managed cluster extension on AKS and Azure Arc | Azure portal, CLI, ARM | Azure-centred Kubernetes estates | Tied to Azure extension versions and support windows |
Fleet's documentation describes it as GitOps at scale that can be installed on any Kubernetes cluster with Helm and comes preinstalled in Rancher, as outlined in the Rancher Fleet overview. GitLab states that it recommends Flux for GitOps, and its Kubernetes integration documentation shows how the agent and Flux work together. Microsoft documents GitOps on AKS and Azure Arc-enabled Kubernetes as a Flux v2 cluster extension in its Azure GitOps documentation. Teams that prefer to buy managed Kubernetes can also compare Kubernetes as a Service options for how much GitOps tooling they bundle.
Argo CD vs Flux
These two dominate Kubernetes GitOps discussions, and neither is universally better. The right choice depends on how your teams work.
Dimension | Argo CD | Flux |
|---|---|---|
Philosophy | Application-centric, with a central view of apps | Toolkit of controllers, composable and API-first |
Architecture | Kubernetes controller plus API server, repo server and UI | Separate source, kustomize, helm, notification and image controllers |
Interface | Web UI, CLI and API | CLI and Kubernetes custom resources; UIs come from the ecosystem |
Configuration tools | Helm, Kustomize, Jsonnet, plain YAML, plugins | Kustomize and Helm controllers, with OCI and Git sources |
Multi-cluster | Manages many clusters from one instance | Per-cluster controllers, or a management cluster pattern |
Multi-tenancy | Projects, RBAC and SSO | Service account impersonation and cross-namespace lockdown |
Image automation | Commonly handled by CI or companion tools | Built-in image reflector and automation controllers that commit to Git |
Progressive delivery | Pairs with Argo Rollouts | Pairs with Flagger |
Learning curve | Gentler visual entry point | Steeper but closer to Kubernetes primitives |
Argo CD's design puts an application at the centre, with sync status, health and diffs visible in a UI, which suits organisations where developers want to see what is deployed. Its ApplicationSet feature generates many applications from templates, which helps with fleets. Flux's design splits responsibilities across controllers that you assemble, so platform teams can build repeatable, API-driven setups, and its security guidance covers impersonation and tenant isolation in detail. Flux also supports image update automation that writes changes back to Git.
Progressive delivery is complementary rather than built into either tool's sync logic. Flux's documentation points to Flagger for canary and related strategies. Argo Rollouts is described in Red Hat's OpenShift GitOps documentation as a controller with custom resources for blue-green, canary and experimentation.
A concise decision guide follows.
Prefer Argo CD when developers need a clear visual interface, when you want one place to see many applications, or when ApplicationSet-style generation suits your fleet.
Prefer Flux when you want composable controllers, strict tenant isolation by namespace and impersonation, OCI-based delivery, or tight integration with GitLab or Azure.
Either is reasonable when your needs are modest. Run a short pilot with representative applications and judge operability, not feature lists.
How to Adopt GitOps Step by Step
A phased approach limits risk and builds evidence.
Assess the current flow. Map how code, configuration and secrets reach each environment today, and list manual steps and known drift. You cannot improve what you cannot see.
Pick a low-risk pilot. Choose one stateless application and one non-production environment so that mistakes are cheap and reversible.
Make configuration declarative. Move manifests, Helm values or Kustomize overlays into version control, and remove hand-edited configuration from the pilot path.
Decide repository boundaries. Separate application source from deployment configuration, and choose a layout you can explain in a paragraph.
Set review and protection rules. Require pull requests, reviews and protected branches before anything reconciles automatically.
Choose a controller. Compare Argo CD, Flux or a platform-integrated option against your interface, tenancy and identity needs.
Define secrets handling. Decide how secrets reach the cluster without plaintext in Git before the first sensitive workload moves.
Add validation and policy. Run schema checks, linting and policy tests before merge so errors appear in review, not in production.
Bootstrap a non-production environment. Install the controller, point it at the pilot path and let it manage itself where the tool supports that.
Add observability and alerts. Track reconciliation failures, sync state and controller health from the first day.
Test drift and recovery. Make a manual change on purpose and watch detection and correction, then rebuild the environment from the repository.
Define promotion. Establish how a change moves from staging to production, by pull request, automation or both.
Write break-glass procedures. Document who may bypass the loop in an emergency, how the bypass is recorded and how the repository catches up.
Expand gradually and measure. Add applications and environments in waves, and compare results with your own baseline before moving production workloads.
GitOps Repository and Environment Design
Repository design shapes everything from permissions to review quality, and no layout suits every organisation.
The Argo CD best-practices guide recommends keeping Kubernetes manifests in a repository separate from application source. Its reasons include a cleaner audit log, separate access for developers and production deployers, and avoiding build loops when automation commits manifest changes. The Flux guide to repository structure describes four patterns: a monorepo, a repository per environment, a repository per team and a repository per application.
A monorepo is simple to browse and to promote across, but everyone with read access can see production configuration, and large pull requests make accidental production changes harder to spot. A repository per environment limits who can see and change production, at the cost of more effort to promote infrastructure changes. A repository per team suits a platform model, where the platform team owns clusters and add-ons and application teams own their own definitions. Whichever you choose, consider code-owner files so the right reviewers approve the right paths.
Environments are usually represented as directories with a shared base and per-environment overlays, using Kustomize or Helm values. Branch-per-environment layouts can work, but they make promotion a merge exercise that tends to accumulate conflicts and silent differences. Machine-generated updates, such as image tag bumps, should arrive as reviewable commits or pull requests, particularly for production.
An illustrative layout, adapted from the pattern in Flux's documentation, looks like this:
platform-config/
├── clusters/
│ ├── staging/
│ └── production/
├── infrastructure/
│ ├── base/
│ ├── staging/
│ └── production/
└── apps/
├── base/
├── staging/
└── production/Separating infrastructure from apps allows ordering, for example installing cluster add-ons before applications. As the estate grows, revisit the layout deliberately instead of letting it sprawl.
GitOps Security, Secrets, and Governance
GitOps does not make a system secure simply because everything is in Git. It concentrates control in a few places, so those places must be hardened.
Treat the repository and controller as critical infrastructure
Protect production branches, require review, and give contributors the least access they need. Scope credentials narrowly and rotate them: GitLab's best-practice documentation advises regularly rotating the keys Flux uses to access manifests and the agent registration token. Use signed commits where your risk profile justifies them. Argo CD documents Git GnuPG verification for source integrity.
Controller permissions deserve particular care. In shared clusters, Flux's security guidance recommends disabling cross-namespace references, setting a default service account so reconciliations impersonate a limited identity, and using network policies so only Flux components reach source artifacts. Kubernetes RBAC and service accounts define what each reconciliation may change. These controls matter for cloud-native security because a compromised controller can alter everything it manages.
Secrets
Never store unencrypted production secrets in Git. Argo CD's documentation distinguishes two styles. It strongly recommends populating secrets on the destination cluster with tools such as Sealed Secrets, External Secrets Operator, the Kubernetes Secrets Store CSI Driver or Vault Secrets Operator, so Argo CD never needs to read them. It cautions against injecting secrets during manifest generation, partly because generated manifests are cached in plaintext in its Redis instance. Flux supports decrypting SOPS-encrypted secrets at reconciliation time and documents Sealed Secrets, and its guidance warns against keeping plaintext credentials in sources. Encrypted secrets are not risk-free. Keys, scopes and access to decryption identities remain attack surfaces, which is why cloud security reviews should include them.
Policy, provenance and separation of duties
Policy as code enforces rules before and at admission. Admission controllers and policy engines such as OPA Gatekeeper and Kyverno can block privileged containers or unsigned images, and Flux's guidance suggests checking that such controls prevent tenants from creating privileged workloads. Supply-chain integrity matters as well: GitLab advises signing generated OCI images and deploying only images that Flux verifies. Keep authors and approvers separate, retain audit logs outside the repository too, and make break-glass use visible. Git history supports governance, but it is not perfect evidence on its own, which is why cloud governance controls and the cybersecurity infrastructure around the repository still apply.
Multi-Cluster GitOps and Platform Engineering
Multi-cluster operation is where GitOps often proves its value. One repository structure can describe a fleet: bootstrap a new cluster from a known baseline, apply the same add-ons and policies everywhere, and vary only what must differ. Argo CD manages many clusters from one instance and generates applications with ApplicationSets. Flux's documentation describes tenant-dedicated clusters managed from a management cluster, and Fleet targets large numbers of clusters. This supports cloud orchestration at fleet scale and works across multi-cloud and hybrid cloud estates, including edge locations where central control is hard.
Platform engineering teams build internal platforms so developers can self-serve environments and deployments. GitOps is a natural delivery mechanism for such platforms, because onboarding an application can mean adding a reviewed entry to a repository. The two are not synonyms, however. Platform engineering covers developer experience, abstractions and support, while GitOps is one way to deliver changes. A thoughtful cloud operating model clarifies which team owns which layer.
Scale introduces risks. Blast radius grows when one change reaches every cluster, so stage rollouts by ring. A single repository can become a bottleneck for reviews and merges. Permissions grow complicated across tenants, and ownership becomes ambiguous without code owners. Reconciliation load rises with the number of objects, and Flux's guidance notes that tenants sharing controllers compete for the same reconciliation queue, so mission-critical lanes may warrant separate instances. Shared controllers also form a security boundary that needs deliberate isolation.
Drift, Rollbacks, Disaster Recovery, and Observability
Drift is any difference between desired and live state. Detection is the baseline capability. Correction is a policy decision. Automatic remediation suits stable, well-understood workloads, while manual approval may suit sensitive production environments where an unexpected overwrite could cause harm, such as when an operator is mid-incident. Self-healing needs safeguards: pruning off or reviewed, protected resources for stateful systems, and clear signals when the controller reverts something.
Reverting a desired-state commit restores prior declarative configuration, and the controller will reconcile toward it. Rollback is not magic. Application and database compatibility still matter, persistent data is separate, and destructive external operations may be irreversible. Argo CD's documentation also notes that rollback cannot be performed against an application with automated sync enabled, so teams should revert in Git instead. Rolling forward with a fix is sometimes safer than reverting, especially after a schema migration. Immutable artifacts help, since a pinned image digest means a rollback reproduces what actually ran.
Recovery of a cluster from Git-managed state is a major benefit. A replacement cluster can be bootstrapped and reconciled from the repository, which is why disaster recovery tests should rebuild an environment end to end. Git is not a backup system. It does not restore databases, volumes, external resources or secrets that live in other stores, so backups and disaster recovery planning remain essential, along with protections against ransomware affecting repositories and backups.
Observability should cover the controller as well as the application. Watch controller logs, reconciliation events, sync and health status and metrics. Argo CD exposes Prometheus metrics, and Flux provides alerts, events, logs and metrics. Alert on failed or stalled reconciliation, long-running drift and controller errors, and feed signals into your cloud telemetry and wider cloud operations practice, including incident tooling such as SIEM where relevant. Common failure modes include authentication to the repository expiring, rendering errors from template changes, dependency ordering problems and controllers starved of resources.
GitOps Best Practices
These practices reflect the guidance cited throughout this article.
Keep desired state declarative and pinned. Reproducibility disappears when a base or chart can change meaning without a commit.
Protect production branches and require reviewed pull or merge requests. Review is the main control GitOps adds.
Separate CI from reconciliation. CI builds and verifies artifacts, and the controller applies approved configuration.
Use immutable artifacts for sensitive environments. Prefer digests or versioned tags over mutable ones.
Keep secrets out of plaintext Git. Use encryption or destination-cluster secret management.
Apply least privilege everywhere. Scope repository credentials, controller identities and tenant permissions.
Keep repositories understandable. A new engineer should be able to find where a change belongs.
Validate before merge. Schema checks, linting and policy tests catch errors cheaply.
Add policy gates at admission. Enforce rules for privileged workloads, registries and signatures.
Make ownership explicit with code owners, so changes reach the people who understand them.
Monitor reconciliation health, not only application health.
Test disaster recovery and drift correction on a schedule.
Maintain a documented break-glass process, and require follow-up commits so Git catches up.
Adopt gradually, minimise manual changes to managed resources, and document exceptions.
Be cautious with automatic updates in sensitive production environments, and measure outcomes against a baseline.
Common GitOps Mistakes and Anti-Patterns
Calling any deployment from Git "GitOps." Without pull and continuous reconciliation, you have version-controlled deployment, which is valuable but different.
Letting CI push directly while assuming a pull model, which keeps broad cluster credentials in the pipeline.
Editing managed production resources by hand, which creates drift the controller will revert or report.
Storing plaintext secrets in Git, where history keeps them even after removal.
Granting the controller excessive privileges, so a compromise or a bad commit has maximum impact.
Overcomplicated repository layouts that nobody can reason about.
Using a branch for every environment without understanding promotion and merge-conflict costs.
Weak review and approval controls, which turn the repository into an unguarded path to production.
Mutable image tags in sensitive environments, which break reproducibility and rollback.
Enabling automatic reconciliation and pruning without safeguards for stateful or critical resources.
Ignoring observability, so failed reconciliation goes unnoticed.
Having no break-glass process, which pushes engineers to bypass the system in an emergency.
Never testing disaster recovery, so rebuild assumptions fail when needed.
Migrating everything at once instead of in waves.
How to Measure Whether GitOps Is Working
Measure against your own baseline before adoption, and avoid expecting predetermined percentage gains. Useful indicators include the following.
Deployment lead time and frequency, where faster, smaller changes are a goal for your team.
Failed change rate and mean time to restore, which show whether changes are safer and recovery quicker.
Reconciliation failure rate and time to reconcile, which reveal controller and configuration health.
Drift frequency and time to detect drift, which show how often reality departs from the repository and how quickly you notice.
Manual production changes, which should fall as the process takes hold.
Configuration-related incidents and rollback or roll-forward time.
Policy violations caught before merge, a sign that validation is working.
Cluster bootstrap and recovery time, tested by rebuilding.
Interpret trends over several weeks and pair numbers with qualitative feedback from the engineers who operate the system.
Frequently Asked Questions About GitOps
What is GitOps in simple terms?
GitOps means you describe what your systems should look like in files stored in version control, review changes the way you review code, and let software continuously make the running environment match those files. Instead of engineers running commands against production, an agent pulls approved configuration and corrects differences. It is defined formally by the four OpenGitOps principles.
Is GitOps the same as DevOps or CI/CD?
No. DevOps is a culture and set of practices, CI/CD automates building, testing and delivering software, and GitOps is a specific operating model for declarative, pulled and reconciled delivery. They overlap and combine well. CI builds and verifies artifacts, while GitOps delivery makes environments match the configuration you approved.
What is the difference between GitOps and Infrastructure as Code?
Infrastructure as Code describes and provisions infrastructure declaratively, often through a plan and apply run. GitOps adds versioned, pulled state and an agent that continuously reconciles. A Git-triggered IaC pipeline is not automatically GitOps because nothing keeps checking between runs. Many teams use IaC for cloud resources and GitOps for what runs on them.
Does GitOps require Kubernetes?
The principles do not require Kubernetes, but the mainstream tooling targets it because Kubernetes is declarative and controller-driven. Outside Kubernetes, you need a declarative interface that an agent can observe and correct, and mature options are fewer. If you do not run Kubernetes, evaluate whether the model fits before assuming it does.
Why does GitOps use a pull model?
According to the OpenGitOps glossary, agents must be able to access desired state at any time, not only when a push event occurs, which is what makes continuous reconciliation possible. Pull also keeps deployment credentials inside the environment. Some products labelled GitOps push changes from a pipeline, which falls outside the strict definition.
Argo CD or Flux: which should you choose?
There is no universal winner. Argo CD offers an application-centric view with a web UI, which many teams value for visibility. Flux offers composable controllers, built-in image automation and strong tenant isolation features. Choose based on interface needs, tenancy model, integrations and team skills, and pilot with representative applications before committing.
Can GitOps manage secrets?
Yes, but not as plaintext in Git. Common patterns include encrypting secrets in the repository with tools such as SOPS or Sealed Secrets, or syncing them from an external store with External Secrets Operator or Vault integrations. Argo CD's documentation strongly favours populating secrets on the destination cluster. Encryption keys and access still need careful control.
How does GitOps handle rollbacks?
Reverting a commit restores earlier declarative configuration, and the controller reconciles toward it. That does not undo data migrations, persistent data changes or destructive external actions, and schema compatibility still matters. Sometimes rolling forward with a fix is safer. Immutable artifacts and tested backups make either path more reliable.
What happens if someone manually changes the cluster?
The agent detects the difference as drift. Depending on policy, it reports it, as Argo CD does by default with an OutOfSync status, or reverts it, as Flux does on the next reconciliation unless suspended. Teams usually define a break-glass procedure for real emergencies and require follow-up commits so Git matches reality.
Can GitOps work with Terraform?
Yes, in several ways. Teams often use Terraform to provision clusters and cloud resources, then use GitOps for workloads. Running Terraform from a Git-triggered pipeline is Git-centred automation but not strict GitOps. Controller-based approaches that reconcile Terraform exist, and they should be assessed for current maintenance and operational fit.
Does GitOps replace CI?
No. CI still builds, tests, scans and packages software and publishes artifacts. GitOps replaces or reduces the push-style deployment step by having an agent pull approved configuration. The two complement each other, and a pipeline commonly ends by proposing a configuration change that the GitOps workflow then reviews and applies.
When should a company not use GitOps?
Delay it when environments are tiny and simple, when workloads are mostly imperative, when configuration is not yet declarative, or when review and secrets practices are weak. A short pipeline may serve better. Revisit the question as environments, clusters and compliance needs grow, using the scorecard in this guide.
Key Takeaways
GitOps is defined by four OpenGitOps principles, and a repository of manifests alone does not meet them.
The continuous reconciliation loop, not Git itself, provides drift detection and correction.
Container images belong in artifact registries, while Git or an OCI store holds deployment configuration.
Benefits such as auditability and recovery are conditional on branch protection, review quality, secrets handling and tested backups.
Rollback restores configuration only, so data, external side effects and schema changes need separate plans.
GitOps can amplify a poorly designed platform, so fix ownership, repository structure and secrets before scaling it.
Argo CD and Flux are both mature options, and the decision should rest on a pilot rather than feature lists.
Security shifts to the repository, credentials and controller, which need least privilege, policy and audit.
Measure results against your own baseline instead of expecting fixed percentage gains.
Actionable Next Steps
Map your current deployment flow, including who can change production and how secrets reach each environment.
Score your organisation with the decision framework and note the two lowest factors.
Check declarative readiness: list which workloads already have versioned manifests, charts or overlays.
Choose a low-risk pilot application and one non-production environment.
Compare Argo CD and Flux, or your platform's integrated option, against interface, tenancy and identity requirements.
Design secrets handling and decide who owns encryption keys or external stores.
Define repository conventions, branch protections, review rules and code owners.
Run a pilot that includes drift tests, a rollback, a roll-forward and a full environment rebuild.
Measure lead time, failed changes, drift and recovery time, then decide whether to expand.
Glossary
GitOps: An operating model in which declarative desired state is versioned, pulled by agents and continuously reconciled.
Desired state: The configuration sufficient to recreate a system, generally excluding persistent application data.
Actual (live) state: What is currently running in the target environment.
Reconciliation: The process of making actual state match desired state.
Drift: A system's actual state moving or having moved away from the desired state.
Declarative configuration: A description of the target state without the steps to reach it.
Controller: A software agent that observes state and works to bring it toward a target.
Source of truth: The authoritative store from which desired state is read.
Continuous integration (CI): Automated building, testing and scanning of each change.
Continuous delivery: Practices that keep software in a releasable state with automated delivery.
Continuous deployment: Automatically deploying changes to production once they pass automated checks.
Infrastructure as Code (IaC): Defining and provisioning infrastructure through code.
Pull request: A proposed change submitted for review before merging; GitLab calls it a merge request.
Helm: A package manager for Kubernetes that renders templated charts.
Kustomize: A tool for customising Kubernetes manifests through bases and overlays.
Argo CD: A declarative GitOps continuous delivery tool for Kubernetes.
Flux: A set of GitOps controllers, the GitOps Toolkit, for Kubernetes.
Policy as code: Rules expressed in code and enforced automatically.
Progressive delivery: Gradual rollout techniques such as canary releases, often handled by tools like Argo Rollouts or Flagger.
RBAC: Role-based access control, which governs what identities may do.
Multi-tenancy: Serving multiple teams or customers on shared infrastructure with isolation.
Immutable artifact: A build output, such as an image digest, that does not change once published.
State store: A system that stores immutable versions of desired state, such as Git or an OCI registry.
Sources & References
GitOps Principles (v1.0.0) and OpenGitOps home page. OpenGitOps, overseen by the GitOps Working Group. Accessed October 12, 2026. https://opengitops.dev/
GitOps Glossary. OpenGitOps. Accessed October 12, 2026. https://github.com/open-gitops/documents/blob/main/GLOSSARY.md
Argo CD Overview. Argo Project. Accessed October 12, 2026. https://argo-cd.readthedocs.io/en/stable/
Automated Sync Policy. Argo CD documentation. Accessed October 12, 2026. https://argo-cd.readthedocs.io/en/stable/user-guide/auto_sync/
Secret Management. Argo CD documentation. Accessed October 12, 2026. https://argo-cd.readthedocs.io/en/stable/operator-manual/secret-management/
Best Practices. Argo CD documentation. Accessed October 12, 2026. https://argo-cd.readthedocs.io/en/stable/user-guide/best_practices/
Core Concepts. Flux documentation. Last modified April 20, 2026. https://fluxcd.io/flux/concepts/
Security Best Practices. Flux documentation. Last modified July 9, 2026. https://fluxcd.io/flux/security/best-practices/
Ways of structuring your repositories. Flux documentation. Last modified May 13, 2024. https://fluxcd.io/flux/guides/repository-structure/
Controllers. Kubernetes documentation, The Kubernetes Authors. Accessed October 12, 2026. https://kubernetes.io/docs/concepts/architecture/controller/
Argo graduates from CNCF Incubator following massive community growth. ITOps Times, republished by CNCF. December 6, 2022. https://www.cncf.io/news/2022/12/06/itops-times-argo-graduates-from-cncf-incubator-following-massive-community-growth/
GitOps hits stride as CNCF graduates Flux CD and Argo CD. TechTarget. Accessed October 12, 2026. https://www.techtarget.com/it-infrastructure/news/252528152/GitOps-hits-stride-as-CNCF-graduates-Flux-CD-and-Argo-CD
Best practices for using the GitLab integration with Kubernetes. GitLab Docs. Accessed October 12, 2026. https://docs.gitlab.com/user/clusters/agent/enterprise_considerations/
Connecting a Kubernetes cluster with GitLab. GitLab Docs. Accessed October 12, 2026. https://docs.gitlab.com/ee/user/clusters/agent/
GitOps with GitLab: What you need to know about the Flux CD integration. GitLab. February 8, 2023. https://about.gitlab.com/blog/why-did-we-choose-to-integrate-fluxcd-with-gitlab/
Fleet overview. SUSE Rancher documentation. Accessed October 12, 2026. https://ranchermanager.docs.rancher.com/integrations-in-rancher/fleet/overview
Application deployments with GitOps (Flux v2) for AKS and Azure Arc-enabled Kubernetes. Microsoft Learn. Accessed October 12, 2026. https://learn.microsoft.com/en-us/azure/azure-arc/kubernetes/conceptual-gitops-flux2
Argo Rollouts overview. Red Hat OpenShift GitOps documentation. Accessed October 12, 2026. https://docs.redhat.com/en/documentation/red_hat_openshift_gitops/1.20/html/argo_rollouts/argo-rollouts-overview


