Enforcing Policy in Kubernetes: OPA Gatekeeper vs. Kyverno, and When to Use Which
RBAC decides who can act, not whether the object they submit is actually valid. Admission-based policy engines close that gap. We compare OPA Gatekeeper, Kyverno, and Kubernetes' own ValidatingAdmissionPolicy, and walk through the operational traps each one carries.
The boundary RBAC never draws
A well-configured RBAC setup tells you exactly who can create a pod in which namespace, and who can read a Secret. What it never asks is whether the object being submitted is actually reasonable: does it pull an image tagged latest, does it run as privileged: true, does it set resource limits, does it carry the labels your platform requires. Those are validation questions, not authorization questions, and RBAC has no opinion on them. If a developer has permission to create pods, nothing stops the content of that pod, from image provenance to privilege level, unless something else is watching.
That gap is closed at the admission layer. Before the Kubernetes API server persists an object, the request passes through authentication, authorization, mutating admission webhooks, and validating admission webhooks, in that order; only after the full chain succeeds does the object reach etcd. Three approaches have matured around this layer: OPA Gatekeeper, Kyverno, and Kubernetes' own native ValidatingAdmissionPolicy. All three solve the same problem, at different costs.
What each one actually looks like
OPA Gatekeeper is the Kubernetes-specific distribution of Open Policy Agent. Policies are written in Rego and split into two pieces: a ConstraintTemplate defines the logic, and a Constraint applies that template to specific resources with specific parameters.
apiVersion: templates.gatekeeper.sh/v1
kind: ConstraintTemplate
metadata:
name: k8srequiredlabels
spec:
crd:
spec:
names:
kind: K8sRequiredLabels
validation:
openAPIV3Schema:
type: object
properties:
labels:
type: array
items: { type: string }
targets:
- target: admission.k8s.gatekeeper.sh
rego: |
package k8srequiredlabels
violation[{"msg": msg}] {
required := input.parameters.labels
provided := input.review.object.metadata.labels
missing := required[_]
not provided[missing]
msg := sprintf("missing label: %v", [missing])
}
---
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sRequiredLabels
metadata:
name: require-team-label
spec:
match:
kinds: [{ apiGroups: [""], kinds: ["Namespace"] }]
parameters:
labels: ["team"]
Rego is a general-purpose query language, and that power pays off in scenarios where you need to validate one resource against another, for instance checking that a Service an Ingress refers to actually exists. The cost is the learning curve: if nobody on the team has written Rego before, the first few policies take longer than expected.
Kyverno solves the same problem using the YAML conventions Kubernetes users already know. Instead of a separate language, policies are written as pattern matching, or, in newer versions, as CEL expressions.
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: require-team-label
spec:
validationFailureAction: Enforce
rules:
- name: check-team-label
match:
any:
- resources: { kinds: ["Namespace"] }
validate:
message: "team label is required"
pattern:
metadata:
labels:
team: "?*"
Kyverno's real differentiator isn't just validate, it's that the same engine also supports mutate and generate rules: it can backfill a missing field, or automatically create a related NetworkPolicy whenever a new Namespace is created. Gatekeeper has no generate capability; it validates and, optionally, mutates, but it does not create new objects.
The third path skips installing a separate component altogether. Since becoming GA in Kubernetes 1.30, ValidatingAdmissionPolicy runs CEL-based validation logic directly inside the API server, with no external webhook pod required.
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicy
metadata:
name: require-team-label
spec:
failurePolicy: Fail
matchConstraints:
resourceRules:
- apiGroups: [""]
apiVersions: ["v1"]
resources: ["namespaces"]
validations:
- expression: "has(object.metadata.labels.team)"
message: "team label is required"
---
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicyBinding
metadata:
name: require-team-label-binding
spec:
policyName: require-team-label
validationActions: ["Deny"]
The win here is operational: there's no external webhook service that can be unavailable, add network latency, or need its own Deployment kept up to date. The cost is scope: no mutation, no generation, validation only, and CEL is less expressive than Rego for cross-resource logic.
The traps
All three approaches share the same core risk: the failurePolicy field. Setting it to Fail means that if the policy engine's webhook is temporarily unreachable, say the pod is restarting or the node is draining, every request touching that resource type is blocked; cluster-wide deploys can stall. Setting it to Ignore means that when the engine crashes, policies silently stop being enforced and nobody notices. The practical fix is to exclude the policy engine's own namespace via namespaceSelector, so the engine can never deadlock on its own webhook during an update, and to only turn on Fail once the engine runs with real high availability, at least two replicas plus a PodDisruptionBudget.
The second trap is shipping a new policy straight into enforce mode. Both Gatekeeper and Kyverno offer an audit or dry-run mode: the policy reports which existing resources violate it without rejecting anything. Skipping this step and flipping a new policy straight to enforce/deny, without first running it in audit mode for a few days to clean up existing violations, routinely produces unexpected deploy blocks in production. ValidatingAdmissionPolicy mimics the same gradual rollout by setting validationActions to Warn or Audit before switching to Deny.
The third trap is performance. Every admission request flows through every webhook in the chain synchronously. A Gatekeeper installation running many complex Rego rules can add measurable API server latency during busy deploy windows. As the number of policies grows, keeping match blocks as narrow as possible, targeting the specific kind and namespace instead of matching everything, keeps that latency under control.
Testing before you deploy
Applying a policy straight to the cluster and watching what breaks in production is a bad habit no matter which engine you pick. Gatekeeper's gator CLI runs sample Kubernetes objects against a ConstraintTemplate and Constraint pair without touching a cluster at all, and shows locally which object would be denied and why. Kyverno offers the equivalent with kyverno apply <policy.yaml> --resource <test-object.yaml>, or the CI-friendly kyverno test command, which lets a policy repository keep test files describing expected pass/fail outcomes and run them on every change. ValidatingAdmissionPolicy has no dedicated CLI for its CEL expressions, but because the expressions tend to be short, most teams simply try them with kubectl create --dry-run=server and watch how the API server evaluates the CEL directly. What all three approaches share is the same habit: see the outcome of a policy change locally or in CI first, instead of learning it from the first production deploy.
When to reach for which
If the team already has OPA experience, or if a policy genuinely needs to validate one resource against another, checking that a Deployment's referenced ConfigMap exists, for example, Gatekeeper's Rego expresses that kind of query naturally. Teams starting from scratch, who don't want the upfront cost of learning a new language, and who need mutate or generate capabilities, backfilling missing fields, auto-creating related resources, reach the same outcome with less friction using Kyverno. If the only need is simple, single-resource validation and the priority is finishing the job without installing an extra component, ValidatingAdmissionPolicy has the fewest moving parts; small clusters can adopt it without adding another Deployment, Service, and certificate to manage.
There's also a case for adopting none of them: in a small, single-namespace cluster run by a handful of people, these problems are usually solved well enough by PR review or a simple CI check, validating manifests against a schema with a tool like kubeconform. The real payoff of an admission-based policy engine starts once multiple teams share the same cluster and a central team needs to enforce standards automatically instead of by hand.