Kubernetes admission policy in 2026: revisit Kyverno, Gatekeeper and ValidatingAdmissionPolicy before you add another policy engine
Picture a health insurance technology company in Chennai (an invented example, not a real client). Its platform team runs eleven Kubernetes clusters. OPA Gatekeeper went in during 2021 with about forty constraint templates copied from the community library, and nobody has changed its webhook settings since. Last year a product team installed Kyverno on two clusters because they wanted image signature checks and default NetworkPolicies in every new namespace. Now an architect has read that Kubernetes can do admission policy on its own with CEL, and asks a simple question in the design review: why are we running two policy engines, and do we need either?
It is a fair question, and the answer is not one word. An admission policy engine decides which API requests are allowed, what gets changed on the way in, what happens when the engine itself is down, who can write exceptions, and what you learn about objects that were already in the cluster before the rule existed. This post walks through those decisions for Kyverno, OPA Gatekeeper and the built-in ValidatingAdmissionPolicy and MutatingAdmissionPolicy, with short notes on image verification tools such as Ratify and the Sigstore policy controller. Facts come from GitHub release pages, official docs, Kubernetes enhancement proposals, CNCF project pages and CVE records, checked on 11 October 2026. The drill near the end is my own run.
The short version
- The built-in policies are now complete for the common case. ValidatingAdmissionPolicy has been stable since Kubernetes 1.30 (April 2024), and MutatingAdmissionPolicy became stable and enabled by default in 1.36 (April 2026). Both run inside the API server, so there is no webhook pod to go down.
- Kyverno 1.19 (20 August 2026) reached feature parity between its CEL policy types and the old
ClusterPolicy, and officially deprecatedClusterPolicy,Policy,CleanupPolicyand the oldkyverno.ioPolicyException. Removal is planned for 1.20, estimated for November 2026. If you still writeClusterPolicyYAML, you are writing migration work for yourself. - Gatekeeper 3.23.1 (28 August 2026) can take a ConstraintTemplate written in CEL and generate a ValidatingAdmissionPolicy and binding from it, so the API server enforces the rule while Gatekeeper still audits. Its Helm chart still installs the validating webhook with
failurePolicy: Ignore. - What the built-ins do not do: audit existing objects, produce policy reports, generate new resources, call registries to verify signatures, or manage exceptions as objects. That is still the job of an engine.
- Kubernetes 1.37 (26 August 2026) turned on manifest-based admission control by default as a beta feature. Policies loaded from files on the control plane are active from API server start and cannot be deleted through the API. In my drill, a static policy blocked deletion of a policy binding.
- Kyverno published five more security advisories with 1.19.1 on 10 September 2026, one rated critical. If you run Kyverno with
apiCallor namespaced policies, 1.19.1 is the minimum.
Where things stand on 11 October 2026
Dates are from GitHub release pages and CNCF project pages, converted to IST.
Project
Latest release
What it is
CNCF status and licence
Kubernetes
1.37.1 on 24 September 2026 (1.37.0 on 26 August; 1.36.5 on 24 September)
ValidatingAdmissionPolicy (stable since 1.30), MutatingAdmissionPolicy (stable since 1.36), admission webhooks
Apache 2.0
Kyverno
1.19.1 on 10 September 2026 (1.19.0 on 20 August; 1.18.2 on 10 July)
Policy engine with CEL policy types, reports, generation, cleanup and image verification
Graduated on 16 March 2026; Apache 2.0
OPA Gatekeeper
3.23.1 on 28 August 2026 (3.23.0 on 9 July; 3.24.0-beta.0 on 13 July)
Admission webhook and audit built on OPA and the Constraint Framework, with Rego and CEL
Part of OPA, Graduated on 29 January 2021; Apache 2.0
Open Policy Agent
1.21.1 on 30 September 2026
General policy engine; the gator 3.23.1 binary I used reports OPA 1.17.1 inside
Graduated; Apache 2.0
Ratify
1.4.6 on 18 September 2026; 2.0.0-beta.2 on the same day
Verification engine for signatures and attestations, used as a Gatekeeper external data provider
Sandbox since 30 August 2024; Apache 2.0
Sigstore policy-controller
0.15.1 on 26 March 2026
Admission webhook that enforces Sigstore image policies
Apache 2.0
cosign
3.1.3 and 2.6.5 on 6 August 2026
Signs and verifies container images and attestations
Apache 2.0
Date
Event
October 2018
Gatekeeper repository created under the Open Policy Agent organisation
4 February 2019
Kyverno repository created at Nirmata
10 November 2020
Kyverno accepted into CNCF at Sandbox level
29 January 2021
OPA, which includes Gatekeeper, reaches CNCF Graduated level
13 July 2022
Kyverno moves to CNCF Incubating
9 December 2022
Kubernetes 1.26: ValidatingAdmissionPolicy alpha
15 August 2023
Kubernetes 1.28: ValidatingAdmissionPolicy beta, still off by default
18 April 2024
Kubernetes 1.30: ValidatingAdmissionPolicy stable and on by default
30 August 2024
Ratify accepted into CNCF at Sandbox level
12 December 2024
Kubernetes 1.32: MutatingAdmissionPolicy alpha
April 2025
Kyverno 1.14 adds the CEL-based ValidatingPolicy and ImageValidatingPolicy
25 July 2025
Gatekeeper 3.20.0: generating ValidatingAdmissionPolicy from templates becomes beta and on by default
July 2025
Kyverno 1.15 adds MutatingPolicy, GeneratingPolicy and DeletingPolicy
27 August 2025
Kubernetes 1.34: MutatingAdmissionPolicy beta, still off by default
2 February 2026
Kyverno 1.17: CEL policy types reach v1; ClusterPolicy marked for deprecation
16 March 2026
Kyverno reaches CNCF Graduated level
22 April 2026
Kubernetes 1.36: MutatingAdmissionPolicy stable and on by default; manifest-based admission control alpha
20 August 2026
Kyverno 1.19: CEL feature parity; ClusterPolicy officially deprecated, removal planned for 1.20
26 August 2026
Kubernetes 1.37: manifest-based admission control beta and on by default; webhooks stop receiving TokenReview and similar virtual resources by default
28 August 2026
Gatekeeper 3.23.1 fixes reconcile loops in generated ValidatingAdmissionPolicies
10 September 2026
Kyverno 1.19.1 fixes five advisories
What actually happens on an admission request
Every option in this post plugs into the same path inside the API server, and most surprises come from not knowing where in that path a rule runs.
- Authentication and authorisation. RBAC decides whether the user may make the request at all. Admission runs only after this.
- Mutating admission. Built-in mutating plugins, MutatingAdmissionPolicies and mutating webhooks can change the object. If a webhook changes it, built-in plugins run again, and webhooks with
reinvocationPolicy: IfNeededmay run again too. - Schema validation. The object is checked against its OpenAPI schema.
- Validating admission. ValidatingAdmissionPolicies and validating webhooks can reject the request, warn the client or write an audit annotation. They cannot change anything.
- Persistence. The object is written to etcd.
Two points follow. A rule that must see the final object belongs in validation, because a later mutating step can still change it. And admission sees only requests: an object created before a rule existed, or a change through a subresource your rule does not match, never passes through it. My drill showed both.
Design: webhooks versus CEL inside the API server
Webhooks: Kyverno and Gatekeeper
Kyverno and Gatekeeper register validating and mutating webhook configurations. For each matching request the API server sends an AdmissionReview over HTTPS to their service and waits. They can run any code, read cached cluster data, call registries and keep reports, but the API server now depends on pods running in the same cluster.
The Kubernetes webhook documentation sets the rules. timeoutSeconds must be between 1 and 30 and defaults to 10. failurePolicy decides what happens on errors such as connection failures, timeouts or bad responses: Fail rejects the request and Ignore lets it through. For admissionregistration.k8s.io/v1 the default is Fail. An explicit “deny” from a working webhook always denies, whatever the failure policy.
The two projects choose different defaults, and this is the most important line in your current setup:
Setting
Kyverno 1.19.1
Gatekeeper 3.23.1 Helm chart
Validating failure policy
Per policy; the ValidatingPolicy CRD says failurePolicy “Defaults to Fail”
validatingWebhookFailurePolicy: Ignore
Mutating failure policy
Per policy
mutatingWebhookFailurePolicy: Ignore
Webhook timeout
Per policy in spec.webhookConfiguration.timeoutSeconds; default 10 seconds
validatingWebhookTimeoutSeconds: 3, mutatingWebhookTimeoutSeconds: 1
Escape hatches
forceFailurePolicyIgnore and excludeBootstrapResources chart options, both off by default
Namespace exemptions; the docs’ emergency step is deleting the webhook configuration
Admission replicas
Admission controller needs at least three replicas for high availability
replicas: 3
Gatekeeper’s “Failing Closed” page explains its choice: when the webhook is down, constraints are not enforced, and audit is expected to catch what slipped in. Switch to Fail and you take on the circular dependency the page describes: if every node disappears, the Gatekeeper pods are gone, so new Node objects cannot be admitted, so the pods cannot return. Webhook configurations themselves are never sent through webhooks, so you can still delete them, unless an operator or GitOps controller puts them straight back.
Kyverno defaults the other way. The ControlPlane threat model it published in July 2026 lists “permissive failurePolicy: Ignore settings” among the most significant risks, so relaxing the default buys an enforcement gap. Fail has its own cost, which is why the chart’s excludeBootstrapResources option, off by default, keeps Node and CertificateSigningRequest objects away from Fail webhooks to avoid this deadlock after a full restart.
Neither default is wrong. A fail-open webhook is a hole during outages; a fail-closed webhook is a dependency for every write it matches. Decide per rule, scope the Fail rules narrowly, and keep the bootstrap path clear.
In-process CEL: ValidatingAdmissionPolicy and MutatingAdmissionPolicy
The built-in policies are evaluated by the API server itself. Each needs a policy (the CEL logic) and a binding (where it applies and what to do on failure). An optional parameter resource, such as a ConfigMap, lets one policy carry different values per namespace.
What you get by moving a rule here:
- No network hop and no webhook pod. The policy keeps working when your engine’s pods are down, which is why the Gatekeeper docs say in-process policies are “able to fail closed without impacting availability”.
- Validation actions on the binding.
Deny,WarnandAudit, withWarnandAuditusable together. That lets one policy deny in production and warn in development, as in my drill. - Type checking. When a policy is created, the expressions are checked against the schemas of the matched types and any problems appear in
status.typeChecking. On my bare API server, with no controller manager running, the status stayed empty, so do not rely on it in a minimal test setup. - Mutation without webhooks. MutatingAdmissionPolicy can return an apply configuration, merged with server-side apply rules, or a JSON Patch.
Before writing your own pod rules, remember that Pod Security Admission, stable since Kubernetes 1.25, already enforces the Pod Security Standards through namespace labels.
What you do not get: no background scan of existing objects, no reports beyond warnings and audit annotations, no network calls (so no registry lookups and no signature verification), no generation of other resources, and no exception objects.
A minimal policy and binding from my drill (this exact policy ran on a Kubernetes 1.37.1 API server):
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicy
metadata:
name: deploy-baseline.revisit.example
spec:
failurePolicy: Fail
matchConstraints:
resourceRules:
- apiGroups: ["apps"]
apiVersions: ["v1"]
operations: ["CREATE", "UPDATE"]
resources: ["deployments"]
variables:
- name: containers
expression: "object.spec.template.spec.containers"
validations:
- expression: "object.metadata.?labels['team'].orValue('') != ''"
message: "every Deployment needs a team label"
- expression: "variables.containers.all(c, c.image.contains('@sha256:') || (c.image.contains(':') && !c.image.endsWith(':latest')))"
messageExpression: "'pin images by tag or digest, not latest: ' + variables.containers.map(c, c.image).join(', ')"
---
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicyBinding
metadata:
name: deploy-baseline-prod.revisit.example
spec:
policyName: deploy-baseline.revisit.example
validationActions: [Deny]
matchResources:
namespaceSelector:
matchLabels:
env: prod
The image check is deliberately simple and has a known gap: an untagged image from a registry with a port, such as registry.example.in:5000/api, contains a colon and passes. Real image rules need proper parsing of the reference, so test them against awkward names before you enforce.
Static policies loaded from files
Kubernetes 1.37 made manifest-based admission control (KEP-5793) beta and on by default. The AdmissionConfiguration file given to --admission-control-config-file can point each admission plugin at a staticManifestsDir. Policies there are active from start and reloaded when files change. Their names must end in .static.k8s.io, a suffix the API refuses for normal objects, and unlike API-created policies they may match admission configuration objects, so they can protect your bindings from deletion.
apiVersion: apiserver.config.k8s.io/v1
kind: AdmissionConfiguration
plugins:
- name: ValidatingAdmissionPolicy
configuration:
apiVersion: apiserver.config.k8s.io/v1
kind: ValidatingAdmissionPolicyConfiguration
staticManifestsDir: "/etc/kubernetes/admission/policies/"
I ran this with one static policy. It matters most for self-managed control planes; on a managed service you usually cannot set API server flags, so check with your provider first.
Kyverno in 2026
From ClusterPolicy to CEL policy types
Kyverno began in 2019 with YAML patterns and JMESPath: ClusterPolicy with validate, mutate, generate and verifyImages rules, which is what most teams still run. From 1.14 it added CEL policy types in the policies.kyverno.io group, shaped like the Kubernetes built-ins. The 1.19 announcement maps them one to one:
Legacy rule type
CEL-based replacement
validate rules
ValidatingPolicy
mutate rules
MutatingPolicy
generate rules
GeneratingPolicy
verifyImages rules
ImageValidatingPolicy
CleanupPolicy
DeletingPolicy
PolicyException (kyverno.io)
PolicyException (policies.kyverno.io)
Each type has a namespaced variant, such as NamespacedValidatingPolicy, for namespace owners. The API reached v1 in 1.17, but the storage version stays v1beta1 until 1.20, when kyverno migrate rewrites stored objects. That is all kyverno migrate does; it does not convert a ClusterPolicy into a ValidatingPolicy. The conversion is manual, guided by the migration guide’s field-by-field table.
The guide is honest about gaps. validate.podSecurity becomes one CEL check per control, patchStrategicMerge becomes an apply configuration or JSON Patch, and failureActionOverrides and allowExistingViolations become policy exceptions. Existing CLI tests can be reused unchanged.
Since 1.19, creating a legacy policy returns an admission warning, a kyverno_deprecated_api_requests_total metric counts such requests, and kyverno apply and kyverno test print the warning and accept --warnings-as-errors for CI. My drill confirmed the CLI part.
Kyverno on top of the built-ins
A Kyverno ValidatingPolicy is, in the docs’ words, “a superset of a ValidatingAdmissionPolicy”. It adds background scanning, policy reports, exceptions, JSON payloads (a Terraform plan, for example) and extra CEL libraries for HTTP calls, cached cluster data, hashing, X.509 and time. It can also generate a native ValidatingAdmissionPolicy from itself with spec.autogen.validatingAdmissionPolicy.enabled: true, so the API server enforces it while Kyverno reports. The Helm chart has generateValidatingAdmissionPolicy on by default and generateMutatingAdmissionPolicy off. One catch from the docs: generating pod controller variants (Deployments, Jobs and so on) and generating a native policy are mutually exclusive for the same policy.
Footprint
Kyverno runs four Deployments: admission, reports, background (generate and mutate-existing) and cleanup. Only the admission controller serves requests on all replicas, and it needs at least three for high availability; the others use leader election, so extra replicas add availability, not throughput. Reports are custom resources in etcd, one per resource, and background scans run hourly by default. Kyverno 1.17 added --allowedResults so you can store, for example, only failures.
Kyverno’s scaling page publishes its own load test. In the vendor’s run (Kyverno 1.18.1, a 32 vCPU machine, a kind cluster with KWOK fake nodes, load from a separate machine), three admission replicas with 16 Pod Security ValidatingPolicies handled 10,000 pod creates from 200 virtual users at 145 ms average, 237 ms p95 and 284 ms p99, measured end to end at the client. That is the vendor’s number on the vendor’s setup, not a prediction for yours.
Support window and Kubernetes versions
Kyverno gives about three months of community patch support per minor release, limited to critical bugs and critical or high CVEs. On 11 October 2026 the supported line is 1.19, with end of life expected at the 1.20 release. One detail to check before you upgrade Kubernetes: the releases page lists Kubernetes 1.33 to 1.35 as tested for 1.19, while Kubernetes itself is at 1.37. Other versions “may work, but are not tested”.
Gatekeeper in 2026
Constraint templates, now in two languages
Gatekeeper splits policy into a ConstraintTemplate (logic and parameter schema) and Constraints (where it applies, with which values). Logic was Rego; since 3.18 a template can also carry a stable K8sNativeValidation engine in the same CEL as ValidatingAdmissionPolicy. If a template has both, CEL wins, with no fallback between engines. In my drill, the messages came from the CEL code.
- engine: K8sNativeValidation
source:
validations:
- expression: "variables.anyObject.metadata.?labels[variables.params.label].orValue('') != ''"
messageExpression: "'every Deployment needs a ' + variables.params.label + ' label'"
Gatekeeper exposes variables.anyObject because it sets object to oldObject on DELETE requests while Kubernetes does not, so policies behave the same in both places.
Generating native policies
Since 3.20, Gatekeeper generates a ValidatingAdmissionPolicy for every template with a CEL engine and a binding for each of its constraints; both defaults are on now that the feature is beta. deny, warn and dryrun map to Deny, Warn and Audit. With enforcementAction: scoped you can enforce at both vap.k8s.io and validation.gatekeeper.sh, so the webhook backs up the in-process path. The 3.23.1 notes show the feature settling: deterministic generated policies to stop a reconcile loop, less status churn and a default failure policy for CEL templates. The --sync-vap-enforcement-scope flag is deprecated and goes in 3.24.
Rego still matters for two cases the docs list explicitly: referential policies that look at other objects in the cluster (through Gatekeeper’s sync cache) and external data. Those cannot move to the API server.
Audit, mutation and external data
Gatekeeper’s audit re-evaluates existing objects every 60 seconds by default and keeps up to 20 violations per constraint in its status (the docs suggest at most 500, because of etcd object size). For larger estates, export results instead.
Mutation, stable since 3.10, uses small declarative CRDs: AssignMetadata (only adds labels and annotations), Assign, ModifySet and AssignImage. Gatekeeper does not generate other resources.
External data (beta since 3.11) lets Rego or mutators call a provider service. The docs recommend providers answer within one or two seconds, and the call is capped by the time the webhook has left, which is little with a 3-second timeout.
gator
gator is the offline CLI: gator test evaluates manifests against templates and constraints, gator verify runs test suites, and since 3.22 there is an alpha gator policy command for installing policies from the community library. The current docs also describe gator bench for timing policy evaluation. The docs say plainly that gator bench measures “compute-only policy evaluation latency” without network, TLS or API server cost. I used it in my drill with that limit in mind.
Mutation and generation compared
Need
Built-in
Kyverno
Gatekeeper
Add a default label or field
MutatingAdmissionPolicy (stable in 1.36)
MutatingPolicy
Assign, AssignMetadata
Change existing objects in the background
No
mutateExisting
No
Create a resource when another is created (NetworkPolicy per namespace, copied pull secret)
No
GeneratingPolicy, with sync
No
Delete resources on a schedule
No
DeletingPolicy
No
Rewrite an image reference to a digest
No
ImageValidatingPolicy mutateDigest
AssignImage changes parts of the image string; resolving a digest means calling out, for example through external data
If the Chennai team’s Kyverno use is mostly generation of namespace defaults, that alone justifies keeping an engine. If it is two labels and a security context default, MutatingAdmissionPolicy now covers it with no extra pods.
Audit, background scanning and reports
This is the gap that surprises teams moving to the built-ins. A ValidatingAdmissionPolicy sees only new requests. When you add a rule today, the hundred Deployments that already break it stay as they are, and nothing tells you so unless someone reads API server audit logs for the annotation.
Built-in
Kyverno
Gatekeeper
Existing objects checked
No
Background scan, hourly by default
Audit, every 60 seconds by default
Where results go
Client warnings, API server audit log
PolicyReport and ClusterPolicyReport (Policy Working Group format), per resource
Constraint status (capped), logs, exporters
Dry run before enforcing
Warn or Audit on the binding
Audit action, reports
dryrun or warn enforcement action
Exceptions
Change the binding’s selectors
PolicyException objects, with their own RBAC risk
Namespace exemptions, match excludes
A sensible pattern is to enforce the simple rules in the API server and keep one engine for audit and reports. Both Kyverno and Gatekeeper now support exactly that from one policy source.
Image verification
Verifying image signatures needs a network call to a registry and often to a transparency log, so it cannot run inside a built-in policy. You need a webhook. There are three common choices.
- Kyverno ImageValidatingPolicy, the CEL replacement for
verifyImages, supports cosign attestors (keys, KMS, keyless identities, certificates, custom trust roots) and Notary attestors. It can check attestations, require verification and rewrite tags to digests. - Gatekeeper with an external data provider. The Gatekeeper docs list Ratify and a cosign provider among community-maintained providers. Ratify’s quick start installs it with Gatekeeper 3.18 or later and shows Gatekeeper denying an unsigned image after a Notation check. Note that Ratify’s main branch is under active v2 development and the README warns it “may be unstable or broken”; the stable releases are on the 1.4 line.
- Sigstore policy-controller, a separate admission webhook from the Sigstore project that enforces image policies.
Whichever you pick, verify by digest, and keep tag resolution and verification in the same place. If your builds already produce signatures and SLSA provenance, as discussed in my image builders post, admission is where that evidence is finally checked. A sketch of a Kyverno policy (not run in my drill; the identity values are placeholders):
apiVersion: policies.kyverno.io/v1
kind: ImageValidatingPolicy
metadata:
name: require-signed-payments-images
spec:
matchConstraints:
resourceRules:
- apiGroups: [""]
apiVersions: ["v1"]
operations: ["CREATE"]
resources: ["pods"]
matchImageReferences:
- glob: "registry.example.in/payments/*"
attestors:
- name: ci-keyless
cosign:
keyless:
identities:
- subject: "https://git.example.in/payments/api/.ci/release.yaml@refs/heads/main"
issuer: "https://git.example.in"
Two cautions. Image verification makes admission depend on the registry, so set a timeout and decide the failure policy for that one rule deliberately. And it has had real bugs: CVE-2022-47633 allowed a verifyImages bypass through a malicious proxy or registry, CVE-2025-29778 made Kyverno ignore subjectRegExp and issuerRegExp in keyless checks, and the 1.19.1 advisory GHSA-5cjf-wwfg-pj4c let ImageValidatingPolicy exceptions bypass verification completely.
Security advisories
Every CVE ID below was checked against the CVE.org API (cveawg.mitre.org) on 11 October 2026 and returned a published record. Fixed versions are from the projects’ GitHub advisories.
CVE
Project
What it is
Fixed in
CVE-2026-54523
Kyverno
A namespace-scoped policy could make the background controller create resources, including RoleBindings, in other namespaces (critical)
1.18.2
CVE-2026-22039
Kyverno
Namespaced Policy apiCall ran with Kyverno’s own service account against any API path (critical)
1.15.3 and 1.16.3
CVE-2026-41068
Kyverno
Incomplete fix for CVE-2026-22039; cross-namespace reads still possible
1.17.2
CVE-2026-40868, CVE-2026-41323
Kyverno
apiCall service calls sent Kyverno’s service account token to the called endpoint
1.16.4
CVE-2026-4789
Kyverno
Server-side request forgery through the CEL HTTP functions, from 1.16.0
1.16.4
CVE-2026-41485
Kyverno
Controller crash through a forEach mutation panic
1.16.4 and 1.17.2
CVE-2026-23881
Kyverno
Denial of service through context variable amplification
1.15.3 and 1.16.3
CVE-2025-46342
Kyverno
Rules using namespace selectors in match could be bypassed
1.13.5 and 1.14.0
CVE-2025-29778
Kyverno
Keyless image verification ignored subjectRegExp and issuerRegExp
1.13.6 and 1.14.0
CVE-2024-48921
Kyverno
PolicyException objects could be created in any namespace by default
1.13.0
CVE-2022-47633
Kyverno
verifyImages bypass through a malicious proxy or registry (1.8.3 and 1.8.4)
1.8.5
CVE-2025-27403
Ratify
Azure authentication providers could send tokens to non-Azure registries
1.2.3 and 1.3.2
CVE-2025-1974
ingress-nginx
Remote code execution through the ingress-nginx admission webhook
1.11.5 and 1.12.1
Kyverno 1.19.1 also fixed five advisories that had GitHub IDs but no CVE IDs on 11 October 2026, so I list them by advisory only: GHSA-5qq8-67g6-4h2w (critical; privilege escalation to cluster admin through Policy apiCall urlPath), GHSA-c5qq-7g2q-cpqp (namespace isolation bypass through percent-encoded paths), GHSA-q825-p383-r9v5 (legacy apiCall and GlobalContextEntry missed the new egress controls), GHSA-59v6-2x73-wfg4 (namespaced policies could read cross-namespace GlobalContextEntry data) and GHSA-5cjf-wwfg-pj4c (ImageValidatingPolicy exception bypass). Gatekeeper’s GitHub security advisories page listed no published advisories when I checked.
The pattern matters more than any one entry. Most Kyverno issues sit in the features that make an engine more than the built-ins: API and URL calls, tenant-written namespaced policies, exceptions and generation. The engine runs with wide rights, so a tenant who can write policy can borrow them if the boundary leaks. Restrict who can create namespaced policies and exceptions, turn off apiCall and HTTP functions you do not need, and patch on a calendar. CVE-2025-1974 is a reminder from a different project, which I covered in the ingress-nginx retirement post: any admission webhook is also a network service that other pods might reach.
Licences and terms
Kubernetes, Kyverno, Gatekeeper, OPA, Ratify, cosign and the Sigstore policy-controller are all Apache 2.0, so the licence is not a deciding factor. Kyverno’s releases page points to commercial distributions for support beyond its roughly three-month community window. Policies copied from the community libraries (the Kyverno policies repository and the Gatekeeper library, both Apache 2.0) become yours to maintain; pin versions and review changes like code.
Migration paths
From Kyverno ClusterPolicy to CEL types. Use the migration guide’s field table, one CEL policy per rule, with your existing CLI tests as the acceptance check and --warnings-as-errors in CI. Watch kyverno_deprecated_api_requests_total before upgrading to 1.20; it shows what in your GitOps repositories still writes legacy objects. If Argo CD or Flux applies them, ordering of policy CRDs matters; my Argo CD and Flux post covers that side.
From Kyverno CEL policies to native policies. Turn on autogen.validatingAdmissionPolicy for the rules that only look at the incoming object. Keep Kyverno for reports, generation and images.
From Gatekeeper Rego to CEL. Add a K8sNativeValidation engine to templates whose Rego does not use sync data or external data, test with gator, and let Gatekeeper generate the native policy and binding. Keep Rego for referential rules.
From either engine to built-ins only. Possible for small, validation-only estates. You give up background audit, reports, generation and image verification, so plan replacements for those first, or accept that new rules apply only to new changes.
A decision guide
Your situation
A sensible starting point
Watch out for
A few simple validation and defaulting rules, clusters on 1.36 or later
ValidatingAdmissionPolicy and MutatingAdmissionPolicy only
No audit of existing objects; subresources; testing without a cluster
Existing Kyverno with ClusterPolicy
Migrate to CEL types on 1.19.1 before 1.20 removes them
Tested Kubernetes range; tenant-written policies; apiCall
Existing Gatekeeper with Rego templates
Add CEL engines and generate native policies; keep Gatekeeper for audit
Webhook still fails open by default; --sync-vap-enforcement-scope removal in 3.24
Need generation of per-namespace defaults or cleanup
Kyverno
Background controller rights; generated resource sync
Need image signature verification at admission
Kyverno ImageValidatingPolicy, or Gatekeeper with Ratify, or Sigstore policy-controller
Registry dependency; failure policy for that rule; verify by digest
Self-managed control plane, rules that must hold even during bootstrap
Static manifest policies on 1.37
Beta feature; files must be identical on every API server
Two engines already, as in Chennai
One engine, chosen by what you need beyond validation; move plain validation to native policies
Duplicate rules with different messages and exceptions
For the Chennai team my recommendation is in three steps. First, this month, check every webhook’s failure policy and timeout, and write down which rules would block a cluster restart. Second, move the plain validation rules, such as labels, image tag rules and replica limits, to CEL. With Gatekeeper this is a K8sNativeValidation engine and generated policies; on the Kyverno side, CEL ValidatingPolicies with native generation. Third, pick one engine for what remains. Because their Kyverno use is generation and image verification, which Gatekeeper does not do on its own, Kyverno is the likelier survivor, but only after they have migrated off ClusterPolicy and restricted who can write namespaced policies. Choose by the features you need beyond validation, not by benchmark tables.
A practical checklist
- List every
ValidatingWebhookConfigurationandMutatingWebhookConfiguration, with owner, failure policy, timeout and what it matches. - Write down what happens to each
Failwebhook if all its pods are down, and test the recovery step once. - Exclude bootstrap resources, the engine’s own namespace and lease objects from
Failwebhooks. - Move rules that look only at the incoming object to ValidatingAdmissionPolicy or MutatingAdmissionPolicy.
- For each moved rule, decide how existing violators will be found, since the built-ins will not report them.
- Check subresources such as
deployments/scale,pods/execandpods/ephemeralcontainers, and match them explicitly where it matters. - On Kyverno, upgrade to 1.19.1 or later, migrate off
ClusterPolicy, and add--warnings-as-errorsto CI. - On Gatekeeper, add CEL engines where Rego is not needed, and review constraint violation limits and audit export.
- Restrict RBAC on policy exceptions, namespaced policies and the engines’ own CRDs.
- Test every policy offline (Kyverno CLI or gator) and against a real API server before enforcing, starting with
WarnorAudit.
Common mistakes
- Assuming a new policy fixes the cluster. Admission only sees new requests; existing objects stay until something audits them.
- Matching
deploymentsbut notdeployments/scale. In my drill, a denied Deployment was still scaled through the subresource. - Leaving
failurePolicy: Ignoreon security rules without knowing it. Gatekeeper’s chart default is fail open. - Setting
Failon everything, including nodes and leases, then finding the cluster cannot heal after an outage. - Writing CEL that assumes fields exist. A missing field is a runtime error, and with
failurePolicy: Failthat error denies every matching request. - Running two engines with overlapping rules, so that teams get two different error messages and two exception processes for the same thing.
- Treating
kyverno migrateas a policy converter. It rewrites storage versions; the rule conversion is yours.
Drill: one policy, three evaluators and a webhook that is down
On 11 October 2026, from about 10:21 PM to 10:32 PM IST, on a shared Linux VM with 8 vCPUs, 15 GiB of RAM, Debian 13 and kernel 6.12, without root, I downloaded official binaries and checked each one: Kyverno CLI 1.19.1 and gator 3.23.1 against the SHA-256 digests GitHub lists for the release assets, kube-apiserver and kubectl 1.37.1 against the .sha256 files on dl.k8s.io, and etcd 3.7.0 from the controller-tools envtest-v1.37.0 archive, whose digest also matched. I ran etcd and a single kube-apiserver on 127.0.0.1 with token authentication and RBAC, and no controller manager, scheduler or nodes. Afterwards I stopped both processes and deleted the state directory and binaries.
Part 1: the same rule in three places. I wrote the “team label and no latest” rule as a ValidatingAdmissionPolicy (above), as a Kyverno ValidatingPolicy, and as a Gatekeeper ConstraintTemplate with both a CEL and a Rego engine. I tested four Deployment manifests: good, missing label, :latest and untagged nginx.
Evaluator
Good
No label
:latest
nginx
API server 1.37.1, binding with Deny (namespace env=prod)
Created
Denied
Denied
Denied
API server 1.37.1, binding with Warn and Audit (namespace env=dev)
Created
Created with warning
Created with warning
Created with warning
Kyverno CLI, ValidatingPolicy
Pass
Fail
Fail
Fail
Kyverno CLI, the native policy file with a binding that has no selector
Pass
Fail
Fail
Fail
gator, template with CEL and Rego
Pass
Violation
Violation
Violation
The verdicts matched. Offline, the Kyverno CLI gave no results for the native policy file with namespace-selector bindings, because it had no namespace labels to match, so I added a binding without a selector. A legacy ClusterPolicy printed the deprecation warning, and with --warnings-as-errors the command exited with status 1.
Part 2: built-in mutation and the limits of admission. A MutatingAdmissionPolicy added cost-centre: shared-platform to the good Deployment before validation ran. Then I relabelled the dev namespace to env=prod. The three violating Deployments created earlier stayed in place. Adding an annotation to one of them was denied, because an UPDATE is checked against the whole object. Running kubectl scale on the same Deployment succeeded and set replicas to 3, because my policy matched deployments and not deployments/scale.
Part 3: a typo with Fail. A policy with the expression object.replicas <= 5 (the field is spec.replicas) and a Deny binding rejected a valid Deployment with “no such key: replicas”. After I set failurePolicy: Ignore, the same Deployment was created. Type checking should have warned me, but on this setup without a controller manager status.typeChecking stayed empty.
Part 4: a webhook that is down. I registered a validating webhook for ConfigMaps in one namespace and pointed it at addresses with nothing behind them.
Webhook target
failurePolicy
timeoutSeconds
Result
Time for kubectl create
Closed port on 127.0.0.1
Fail
10
Rejected: connection refused
0.04 s
Closed port on 127.0.0.1
Ignore
10
Created
0.04 s
Unreachable address 10.255.255.1
Fail
2
Rejected: context deadline exceeded
2.04 s
Unreachable address 10.255.255.1
Ignore
2
Created
2.04 s
Unreachable address 10.255.255.1
Fail
10
Rejected: EOF
5.05 s
Unreachable address 10.255.255.1
Ignore
10
Created
5.05 s
No webhook
Created
0.03 s
A dead webhook that refuses connections fails fast. A webhook behind an address that does not answer costs every matching write its full timeout, whether the policy is Fail or Ignore. Ignore does not make an outage free; it only changes the answer. The 10-second rows ended at about 5 seconds with EOF, which means something in this VM’s network path closed the connection first, so they do not show the full default timeout.
Part 5: a static policy. I restarted the API server with an AdmissionConfiguration pointing the ValidatingAdmissionPolicy plugin at a directory containing one policy and binding named *.static.k8s.io. The log said “Loaded manifest-based admission policy configurations” with count 1. The static policy did not appear in kubectl get validatingadmissionpolicies, yet it denied my attempt to delete the dev policy binding. Creating an API object named fake.static.k8s.io was refused because the suffix is reserved.
Part 6: gator bench. I ran gator bench --engine=all --iterations=2000 on the template, constraint and four Deployments, three times. Mean latency per review was 125 to 127 microseconds for Rego and 54 to 55 for CEL; p99 was 598 to 611 for Rego and 260 to 290 for CEL. As gator itself warns, these are compute-only numbers for one tiny policy on a shared VM, not webhook latency.
Honest limits: one VM, one API server with no controller manager, nodes or high availability, and no real workload. I did not install Kyverno or Gatekeeper as webhooks, so I measured neither their admission latency nor their audit or reports, and I did not test image verification, generation, external data, Ratify or the Sigstore policy-controller. Each scenario ran once, except the bench. The webhook timings show behaviour, not performance; the only vendor numbers here are Kyverno’s, labelled as such.
What to unlearn and re-learn
- Unlearn “admission policy needs a policy engine”. Re-learn that validation and simple mutation now run inside the API server on 1.36 and later, and that engines earn their place through audit, reports, generation, exceptions and image verification.
- Unlearn “Kyverno policies are
ClusterPolicyYAML”. Re-learn thatClusterPolicyis deprecated in 1.19 and due for removal in 1.20, and that new Kyverno policies are CEL types shaped like the Kubernetes built-ins. - Unlearn “Gatekeeper means Rego”. Re-learn that templates can carry CEL, take priority over Rego, and be turned into native policies automatically.
- Unlearn “
failurePolicy: Ignoreis the safe choice”. Re-learn that it is an enforcement gap during outages, thatFailis a dependency, and that an unreachable webhook still costs every write its timeout. - Unlearn “once the policy is in, the cluster is compliant”. Re-learn that admission only sees new requests and matched resources, so audit and subresource rules are separate work.
Revisit your admission policy before you add another engine
The Chennai team does not need to rip anything out this week; it needs to know which rules it has, where they run and what happens when the thing running them is down. Learn the admission path and where each of your rules sits in it. Unlearn the habit of adding a new engine for each new team’s need, and the comfort of webhook settings nobody has looked at since 2021. Re-learn CEL, the built-in policies and the current Kyverno and Gatekeeper models from their release notes and docs, because both changed a lot this year. Practise by running one rule through the Kyverno CLI, gator and a real API server, and by pointing a test webhook at nothing, as I did on one VM. Then apply what you find: simple rules in the API server, one engine for what needs more, failure policies chosen on purpose and an audit that tells you about the objects admission never saw. Revisit it when you upgrade Kubernetes, when Kyverno 1.20 arrives, or when a new team asks for “just one more” policy tool.
Sources
- Kubernetes: ValidatingAdmissionPolicy, MutatingAdmissionPolicy, dynamic admission control (webhooks), admission controllers, feature gates, Pod Security Admission, KEP-3488 CEL admission control, KEP-3962 mutating admission policies, KEP-5793 manifest-based admission control, 1.37.1 release, 1.37.0 release, 1.36.0 release, 1.30.0 release, envtest binaries
- Kyverno: repository, 1.19.1 release, 1.19.0 release, Announcing Kyverno 1.19, Announcing Kyverno 1.17, policy types overview, ValidatingPolicy, MutatingPolicy, GeneratingPolicy, ImageValidatingPolicy, migration to CEL, reports, high availability, scaling, releases and support, kyverno migrate, kyverno apply, Helm values for 1.19.1, ValidatingPolicy CRD, ControlPlane threat model announcement, security advisories, policies library, LICENSE, CNCF: Kyverno
- Gatekeeper and OPA: Gatekeeper repository, 3.23.1 release, 3.23.0 release, 3.20.0 release, ValidatingAdmissionPolicy integration, failing closed, audit, mutation, external data, gator CLI, Helm values for 3.23.1, security advisories, Gatekeeper library, Gatekeeper LICENSE, OPA 1.21.1 release, CNCF: Open Policy Agent
- Image verification: Ratify repository, Ratify 1.4.6 release, Ratify 2.0.0-beta.2 release, Ratify quick start, Ratify security advisories, CNCF: Ratify, Sigstore policy-controller, policy-controller 0.15.1 release, cosign, cosign 3.1.3 release
- CVE records (cveawg.mitre.org API): CVE-2026-54523, CVE-2026-22039, CVE-2026-41068, CVE-2026-40868, CVE-2026-41323, CVE-2026-4789, CVE-2026-41485, CVE-2026-23881, CVE-2025-46342, CVE-2025-29778, CVE-2024-48921, CVE-2022-47633, CVE-2025-27403, CVE-2025-1974
Releted Posts
Container image builds in 2026: revisit BuildKit, Buildah and kaniko before you change your CI builder
Picture a payments software company in Pune (an invented example, not a real client). Since 2021 its GitLab runners on Kubernetes have built every service image with gcr.
Read morePostgreSQL high availability: revisit Patroni and CloudNativePG before your next failover
Picture a lending company in Bengaluru (an invented example, not a real client). Its loan ledger runs on PostgreSQL 16 across three VMs, managed by Patroni with a three-node etcd cluster and HAProxy in front.
Read moreFluent Bit, Vector and the OpenTelemetry Collector: revisit your log pipeline before you ship more logs
A logistics company in Pune runs about 600 pods across three Kubernetes clusters. A Fluentd DaemonSet, set up in 2019 with a dozen Ruby plugins, ships about 25 crore log lines a day to Loki.
Read more