Kubernetes admission policy in 2026: revisit Kyverno, Gatekeeper and ValidatingAdmissionPolicy before you add another policy engine

Picture a health insurance technology company in Chennai (an invented example, not a real client). Its platform team runs eleven Kubernetes clusters. OPA Gatekeeper went in during 2021 with about forty constraint templates copied from the community library, and nobody has changed its webhook settings since. Last year a product team installed Kyverno on two clusters because they wanted image signature checks and default NetworkPolicies in every new namespace. Now an architect has read that Kubernetes can do admission policy on its own with CEL, and asks a simple question in the design review: why are we running two policy engines, and do we need either?

It is a fair question, and the answer is not one word. An admission policy engine decides which API requests are allowed, what gets changed on the way in, what happens when the engine itself is down, who can write exceptions, and what you learn about objects that were already in the cluster before the rule existed. This post walks through those decisions for Kyverno, OPA Gatekeeper and the built-in ValidatingAdmissionPolicy and MutatingAdmissionPolicy, with short notes on image verification tools such as Ratify and the Sigstore policy controller. Facts come from GitHub release pages, official docs, Kubernetes enhancement proposals, CNCF project pages and CVE records, checked on 11 October 2026. The drill near the end is my own run.

The short version

  • The built-in policies are now complete for the common case. ValidatingAdmissionPolicy has been stable since Kubernetes 1.30 (April 2024), and MutatingAdmissionPolicy became stable and enabled by default in 1.36 (April 2026). Both run inside the API server, so there is no webhook pod to go down.
  • Kyverno 1.19 (20 August 2026) reached feature parity between its CEL policy types and the old ClusterPolicy, and officially deprecated ClusterPolicy, Policy, CleanupPolicy and the old kyverno.io PolicyException. Removal is planned for 1.20, estimated for November 2026. If you still write ClusterPolicy YAML, you are writing migration work for yourself.
  • Gatekeeper 3.23.1 (28 August 2026) can take a ConstraintTemplate written in CEL and generate a ValidatingAdmissionPolicy and binding from it, so the API server enforces the rule while Gatekeeper still audits. Its Helm chart still installs the validating webhook with failurePolicy: Ignore.
  • What the built-ins do not do: audit existing objects, produce policy reports, generate new resources, call registries to verify signatures, or manage exceptions as objects. That is still the job of an engine.
  • Kubernetes 1.37 (26 August 2026) turned on manifest-based admission control by default as a beta feature. Policies loaded from files on the control plane are active from API server start and cannot be deleted through the API. In my drill, a static policy blocked deletion of a policy binding.
  • Kyverno published five more security advisories with 1.19.1 on 10 September 2026, one rated critical. If you run Kyverno with apiCall or namespaced policies, 1.19.1 is the minimum.

Where things stand on 11 October 2026

Dates are from GitHub release pages and CNCF project pages, converted to IST.

Project

Latest release

What it is

CNCF status and licence

Kubernetes

1.37.1 on 24 September 2026 (1.37.0 on 26 August; 1.36.5 on 24 September)

ValidatingAdmissionPolicy (stable since 1.30), MutatingAdmissionPolicy (stable since 1.36), admission webhooks

Apache 2.0

Kyverno

1.19.1 on 10 September 2026 (1.19.0 on 20 August; 1.18.2 on 10 July)

Policy engine with CEL policy types, reports, generation, cleanup and image verification

Graduated on 16 March 2026; Apache 2.0

OPA Gatekeeper

3.23.1 on 28 August 2026 (3.23.0 on 9 July; 3.24.0-beta.0 on 13 July)

Admission webhook and audit built on OPA and the Constraint Framework, with Rego and CEL

Part of OPA, Graduated on 29 January 2021; Apache 2.0

Open Policy Agent

1.21.1 on 30 September 2026

General policy engine; the gator 3.23.1 binary I used reports OPA 1.17.1 inside

Graduated; Apache 2.0

Ratify

1.4.6 on 18 September 2026; 2.0.0-beta.2 on the same day

Verification engine for signatures and attestations, used as a Gatekeeper external data provider

Sandbox since 30 August 2024; Apache 2.0

Sigstore policy-controller

0.15.1 on 26 March 2026

Admission webhook that enforces Sigstore image policies

Apache 2.0

cosign

3.1.3 and 2.6.5 on 6 August 2026

Signs and verifies container images and attestations

Apache 2.0

Date

Event

October 2018

Gatekeeper repository created under the Open Policy Agent organisation

4 February 2019

Kyverno repository created at Nirmata

10 November 2020

Kyverno accepted into CNCF at Sandbox level

29 January 2021

OPA, which includes Gatekeeper, reaches CNCF Graduated level

13 July 2022

Kyverno moves to CNCF Incubating

9 December 2022

Kubernetes 1.26: ValidatingAdmissionPolicy alpha

15 August 2023

Kubernetes 1.28: ValidatingAdmissionPolicy beta, still off by default

18 April 2024

Kubernetes 1.30: ValidatingAdmissionPolicy stable and on by default

30 August 2024

Ratify accepted into CNCF at Sandbox level

12 December 2024

Kubernetes 1.32: MutatingAdmissionPolicy alpha

April 2025

Kyverno 1.14 adds the CEL-based ValidatingPolicy and ImageValidatingPolicy

25 July 2025

Gatekeeper 3.20.0: generating ValidatingAdmissionPolicy from templates becomes beta and on by default

July 2025

Kyverno 1.15 adds MutatingPolicy, GeneratingPolicy and DeletingPolicy

27 August 2025

Kubernetes 1.34: MutatingAdmissionPolicy beta, still off by default

2 February 2026

Kyverno 1.17: CEL policy types reach v1; ClusterPolicy marked for deprecation

16 March 2026

Kyverno reaches CNCF Graduated level

22 April 2026

Kubernetes 1.36: MutatingAdmissionPolicy stable and on by default; manifest-based admission control alpha

20 August 2026

Kyverno 1.19: CEL feature parity; ClusterPolicy officially deprecated, removal planned for 1.20

26 August 2026

Kubernetes 1.37: manifest-based admission control beta and on by default; webhooks stop receiving TokenReview and similar virtual resources by default

28 August 2026

Gatekeeper 3.23.1 fixes reconcile loops in generated ValidatingAdmissionPolicies

10 September 2026

Kyverno 1.19.1 fixes five advisories

What actually happens on an admission request

Every option in this post plugs into the same path inside the API server, and most surprises come from not knowing where in that path a rule runs.

  1. Authentication and authorisation. RBAC decides whether the user may make the request at all. Admission runs only after this.
  2. Mutating admission. Built-in mutating plugins, MutatingAdmissionPolicies and mutating webhooks can change the object. If a webhook changes it, built-in plugins run again, and webhooks with reinvocationPolicy: IfNeeded may run again too.
  3. Schema validation. The object is checked against its OpenAPI schema.
  4. Validating admission. ValidatingAdmissionPolicies and validating webhooks can reject the request, warn the client or write an audit annotation. They cannot change anything.
  5. Persistence. The object is written to etcd.

Two points follow. A rule that must see the final object belongs in validation, because a later mutating step can still change it. And admission sees only requests: an object created before a rule existed, or a change through a subresource your rule does not match, never passes through it. My drill showed both.

Design: webhooks versus CEL inside the API server

Webhooks: Kyverno and Gatekeeper

Kyverno and Gatekeeper register validating and mutating webhook configurations. For each matching request the API server sends an AdmissionReview over HTTPS to their service and waits. They can run any code, read cached cluster data, call registries and keep reports, but the API server now depends on pods running in the same cluster.

The Kubernetes webhook documentation sets the rules. timeoutSeconds must be between 1 and 30 and defaults to 10. failurePolicy decides what happens on errors such as connection failures, timeouts or bad responses: Fail rejects the request and Ignore lets it through. For admissionregistration.k8s.io/v1 the default is Fail. An explicit “deny” from a working webhook always denies, whatever the failure policy.

The two projects choose different defaults, and this is the most important line in your current setup:

Setting

Kyverno 1.19.1

Gatekeeper 3.23.1 Helm chart

Validating failure policy

Per policy; the ValidatingPolicy CRD says failurePolicy “Defaults to Fail”

validatingWebhookFailurePolicy: Ignore

Mutating failure policy

Per policy

mutatingWebhookFailurePolicy: Ignore

Webhook timeout

Per policy in spec.webhookConfiguration.timeoutSeconds; default 10 seconds

validatingWebhookTimeoutSeconds: 3, mutatingWebhookTimeoutSeconds: 1

Escape hatches

forceFailurePolicyIgnore and excludeBootstrapResources chart options, both off by default

Namespace exemptions; the docs’ emergency step is deleting the webhook configuration

Admission replicas

Admission controller needs at least three replicas for high availability

replicas: 3

Gatekeeper’s “Failing Closed” page explains its choice: when the webhook is down, constraints are not enforced, and audit is expected to catch what slipped in. Switch to Fail and you take on the circular dependency the page describes: if every node disappears, the Gatekeeper pods are gone, so new Node objects cannot be admitted, so the pods cannot return. Webhook configurations themselves are never sent through webhooks, so you can still delete them, unless an operator or GitOps controller puts them straight back.

Kyverno defaults the other way. The ControlPlane threat model it published in July 2026 lists “permissive failurePolicy: Ignore settings” among the most significant risks, so relaxing the default buys an enforcement gap. Fail has its own cost, which is why the chart’s excludeBootstrapResources option, off by default, keeps Node and CertificateSigningRequest objects away from Fail webhooks to avoid this deadlock after a full restart.

Neither default is wrong. A fail-open webhook is a hole during outages; a fail-closed webhook is a dependency for every write it matches. Decide per rule, scope the Fail rules narrowly, and keep the bootstrap path clear.

In-process CEL: ValidatingAdmissionPolicy and MutatingAdmissionPolicy

The built-in policies are evaluated by the API server itself. Each needs a policy (the CEL logic) and a binding (where it applies and what to do on failure). An optional parameter resource, such as a ConfigMap, lets one policy carry different values per namespace.

What you get by moving a rule here:

  • No network hop and no webhook pod. The policy keeps working when your engine’s pods are down, which is why the Gatekeeper docs say in-process policies are “able to fail closed without impacting availability”.
  • Validation actions on the binding. Deny, Warn and Audit, with Warn and Audit usable together. That lets one policy deny in production and warn in development, as in my drill.
  • Type checking. When a policy is created, the expressions are checked against the schemas of the matched types and any problems appear in status.typeChecking. On my bare API server, with no controller manager running, the status stayed empty, so do not rely on it in a minimal test setup.
  • Mutation without webhooks. MutatingAdmissionPolicy can return an apply configuration, merged with server-side apply rules, or a JSON Patch.

Before writing your own pod rules, remember that Pod Security Admission, stable since Kubernetes 1.25, already enforces the Pod Security Standards through namespace labels.

What you do not get: no background scan of existing objects, no reports beyond warnings and audit annotations, no network calls (so no registry lookups and no signature verification), no generation of other resources, and no exception objects.

A minimal policy and binding from my drill (this exact policy ran on a Kubernetes 1.37.1 API server):

apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicy
metadata:
  name: deploy-baseline.revisit.example
spec:
  failurePolicy: Fail
  matchConstraints:
    resourceRules:
    - apiGroups: ["apps"]
      apiVersions: ["v1"]
      operations: ["CREATE", "UPDATE"]
      resources: ["deployments"]
  variables:
  - name: containers
    expression: "object.spec.template.spec.containers"
  validations:
  - expression: "object.metadata.?labels['team'].orValue('') != ''"
    message: "every Deployment needs a team label"
  - expression: "variables.containers.all(c, c.image.contains('@sha256:') || (c.image.contains(':') && !c.image.endsWith(':latest')))"
    messageExpression: "'pin images by tag or digest, not latest: ' + variables.containers.map(c, c.image).join(', ')"
---
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicyBinding
metadata:
  name: deploy-baseline-prod.revisit.example
spec:
  policyName: deploy-baseline.revisit.example
  validationActions: [Deny]
  matchResources:
    namespaceSelector:
      matchLabels:
        env: prod

The image check is deliberately simple and has a known gap: an untagged image from a registry with a port, such as registry.example.in:5000/api, contains a colon and passes. Real image rules need proper parsing of the reference, so test them against awkward names before you enforce.

Static policies loaded from files

Kubernetes 1.37 made manifest-based admission control (KEP-5793) beta and on by default. The AdmissionConfiguration file given to --admission-control-config-file can point each admission plugin at a staticManifestsDir. Policies there are active from start and reloaded when files change. Their names must end in .static.k8s.io, a suffix the API refuses for normal objects, and unlike API-created policies they may match admission configuration objects, so they can protect your bindings from deletion.

apiVersion: apiserver.config.k8s.io/v1
kind: AdmissionConfiguration
plugins:
- name: ValidatingAdmissionPolicy
  configuration:
    apiVersion: apiserver.config.k8s.io/v1
    kind: ValidatingAdmissionPolicyConfiguration
    staticManifestsDir: "/etc/kubernetes/admission/policies/"

I ran this with one static policy. It matters most for self-managed control planes; on a managed service you usually cannot set API server flags, so check with your provider first.

Kyverno in 2026

From ClusterPolicy to CEL policy types

Kyverno began in 2019 with YAML patterns and JMESPath: ClusterPolicy with validate, mutate, generate and verifyImages rules, which is what most teams still run. From 1.14 it added CEL policy types in the policies.kyverno.io group, shaped like the Kubernetes built-ins. The 1.19 announcement maps them one to one:

Legacy rule type

CEL-based replacement

validate rules

ValidatingPolicy

mutate rules

MutatingPolicy

generate rules

GeneratingPolicy

verifyImages rules

ImageValidatingPolicy

CleanupPolicy

DeletingPolicy

PolicyException (kyverno.io)

PolicyException (policies.kyverno.io)

Each type has a namespaced variant, such as NamespacedValidatingPolicy, for namespace owners. The API reached v1 in 1.17, but the storage version stays v1beta1 until 1.20, when kyverno migrate rewrites stored objects. That is all kyverno migrate does; it does not convert a ClusterPolicy into a ValidatingPolicy. The conversion is manual, guided by the migration guide’s field-by-field table.

The guide is honest about gaps. validate.podSecurity becomes one CEL check per control, patchStrategicMerge becomes an apply configuration or JSON Patch, and failureActionOverrides and allowExistingViolations become policy exceptions. Existing CLI tests can be reused unchanged.

Since 1.19, creating a legacy policy returns an admission warning, a kyverno_deprecated_api_requests_total metric counts such requests, and kyverno apply and kyverno test print the warning and accept --warnings-as-errors for CI. My drill confirmed the CLI part.

Kyverno on top of the built-ins

A Kyverno ValidatingPolicy is, in the docs’ words, “a superset of a ValidatingAdmissionPolicy”. It adds background scanning, policy reports, exceptions, JSON payloads (a Terraform plan, for example) and extra CEL libraries for HTTP calls, cached cluster data, hashing, X.509 and time. It can also generate a native ValidatingAdmissionPolicy from itself with spec.autogen.validatingAdmissionPolicy.enabled: true, so the API server enforces it while Kyverno reports. The Helm chart has generateValidatingAdmissionPolicy on by default and generateMutatingAdmissionPolicy off. One catch from the docs: generating pod controller variants (Deployments, Jobs and so on) and generating a native policy are mutually exclusive for the same policy.

Footprint

Kyverno runs four Deployments: admission, reports, background (generate and mutate-existing) and cleanup. Only the admission controller serves requests on all replicas, and it needs at least three for high availability; the others use leader election, so extra replicas add availability, not throughput. Reports are custom resources in etcd, one per resource, and background scans run hourly by default. Kyverno 1.17 added --allowedResults so you can store, for example, only failures.

Kyverno’s scaling page publishes its own load test. In the vendor’s run (Kyverno 1.18.1, a 32 vCPU machine, a kind cluster with KWOK fake nodes, load from a separate machine), three admission replicas with 16 Pod Security ValidatingPolicies handled 10,000 pod creates from 200 virtual users at 145 ms average, 237 ms p95 and 284 ms p99, measured end to end at the client. That is the vendor’s number on the vendor’s setup, not a prediction for yours.

Support window and Kubernetes versions

Kyverno gives about three months of community patch support per minor release, limited to critical bugs and critical or high CVEs. On 11 October 2026 the supported line is 1.19, with end of life expected at the 1.20 release. One detail to check before you upgrade Kubernetes: the releases page lists Kubernetes 1.33 to 1.35 as tested for 1.19, while Kubernetes itself is at 1.37. Other versions “may work, but are not tested”.

Gatekeeper in 2026

Constraint templates, now in two languages

Gatekeeper splits policy into a ConstraintTemplate (logic and parameter schema) and Constraints (where it applies, with which values). Logic was Rego; since 3.18 a template can also carry a stable K8sNativeValidation engine in the same CEL as ValidatingAdmissionPolicy. If a template has both, CEL wins, with no fallback between engines. In my drill, the messages came from the CEL code.

- engine: K8sNativeValidation
  source:
    validations:
    - expression: "variables.anyObject.metadata.?labels[variables.params.label].orValue('') != ''"
      messageExpression: "'every Deployment needs a ' + variables.params.label + ' label'"

Gatekeeper exposes variables.anyObject because it sets object to oldObject on DELETE requests while Kubernetes does not, so policies behave the same in both places.

Generating native policies

Since 3.20, Gatekeeper generates a ValidatingAdmissionPolicy for every template with a CEL engine and a binding for each of its constraints; both defaults are on now that the feature is beta. deny, warn and dryrun map to Deny, Warn and Audit. With enforcementAction: scoped you can enforce at both vap.k8s.io and validation.gatekeeper.sh, so the webhook backs up the in-process path. The 3.23.1 notes show the feature settling: deterministic generated policies to stop a reconcile loop, less status churn and a default failure policy for CEL templates. The --sync-vap-enforcement-scope flag is deprecated and goes in 3.24.

Rego still matters for two cases the docs list explicitly: referential policies that look at other objects in the cluster (through Gatekeeper’s sync cache) and external data. Those cannot move to the API server.

Audit, mutation and external data

Gatekeeper’s audit re-evaluates existing objects every 60 seconds by default and keeps up to 20 violations per constraint in its status (the docs suggest at most 500, because of etcd object size). For larger estates, export results instead.

Mutation, stable since 3.10, uses small declarative CRDs: AssignMetadata (only adds labels and annotations), Assign, ModifySet and AssignImage. Gatekeeper does not generate other resources.

External data (beta since 3.11) lets Rego or mutators call a provider service. The docs recommend providers answer within one or two seconds, and the call is capped by the time the webhook has left, which is little with a 3-second timeout.

gator

gator is the offline CLI: gator test evaluates manifests against templates and constraints, gator verify runs test suites, and since 3.22 there is an alpha gator policy command for installing policies from the community library. The current docs also describe gator bench for timing policy evaluation. The docs say plainly that gator bench measures “compute-only policy evaluation latency” without network, TLS or API server cost. I used it in my drill with that limit in mind.

Mutation and generation compared

Need

Built-in

Kyverno

Gatekeeper

Add a default label or field

MutatingAdmissionPolicy (stable in 1.36)

MutatingPolicy

Assign, AssignMetadata

Change existing objects in the background

No

mutateExisting

No

Create a resource when another is created (NetworkPolicy per namespace, copied pull secret)

No

GeneratingPolicy, with sync

No

Delete resources on a schedule

No

DeletingPolicy

No

Rewrite an image reference to a digest

No

ImageValidatingPolicy mutateDigest

AssignImage changes parts of the image string; resolving a digest means calling out, for example through external data

If the Chennai team’s Kyverno use is mostly generation of namespace defaults, that alone justifies keeping an engine. If it is two labels and a security context default, MutatingAdmissionPolicy now covers it with no extra pods.

Audit, background scanning and reports

This is the gap that surprises teams moving to the built-ins. A ValidatingAdmissionPolicy sees only new requests. When you add a rule today, the hundred Deployments that already break it stay as they are, and nothing tells you so unless someone reads API server audit logs for the annotation.

Built-in

Kyverno

Gatekeeper

Existing objects checked

No

Background scan, hourly by default

Audit, every 60 seconds by default

Where results go

Client warnings, API server audit log

PolicyReport and ClusterPolicyReport (Policy Working Group format), per resource

Constraint status (capped), logs, exporters

Dry run before enforcing

Warn or Audit on the binding

Audit action, reports

dryrun or warn enforcement action

Exceptions

Change the binding’s selectors

PolicyException objects, with their own RBAC risk

Namespace exemptions, match excludes

A sensible pattern is to enforce the simple rules in the API server and keep one engine for audit and reports. Both Kyverno and Gatekeeper now support exactly that from one policy source.

Image verification

Verifying image signatures needs a network call to a registry and often to a transparency log, so it cannot run inside a built-in policy. You need a webhook. There are three common choices.

  • Kyverno ImageValidatingPolicy, the CEL replacement for verifyImages, supports cosign attestors (keys, KMS, keyless identities, certificates, custom trust roots) and Notary attestors. It can check attestations, require verification and rewrite tags to digests.
  • Gatekeeper with an external data provider. The Gatekeeper docs list Ratify and a cosign provider among community-maintained providers. Ratify’s quick start installs it with Gatekeeper 3.18 or later and shows Gatekeeper denying an unsigned image after a Notation check. Note that Ratify’s main branch is under active v2 development and the README warns it “may be unstable or broken”; the stable releases are on the 1.4 line.
  • Sigstore policy-controller, a separate admission webhook from the Sigstore project that enforces image policies.

Whichever you pick, verify by digest, and keep tag resolution and verification in the same place. If your builds already produce signatures and SLSA provenance, as discussed in my image builders post, admission is where that evidence is finally checked. A sketch of a Kyverno policy (not run in my drill; the identity values are placeholders):

apiVersion: policies.kyverno.io/v1
kind: ImageValidatingPolicy
metadata:
  name: require-signed-payments-images
spec:
  matchConstraints:
    resourceRules:
    - apiGroups: [""]
      apiVersions: ["v1"]
      operations: ["CREATE"]
      resources: ["pods"]
  matchImageReferences:
  - glob: "registry.example.in/payments/*"
  attestors:
  - name: ci-keyless
    cosign:
      keyless:
        identities:
        - subject: "https://git.example.in/payments/api/.ci/release.yaml@refs/heads/main"
          issuer: "https://git.example.in"

Two cautions. Image verification makes admission depend on the registry, so set a timeout and decide the failure policy for that one rule deliberately. And it has had real bugs: CVE-2022-47633 allowed a verifyImages bypass through a malicious proxy or registry, CVE-2025-29778 made Kyverno ignore subjectRegExp and issuerRegExp in keyless checks, and the 1.19.1 advisory GHSA-5cjf-wwfg-pj4c let ImageValidatingPolicy exceptions bypass verification completely.

Security advisories

Every CVE ID below was checked against the CVE.org API (cveawg.mitre.org) on 11 October 2026 and returned a published record. Fixed versions are from the projects’ GitHub advisories.

CVE

Project

What it is

Fixed in

CVE-2026-54523

Kyverno

A namespace-scoped policy could make the background controller create resources, including RoleBindings, in other namespaces (critical)

1.18.2

CVE-2026-22039

Kyverno

Namespaced Policy apiCall ran with Kyverno’s own service account against any API path (critical)

1.15.3 and 1.16.3

CVE-2026-41068

Kyverno

Incomplete fix for CVE-2026-22039; cross-namespace reads still possible

1.17.2

CVE-2026-40868, CVE-2026-41323

Kyverno

apiCall service calls sent Kyverno’s service account token to the called endpoint

1.16.4

CVE-2026-4789

Kyverno

Server-side request forgery through the CEL HTTP functions, from 1.16.0

1.16.4

CVE-2026-41485

Kyverno

Controller crash through a forEach mutation panic

1.16.4 and 1.17.2

CVE-2026-23881

Kyverno

Denial of service through context variable amplification

1.15.3 and 1.16.3

CVE-2025-46342

Kyverno

Rules using namespace selectors in match could be bypassed

1.13.5 and 1.14.0

CVE-2025-29778

Kyverno

Keyless image verification ignored subjectRegExp and issuerRegExp

1.13.6 and 1.14.0

CVE-2024-48921

Kyverno

PolicyException objects could be created in any namespace by default

1.13.0

CVE-2022-47633

Kyverno

verifyImages bypass through a malicious proxy or registry (1.8.3 and 1.8.4)

1.8.5

CVE-2025-27403

Ratify

Azure authentication providers could send tokens to non-Azure registries

1.2.3 and 1.3.2

CVE-2025-1974

ingress-nginx

Remote code execution through the ingress-nginx admission webhook

1.11.5 and 1.12.1

Kyverno 1.19.1 also fixed five advisories that had GitHub IDs but no CVE IDs on 11 October 2026, so I list them by advisory only: GHSA-5qq8-67g6-4h2w (critical; privilege escalation to cluster admin through Policy apiCall urlPath), GHSA-c5qq-7g2q-cpqp (namespace isolation bypass through percent-encoded paths), GHSA-q825-p383-r9v5 (legacy apiCall and GlobalContextEntry missed the new egress controls), GHSA-59v6-2x73-wfg4 (namespaced policies could read cross-namespace GlobalContextEntry data) and GHSA-5cjf-wwfg-pj4c (ImageValidatingPolicy exception bypass). Gatekeeper’s GitHub security advisories page listed no published advisories when I checked.

The pattern matters more than any one entry. Most Kyverno issues sit in the features that make an engine more than the built-ins: API and URL calls, tenant-written namespaced policies, exceptions and generation. The engine runs with wide rights, so a tenant who can write policy can borrow them if the boundary leaks. Restrict who can create namespaced policies and exceptions, turn off apiCall and HTTP functions you do not need, and patch on a calendar. CVE-2025-1974 is a reminder from a different project, which I covered in the ingress-nginx retirement post: any admission webhook is also a network service that other pods might reach.

Licences and terms

Kubernetes, Kyverno, Gatekeeper, OPA, Ratify, cosign and the Sigstore policy-controller are all Apache 2.0, so the licence is not a deciding factor. Kyverno’s releases page points to commercial distributions for support beyond its roughly three-month community window. Policies copied from the community libraries (the Kyverno policies repository and the Gatekeeper library, both Apache 2.0) become yours to maintain; pin versions and review changes like code.

Migration paths

From Kyverno ClusterPolicy to CEL types. Use the migration guide’s field table, one CEL policy per rule, with your existing CLI tests as the acceptance check and --warnings-as-errors in CI. Watch kyverno_deprecated_api_requests_total before upgrading to 1.20; it shows what in your GitOps repositories still writes legacy objects. If Argo CD or Flux applies them, ordering of policy CRDs matters; my Argo CD and Flux post covers that side.

From Kyverno CEL policies to native policies. Turn on autogen.validatingAdmissionPolicy for the rules that only look at the incoming object. Keep Kyverno for reports, generation and images.

From Gatekeeper Rego to CEL. Add a K8sNativeValidation engine to templates whose Rego does not use sync data or external data, test with gator, and let Gatekeeper generate the native policy and binding. Keep Rego for referential rules.

From either engine to built-ins only. Possible for small, validation-only estates. You give up background audit, reports, generation and image verification, so plan replacements for those first, or accept that new rules apply only to new changes.

A decision guide

Your situation

A sensible starting point

Watch out for

A few simple validation and defaulting rules, clusters on 1.36 or later

ValidatingAdmissionPolicy and MutatingAdmissionPolicy only

No audit of existing objects; subresources; testing without a cluster

Existing Kyverno with ClusterPolicy

Migrate to CEL types on 1.19.1 before 1.20 removes them

Tested Kubernetes range; tenant-written policies; apiCall

Existing Gatekeeper with Rego templates

Add CEL engines and generate native policies; keep Gatekeeper for audit

Webhook still fails open by default; --sync-vap-enforcement-scope removal in 3.24

Need generation of per-namespace defaults or cleanup

Kyverno

Background controller rights; generated resource sync

Need image signature verification at admission

Kyverno ImageValidatingPolicy, or Gatekeeper with Ratify, or Sigstore policy-controller

Registry dependency; failure policy for that rule; verify by digest

Self-managed control plane, rules that must hold even during bootstrap

Static manifest policies on 1.37

Beta feature; files must be identical on every API server

Two engines already, as in Chennai

One engine, chosen by what you need beyond validation; move plain validation to native policies

Duplicate rules with different messages and exceptions

For the Chennai team my recommendation is in three steps. First, this month, check every webhook’s failure policy and timeout, and write down which rules would block a cluster restart. Second, move the plain validation rules, such as labels, image tag rules and replica limits, to CEL. With Gatekeeper this is a K8sNativeValidation engine and generated policies; on the Kyverno side, CEL ValidatingPolicies with native generation. Third, pick one engine for what remains. Because their Kyverno use is generation and image verification, which Gatekeeper does not do on its own, Kyverno is the likelier survivor, but only after they have migrated off ClusterPolicy and restricted who can write namespaced policies. Choose by the features you need beyond validation, not by benchmark tables.

A practical checklist

  1. List every ValidatingWebhookConfiguration and MutatingWebhookConfiguration, with owner, failure policy, timeout and what it matches.
  2. Write down what happens to each Fail webhook if all its pods are down, and test the recovery step once.
  3. Exclude bootstrap resources, the engine’s own namespace and lease objects from Fail webhooks.
  4. Move rules that look only at the incoming object to ValidatingAdmissionPolicy or MutatingAdmissionPolicy.
  5. For each moved rule, decide how existing violators will be found, since the built-ins will not report them.
  6. Check subresources such as deployments/scale, pods/exec and pods/ephemeralcontainers, and match them explicitly where it matters.
  7. On Kyverno, upgrade to 1.19.1 or later, migrate off ClusterPolicy, and add --warnings-as-errors to CI.
  8. On Gatekeeper, add CEL engines where Rego is not needed, and review constraint violation limits and audit export.
  9. Restrict RBAC on policy exceptions, namespaced policies and the engines’ own CRDs.
  10. Test every policy offline (Kyverno CLI or gator) and against a real API server before enforcing, starting with Warn or Audit.

Common mistakes

  • Assuming a new policy fixes the cluster. Admission only sees new requests; existing objects stay until something audits them.
  • Matching deployments but not deployments/scale. In my drill, a denied Deployment was still scaled through the subresource.
  • Leaving failurePolicy: Ignore on security rules without knowing it. Gatekeeper’s chart default is fail open.
  • Setting Fail on everything, including nodes and leases, then finding the cluster cannot heal after an outage.
  • Writing CEL that assumes fields exist. A missing field is a runtime error, and with failurePolicy: Fail that error denies every matching request.
  • Running two engines with overlapping rules, so that teams get two different error messages and two exception processes for the same thing.
  • Treating kyverno migrate as a policy converter. It rewrites storage versions; the rule conversion is yours.

Drill: one policy, three evaluators and a webhook that is down

On 11 October 2026, from about 10:21 PM to 10:32 PM IST, on a shared Linux VM with 8 vCPUs, 15 GiB of RAM, Debian 13 and kernel 6.12, without root, I downloaded official binaries and checked each one: Kyverno CLI 1.19.1 and gator 3.23.1 against the SHA-256 digests GitHub lists for the release assets, kube-apiserver and kubectl 1.37.1 against the .sha256 files on dl.k8s.io, and etcd 3.7.0 from the controller-tools envtest-v1.37.0 archive, whose digest also matched. I ran etcd and a single kube-apiserver on 127.0.0.1 with token authentication and RBAC, and no controller manager, scheduler or nodes. Afterwards I stopped both processes and deleted the state directory and binaries.

Part 1: the same rule in three places. I wrote the “team label and no latest” rule as a ValidatingAdmissionPolicy (above), as a Kyverno ValidatingPolicy, and as a Gatekeeper ConstraintTemplate with both a CEL and a Rego engine. I tested four Deployment manifests: good, missing label, :latest and untagged nginx.

Evaluator

Good

No label

:latest

nginx

API server 1.37.1, binding with Deny (namespace env=prod)

Created

Denied

Denied

Denied

API server 1.37.1, binding with Warn and Audit (namespace env=dev)

Created

Created with warning

Created with warning

Created with warning

Kyverno CLI, ValidatingPolicy

Pass

Fail

Fail

Fail

Kyverno CLI, the native policy file with a binding that has no selector

Pass

Fail

Fail

Fail

gator, template with CEL and Rego

Pass

Violation

Violation

Violation

The verdicts matched. Offline, the Kyverno CLI gave no results for the native policy file with namespace-selector bindings, because it had no namespace labels to match, so I added a binding without a selector. A legacy ClusterPolicy printed the deprecation warning, and with --warnings-as-errors the command exited with status 1.

Part 2: built-in mutation and the limits of admission. A MutatingAdmissionPolicy added cost-centre: shared-platform to the good Deployment before validation ran. Then I relabelled the dev namespace to env=prod. The three violating Deployments created earlier stayed in place. Adding an annotation to one of them was denied, because an UPDATE is checked against the whole object. Running kubectl scale on the same Deployment succeeded and set replicas to 3, because my policy matched deployments and not deployments/scale.

Part 3: a typo with Fail. A policy with the expression object.replicas <= 5 (the field is spec.replicas) and a Deny binding rejected a valid Deployment with “no such key: replicas”. After I set failurePolicy: Ignore, the same Deployment was created. Type checking should have warned me, but on this setup without a controller manager status.typeChecking stayed empty.

Part 4: a webhook that is down. I registered a validating webhook for ConfigMaps in one namespace and pointed it at addresses with nothing behind them.

Webhook target

failurePolicy

timeoutSeconds

Result

Time for kubectl create

Closed port on 127.0.0.1

Fail

10

Rejected: connection refused

0.04 s

Closed port on 127.0.0.1

Ignore

10

Created

0.04 s

Unreachable address 10.255.255.1

Fail

2

Rejected: context deadline exceeded

2.04 s

Unreachable address 10.255.255.1

Ignore

2

Created

2.04 s

Unreachable address 10.255.255.1

Fail

10

Rejected: EOF

5.05 s

Unreachable address 10.255.255.1

Ignore

10

Created

5.05 s

No webhook

Created

0.03 s

A dead webhook that refuses connections fails fast. A webhook behind an address that does not answer costs every matching write its full timeout, whether the policy is Fail or Ignore. Ignore does not make an outage free; it only changes the answer. The 10-second rows ended at about 5 seconds with EOF, which means something in this VM’s network path closed the connection first, so they do not show the full default timeout.

Part 5: a static policy. I restarted the API server with an AdmissionConfiguration pointing the ValidatingAdmissionPolicy plugin at a directory containing one policy and binding named *.static.k8s.io. The log said “Loaded manifest-based admission policy configurations” with count 1. The static policy did not appear in kubectl get validatingadmissionpolicies, yet it denied my attempt to delete the dev policy binding. Creating an API object named fake.static.k8s.io was refused because the suffix is reserved.

Part 6: gator bench. I ran gator bench --engine=all --iterations=2000 on the template, constraint and four Deployments, three times. Mean latency per review was 125 to 127 microseconds for Rego and 54 to 55 for CEL; p99 was 598 to 611 for Rego and 260 to 290 for CEL. As gator itself warns, these are compute-only numbers for one tiny policy on a shared VM, not webhook latency.

Honest limits: one VM, one API server with no controller manager, nodes or high availability, and no real workload. I did not install Kyverno or Gatekeeper as webhooks, so I measured neither their admission latency nor their audit or reports, and I did not test image verification, generation, external data, Ratify or the Sigstore policy-controller. Each scenario ran once, except the bench. The webhook timings show behaviour, not performance; the only vendor numbers here are Kyverno’s, labelled as such.

What to unlearn and re-learn

  • Unlearn “admission policy needs a policy engine”. Re-learn that validation and simple mutation now run inside the API server on 1.36 and later, and that engines earn their place through audit, reports, generation, exceptions and image verification.
  • Unlearn “Kyverno policies are ClusterPolicy YAML”. Re-learn that ClusterPolicy is deprecated in 1.19 and due for removal in 1.20, and that new Kyverno policies are CEL types shaped like the Kubernetes built-ins.
  • Unlearn “Gatekeeper means Rego”. Re-learn that templates can carry CEL, take priority over Rego, and be turned into native policies automatically.
  • Unlearn “failurePolicy: Ignore is the safe choice”. Re-learn that it is an enforcement gap during outages, that Fail is a dependency, and that an unreachable webhook still costs every write its timeout.
  • Unlearn “once the policy is in, the cluster is compliant”. Re-learn that admission only sees new requests and matched resources, so audit and subresource rules are separate work.

Revisit your admission policy before you add another engine

The Chennai team does not need to rip anything out this week; it needs to know which rules it has, where they run and what happens when the thing running them is down. Learn the admission path and where each of your rules sits in it. Unlearn the habit of adding a new engine for each new team’s need, and the comfort of webhook settings nobody has looked at since 2021. Re-learn CEL, the built-in policies and the current Kyverno and Gatekeeper models from their release notes and docs, because both changed a lot this year. Practise by running one rule through the Kyverno CLI, gator and a real API server, and by pointing a test webhook at nothing, as I did on one VM. Then apply what you find: simple rules in the API server, one engine for what needs more, failure policies chosen on purpose and an audit that tells you about the objects admission never saw. Revisit it when you upgrade Kubernetes, when Kyverno 1.20 arrives, or when a new team asks for “just one more” policy tool.

Sources

comments powered by Disqus

Releted Posts

Container image builds in 2026: revisit BuildKit, Buildah and kaniko before you change your CI builder

Picture a payments software company in Pune (an invented example, not a real client). Since 2021 its GitLab runners on Kubernetes have built every service image with gcr.

Read more

PostgreSQL high availability: revisit Patroni and CloudNativePG before your next failover

Picture a lending company in Bengaluru (an invented example, not a real client). Its loan ledger runs on PostgreSQL 16 across three VMs, managed by Patroni with a three-node etcd cluster and HAProxy in front.

Read more

Fluent Bit, Vector and the OpenTelemetry Collector: revisit your log pipeline before you ship more logs

A logistics company in Pune runs about 600 pods across three Kubernetes clusters. A Fluentd DaemonSet, set up in 2019 with a dozen Ruby plugins, ships about 25 crore log lines a day to Loki.

Read more