Cilium and eBPF: What Changes in Network Policy and Observability
Why iptables-based network policy runs into scaling walls in large clusters, how eBPF fixes that, and what Cilium's identity-based L7 policies and Hubble's sidecar-free observability actually buy you in practice.
The problem: IP-based network policy does not match how pods actually behave
When a pod in Kubernetes is deleted and recreated, its IP address changes. For an iptables-based CNI this is a real scaling problem: every pod addition or removal rewrites iptables rules, the rule count grows roughly linearly with the number of pods, and netfilter evaluates those rules sequentially. In a cluster with a few hundred pods this is invisible. Past a few thousand pods, per-packet rule matching becomes a measurable source of latency, and reloading the rule set can take seconds.
The standard Kubernetes NetworkPolicy object is also limited by design: it only expresses allow/deny at L3 (IP/CIDR) and L4 (port/protocol). Rules like "the backend service should only accept GET, not POST" or "this pod may only reach api.example.com" fall outside its scope, because they require L7 (application layer) awareness that NetworkPolicy was never built to carry.
eBPF (extended Berkeley Packet Filter) addresses both problems at once. It loads small, verified programs into the kernel that process packets without a kernel/user-space round trip, and rule lookups happen through hash maps rather than iptables' sequential scan, which is far faster at scale. Cilium attaches these eBPF programs to network interfaces and uses them for service routing, network policy enforcement, and observability, all in a single data plane.
Identity-based security: labels, not IP addresses
Cilium's core departure from IP-centric networking is that the unit of policy is a security identity derived from pod labels, not the pod's IP. Pods sharing the same label set share the same identity; when a pod is recreated and gets a new IP, its identity stays the same, so policy does not need to be recomputed. This is the actual mechanism that keeps Cilium from hitting the rule-explosion problem that trips up iptables in large, frequently redeployed clusters.
The CiliumNetworkPolicy CRD is built on this identity model and extends the standard NetworkPolicy with L7 rules: HTTP method and path, gRPC service and method, Kafka topic-level permissions, and DNS-aware egress rules (toFQDNs) that let you say "this pod may only reach this specific domain." L7 filtering runs through an embedded Envoy proxy alongside eBPF's fast path: L3/L4 traffic stays entirely in the kernel, while traffic that needs L7 inspection is steered to the user-space proxy.
Example: allowing only one HTTP method
apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
name: backend-http-read-only
spec:
endpointSelector:
matchLabels:
app: backend
ingress:
- fromEndpoints:
- matchLabels:
app: frontend
toPorts:
- ports:
- port: "8080"
protocol: TCP
rules:
http:
- method: "GET"
path: "/api/v1/.*"
This policy allows only GET requests from frontend-labeled pods to backend; a POST or DELETE attempt over the same connection is rejected at the eBPF+Envoy layer. Standard NetworkPolicy cannot express this at all, since it can only open or close port 8080 as a whole.
An FQDN-based egress rule looks like this:
apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
name: worker-egress-payment-api
spec:
endpointSelector:
matchLabels:
app: worker
egress:
- toFQDNs:
- matchName: "api.payment-provider.example"
toPorts:
- ports:
- port: "443"
protocol: TCP
This is how you express "worker pods may only reach this one domain over 443" without hardcoding an IP range. Cilium watches DNS responses and updates the resolved IPs behind the policy automatically.
Hubble: the data eBPF already has, without a sidecar
Most service meshes get their observability by injecting a sidecar proxy into every pod, which costs extra resources and adds latency. Cilium's Hubble component pulls the same kind of data straight from the eBPF programs that are already processing the traffic, with no extra proxy involved. The hubble observe command shows live flows down to source/destination identity, HTTP method and path, DNS queries, and verdict (FORWARDED/DROPPED); Hubble Relay aggregates this across the cluster, and Hubble UI renders it as a service dependency map.
In practice this turns "why is this pod getting a 403" from a log-grepping exercise into a flow-filtering one: hubble observe --namespace backend --verdict DROPPED returns every rejected connection along with the policy that rejected it.
Pitfalls
Kernel version imposes limits you will not see coming. The feature set available to eBPF depends on the kernel version; advanced features like kube-proxy replacement and socket-based load balancing either do not work or work in a reduced form on older kernels. Assuming full Cilium functionality on a managed Kubernetes service without checking the node image's kernel version leads to a post-install "why isn't this feature available" surprise.
Replacing kube-proxy is hard to walk back. Configuring Cilium to fully replace kube-proxy moves service routing into eBPF; doing this migration on a cluster that is already serving traffic risks having two routing mechanisms fight each other during the transition. Starting a new cluster in this mode from day one is far safer than migrating a running one later.
Pushing policy straight to enforce mode is risky. Cilium lets you run new policies in audit/log mode first, observe real traffic against them, and only then switch to enforce. Skipping that step and applying a new CiliumNetworkPolicy directly in enforce mode can silently cut off a dependency you did not think about, such as a health check or a metrics scrape path. In production, audit first, observe, then enforce, in that order.
eBPF maps have a fixed, finite size. Connection tracking, service, and identity information all live in fixed-size eBPF maps; default map sizes can run out in large clusters, and a full map causes new connections to be dropped silently. This failure mode looks nothing like an iptables failure, so map occupancy needs to be checked separately with dedicated tooling.
The debugging model is not iptables. A team used to reading rules with iptables -L has to learn a different toolchain, the cilium CLI, hubble, bpftool, to see what the eBPF programs are actually doing in the kernel. That learning curve is real and should not be underestimated for a small team.
When not to reach for this
Cilium's real value shows up when you need L7 policy or sidecar-free observability. In a cluster with a few dozen pods where simple L3/L4 allow rules are enough, the operational complexity eBPF brings, kernel compatibility, a different debugging toolchain, map sizing, can outweigh the benefit, and a standard CNI's default NetworkPolicy support is enough. Similarly, if the CNI choice is locked in by a managed platform, since some managed Kubernetes offerings only support specific CNIs, verify which mode that platform actually supports Cilium in before designing around features that mode may not provide.