Disaggregated Kubernetes: Contain the Blast Radius

Disaggregated Kubernetes: Contain the Blast Radius

Kubernetes gives us a basic namespace-based isolation boundary around workloads, but much of the infrastructure underneath those workloads remains highly privileged. Networking, storage, device emulation, lifecycle management, and control-plane services all process data originating from workloads. When those functions are concentrated in a small number of privileged components, a vulnerability in one can unfortunately have a large blast radius. Disaggregation takes a somewhat different approach.

How Disaggregation Isolates Infrastructure Services

Instead of putting infrastructure functionality into one large privileged component, Edera separates it into independently isolated services. Kubernetes can run within its own zone, while services such as networking and storage can run in separate trust domains. The hypervisor mediates access between them. This does a lot more for platform teams than simply running several individual processes. Each service in this architecture is only getting access to the resources it needs.

For example, Edera's storage backend uses grant tables to access specific pages of guest memory that have been explicitly shared for I/O. It does not automatically have access to the workload's entire address space. By turning page access into a handle, we reduce the risk of leaking unintended page contents. A network backend has a similarly limited view. That changes the consequences of a vulnerability. A bug in storage processing should compromise the storage service, not automatically give an attacker access to app memory or the networking stack. A network-related vulnerability should not automatically provide access to storage. The objective is not to eliminate bugs but instead to contain them to their respective zones.

What the Firecracker and KVM CVEs Reveal About VMM Boundaries

When considering vulnerabilities in VM device emulation, this is an important point to make. The recent Firecracker and KVM-specific vulnerabilities (CVE-2026-5747, CVE-2026-53359) have highlighted a broader architectural question of whether or not a vulnerability in that device-emulation layer or the virtualization kernel can cross a much larger trust boundary. In the case of Firecracker, it uses KVM for CPU and memory virtualization, while device emulation lives within the Firecracker VMM. This is not an argument that KVM or Firecracker are inherently insecure. It’s worth stating that monolithic VMMs have real advantages in simplicity and performance. The difference here is where the architecture chooses to put its security boundaries. When the KVM hypervisor is tied to a large set of complex services, it presents a security challenge that becomes difficult to handle. Linux’s process-based architecture for virtualization makes it tempting to solve this with namespaces and cgroups, but ultimately the shared kernel is still a large trusted computing base that is one bug away from breaking the entire host.

We chose our hypervisor architecture to put those boundaries around individual infrastructure services to limit how far these kinds of bugs can reach.

Disaggregated Devices to Disaggregated Kubernetes

The same idea can be applied above the device layer. Rather than giving Kubernetes broad access to a privileged infrastructure API, Edera is moving toward a model where Kubernetes consumes specific infrastructure capabilities. Networking is a capability. Storage is a capability. Control operations are also capabilities. This is where Edera's object-capability work becomes important.

Object Capabilities and Scoped Attestation

The earlier SPIRE implementation relied on broad access to the daemon Control API. That meant a privileged zone effectively received access to a rather large admin surface. With disaggregation combined with capabilities, SPIRE can instead receive a specific attestation capability. The SPIRE agent can ask for the identity of its own zone because the communication channel establishes who the caller is. It does not receive general access to the daemon, and it cannot simply claim to be another zone by supplying a different identifier.

The same model can apply to Kubernetes. A Kubernetes workload should not necessarily receive access to everything a node can do. It should receive the specific capabilities required to perform its job. That gives us a useful way of describing disaggregated Kubernetes. It’s Kubernetes where the infrastructure services are separated into independent trust domains and workloads receive narrowly-scoped capabilities rather than broad access to a privileged control plane.

The benefit is containment. A vulnerability in one service does not automatically become a vulnerability everywhere. Containers gave Kubernetes application isolation. VMs added another boundary around workloads. Disaggregation applies the same principle to the infrastructure underneath them. Edera is building that model with a low-level hypervisor, isolated device services, and object capabilities. The goal is simply to make sure a bug has fewer places to go when something inevitably breaks.

Disaggregating the Kubelet

This same question led us to revisit the Kubernetes node agent itself. At one point, we considered whether Edera should simply rewrite Kubelet from scratch. There is precedent for that approach: projects such as Krustlet explored what a Kubelet implemented in Rust might look like when applied to WASM, although that project is no longer active.

But rewriting Kubelet is not, by itself, a security architecture. The more useful question is: if you own the node-agent functions, how should they be divided? And that brings us back to the principle behind disaggregation: a component should not need to both parse attacker-controlled data and hold broad host privilege. Kubelet currently combines both responsibilities in a single, highly privileged process. When an input-processing vulnerability becomes a host-escape vulnerability, that combination is what makes the blast radius so large.

Looking at Kubelet through that lens also makes the existing Edera architecture more interesting. Two areas account for a significant portion of Kubelet's security-sensitive input handling: volume and mount-path processing, and container image management. Edera's architecture has already moved both of these responsibilities away from the privileged broker process. In particular, the volume path is handled outside the broker in a way that provides a stronger isolation boundary than the corresponding path in upstream Kubelet. It’s definitely an important architectural advantage, but isolation and confinement are simply not the same thing.

Today, the remaining Edera services are separated into multiple processes, but the systemd units themselves do not yet apply meaningful service hardening. The components run as root with the default capability set. The image provider is a separate address space, which limits the direct impact of a memory-safety bug in the image provider, but it is still a privileged root process. In other words, the architecture has taken the first step of separating the components, without yet taking the next step of constraining what each component is allowed to do.

The mechanism for going further already exists. Edera can publish capabilities from inside a zone, allowing a service to receive a narrowly scoped authority instead of relying on the daemon's full privilege. That mechanism was designed for precisely this kind of decomposition; the remaining work is to apply it consistently. 

An Incremental Path Instead of a Rewrite

This suggests a more incremental path than rewriting Kubelet. First, harden the services that already exist. Second, move image fetching and packing into an appropriately isolated zone. Third, separate the streaming server from the CRI process. None of those changes requires replacing Kubelet. They apply the same disaggregation principle to the node agent while preserving the parts of Kubernetes that already work.

A full Kubelet rewrite is still a possible direction, but it should be the consequence of an architectural decision rather than the starting point. If we do eventually rewrite it, the value add becomes much more than doing a Kubelet rewrite in a memory-safe language. We would instead be deciding which node-agent responsibilities deserve their own trust domains and which capabilities each of those services actually needs. There are two important limits to this approach.

First, disaggregation does not by itself reduce authority. A confined streaming service that can still ask the daemon to execute arbitrary operations in any zone is a smaller target, but it is not a harmless one. Reducing blast radius requires both process isolation and narrower authority. The service must not only be separated from the rest of the system; it must also receive only the capabilities required for its job.

Disaggregation vs. Reduced Privilege (KEP-2033)

Second, upstream Kubernetes is pursuing a different but complementary approach to Kubelet privilege. KEP-2033 introduced the concept of a rootless Kubelet, reducing the privileges available to the node agent itself rather than decomposing its responsibilities into separate trust domains. These approaches address different dimensions of the problem: reducing privilege limits what a compromised component can do, while disaggregation limits how far a compromise can spread between components and tenants.

The important point is that we do not need to choose between them. A less-privileged Kubelet can still benefit from being decomposed, and a decomposed node agent can still benefit from running with fewer privileges. The architectural goal is the same one that motivates disaggregated devices and capabilities throughout Edera: when something inevitably breaks, there should be fewer places for the failure to go.

FAQ

What is disaggregated Kubernetes?

Disaggregated Kubernetes separates infrastructure functions like networking and storage into independent trust domains and gives workloads narrowly scoped capabilities instead of broad access to a privileged control plane. A hypervisor mediates access, so a bug in one service stays contained rather than compromising the whole node.

Why does a shared kernel create a large blast radius in Kubernetes?

Namespaces and cgroups still rely on one shared Linux kernel, which is a large trusted computing base. An input-processing bug in a privileged component can become a host escape, giving an attacker reach across workloads and tenants. Disaggregation limits how far such a compromise can spread between components.

How do CVE-2026-5747 and CVE-2026-53359 relate to VMM device emulation?

These Firecracker and KVM vulnerabilities raise the question of whether a bug in the device-emulation layer or virtualization kernel can cross a larger trust boundary. Monolithic VMMs tie the hypervisor to complex services, so the architectural choice is where to place security boundaries, not whether KVM or Firecracker are insecure.

Does disaggregating the Kubelet require rewriting it?

No. Edera favors an incremental path: harden existing services, move image fetching and packing into an isolated zone, and separate the streaming server from the CRI process. These apply the disaggregation principle to the node agent while preserving the parts of Kubernetes that already work. A full rewrite remains a possible later decision.

‍

Cute cartoon axolotl with a light blue segmented body, big eyes, and dark gray external gills.

You know you wanna

Let’s solve this together