Binpacking

4 min read · Intermediate


Overview

Binpacking is a resource optimization technique that allows Kubernetes workloads to be consolidated onto fewer nodes by making efficient use of available CPU, memory, and other resources.

In a traditional Kubernetes cluster, the scheduler places pods according to resource requests, constraints, affinity rules, and scheduling policies. Once workloads are running, however, there may be significant unused capacity distributed across nodes. Consolidating workloads onto fewer nodes can increase resource utilization and allow empty nodes to be removed, reducing infrastructure costs.

With Edera, workloads run in isolated zones while retaining the ability to use virtualization primitives such as dynamic resource adjustment, suspend/restore, and copy-on-write forking. These capabilities provide additional mechanisms for increasing workload density beyond what is possible through Kubernetes scheduling alone.

How does binpacking work?

Kubernetes scheduling is designed primarily around placing workloads safely and fairly, rather than continuously optimizing the cluster for maximum node utilization.

For example, a cluster might have several nodes where each node has enough capacity for its currently assigned workloads, but none has enough spare capacity to absorb another workload. The result is fragmented capacity:

Node 1       Node 2       Node 3
┌───────┐    ┌───────┐    ┌───────┐
│ Pod A │    │ Pod D │    │ Pod G │
│ Pod B │    │ Pod E │    │       │
│       │    │       │    │       │
└───────┘    └───────┘    └───────┘

     Capacity is available,
     but distributed across nodes.

In this case, Binpacking attempts to consolidate those workloads:

Node 1              Node 2
┌─────────────┐     ┌─────────┐
│ Pod A       │     │         │
│ Pod B       │     │  empty  │
│ Pod D       │     │         │
│ Pod E       │     │         │
│ Pod G       │     │         │
└─────────────┘     └─────────┘

                   → remove Node 2

For Kubernetes workloads, this is more noticeable when resource usage changes over time. A workload may request a large amount of memory or CPU but use only a fraction of it for much of its lifetime. Efficient binpacking requires both accurate resource accounting and mechanisms for reclaiming or sharing resources when they are not actively being used.

How Edera approaches binpacking

Edera runs each Kubernetes pod in its own security-isolated zone. This provides each workload with an independent kernel, but it also means that traditional host-level resource accounting does not always provide the complete picture of workload utilization. Edera can address this through several mechanisms.

1. Dynamic resource adjustment

Memory ballooning allows the amount of memory assigned to a zone to change based on actual workload demand. Unused memory can be reclaimed from one zone and made available to another without disrupting the pod. This allows the effective memory footprint of zones to more closely follow actual usage rather than remaining fixed at their maximum allocation.

2. Copy-on-write forking

Edera zones can be suspended and restored, and Xen PVH zones can also be forked from an existing suspended zone using copy-on-write (COW). Instead of allocating a completely independent copy of a workload’s memory, multiple zones can initially share memory pages. A page is copied only when one of the zones needs to modify it. This can significantly reduce the memory required when many workloads start from the same state, making this feature useful for highly-parallel workloads seen in reinforcement learning jobs.

3. Resource-aware placement

Binpacking also depends on understanding the resources available to each zone. CPU, memory, NUMA topology, and other placement constraints can affect whether workloads can safely share a node. Edera provides controls for CPU and memory sizing as well as NUMA-aware placement. The result is a combination of Kubernetes scheduling and hypervisor-level resource management, as seen below:

Kubernetes
    │
    │ places workloads
    ▼
┌─────────────────────────────┐
│           Node              │
│                             │
│  ┌──────┐ ┌──────┐ ┌──────┐ │
│  │Zone A│ │Zone B│ │Zone C│ │
│  └──────┘ └──────┘ └──────┘ │
│       ▲       ▲       ▲     │
│       │       │       │     │
│   dynamic resource          │
│       adjustment            │
└─────────────────────────────┘

4. Density considerations

There is often a tradeoff between workload isolation and workload density. Each Edera zone has overhead associated with running its own kernel and userspace environment. In particular, zones have a minimum memory footprint, so small workloads may achieve lower raw density than containers sharing a single host kernel. It’s worth noting that binpacking on Kubernetes is best done when memory requests are not specified or are specified really low.

Edera’s resource-management primitives were intended to reduce this overhead where possible by allowing memory to be dynamically reclaimed and, in some workloads, shared through copy-on-write. The appropriate comparison is therefore not simply containers versus zones, but the total node capacity consumed by the workloads being consolidated and the degree to which their resources can be dynamically shared.

Last updated on