11 min read

On cgroups


Introduction

Last time we met, we were building a cluster from Linux primitives like a couple of old pals. It was fun, and we shared some big laughs.

However, there is a bit more to do. For instance, I lobbed a fork bomb into one of the pods, and it crashed the entire virtual machine. Why would an accumulation of running processes consume all of the resources on the virtual machine? Why doesn’t it affect just the container?

Recall that containers are very different from virtual machines. I have written on this topic before (see On Virtualization And Virtual Machines), so I won’t go into detail here, but the main difference that you should be aware of is that all of the containers running on a (virtual) machine share that machine’s kernel. This is different from many virtual machines running on a hypervisor, where each virtual machine has its own (guest) kernel.

That difference is huge though, and it is very important to know when thinking about security and, in our case, resource sharing and consumption. The latter is what the aforementioned fork bomb exposed: since all of the containers share the same kernel, that means that they are all fighting over resources such as CPU, memory and I/O. When there are no resource limits on any container, then it can continue to use resources and starve other containers on the same virtual machine. So, left to their own devices, containers don’t play well together.

This hogging of resources, like greedy billionaires hoovering up cash at the expense of everyone else, is known as the noisy neighbor problem, and it describes the disruption of quality of service (QoS) of all the other hosts on a node that share the same kernel as the container (the neighbor) that is consuming all of the resources. It makes everyone involved sad.

The solution to this is to constrain the available resources such as CPU and memory by limiting the amount of consumption by any process. The Linux kernel subsystem that allows for this is control groups, or cgroups. Conveniently, the processes are organized into hierarchical groups, which allows for easier control and maintainability.

And, tie to a neat bow on the fork bomb example, cgroups can easily solve this issue by setting the maximum number of processes that a container can spawn. Easy peasy.

In this article, we’re going to look at adding any process in all pods created in the cluster to one of the cgroups. We’ll show how to do it manually for learning purposes, and then we’ll look at using systemd units to manage a cgroup hierarchy that makes maintainability easy.

This article will assume that you’ve read the previous one. The shell scripts and commands will have been described there. See On Building A Cluster From Linux Primitives.

cgroups

There are currently two versions of cgroups, but we will only be talking about cgroups-v2. This version contains big improvements over cgroups-v1 (which is now deprecated), notably a single kernel interface and the ability to organize in hierarchical groups.

The kernel interface is a virtual filesystem called cgroupfs and is mounted at /sys/fs/cgroup. Its type is cgroup2.

$ mount | grep cgroup
cgroup2 on /sys/fs/cgroup type cgroup2 (rw,nosuid,nodev,noexec,relatime,nsdelegate,memory_recursiveprot)

If you list out the contents of the cgroupfs, you’ll see a mixture of both versions:

$ ls -F /sys/fs/cgroup/
cgroup.controllers      cgroup.stat             cpuset.mems.effective  io.cost.model   memory.numa_stat  proc-fs-nfsd.mount/             sys-kernel-tracing.mount/
cgroup.max.depth        cgroup.subtree_control  cpu.stat               io.cost.qos     memory.pressure   proc-sys-fs-binfmt_misc.mount/  system.slice/
cgroup.max.descendants  cgroup.threads          dev-hugepages.mount/   io.pressure     memory.reclaim    sys-fs-fuse-connections.mount/  user.slice/
cgroup.pressure         cpu.pressure            dev-mqueue.mount/      io.stat         memory.stat       sys-kernel-config.mount/
cgroup.procs            cpuset.cpus.effective   init.scope/            machine.slice/  misc.capacity     sys-kernel-debug.mount/

Most of the files in this directory are part of cgroups-v1, and we’re not going to worry about it. Instead, let’s look very briefly at the following directories only:

  • machine.slice
  • system.slice
    • This contains directories for every service that is running.
  • user.slice
    • This directory will contain services for every user on the system, delineated by uid.

Creating a new cgroup is simply a matter of making a new directory in /sys/fs/cgroup.

The directories ending in .slice are systemd units, and we’ll get to them in the systemd section.

$ sudo bash /mnt/shared/cluster/pod.sh --nodens node0 --pods 1 --property MemoryMax=128M --property TasksMax=50
Running as unit: node0_pod0.service

Let’s get the PID of the reaper (the init process).

$ sudo bash /mnt/shared/cluster/ctl.sh --get pids --process sandbox-init
NODE    | NODE IP      | POD          | POD IP         | PIDS       | PROCESS
  node0 |   10.0.0.101 |   node0-pod0 |   172.16.0.100 |     1172,1 | sandbox-init

If you’re not using the ctl.sh script, you can find it like this:

$ pidof sandbox-init
1172

Note that pidof only returns one PID in this example, because there is only one sandbox-init process currently executing. pidof can return multiple values.

$ sudo nsenter --target 1172 --net --pid --mount --cgroup --ipc --uts --time -- bash

Ok, the next step is the most important. You must put the PID of the bash process in the cgroup or it won’t have any of the limitations inherited limits or the ones that may have been explicitly defined on the process.

You may be asking why this step is necessary, and it’s a good question! After all this is running in the same pid namespace as the other containers in the pod. The difference is subtle: unlike the other processes in the pod, this one was not able to be reparented to PID 1 (sandbox-init) in the pod. The reason is because I need a pty to run the fork bomb.

This helps demonstrate another reason why it is crucial to reparent a process to PID 1 in the new namespace. Not only can PID 1 then reap all of its children, but it will also have the new process join the parent’s cgroup. Of course, this will have the new process inherit all of the limits, which is exactly what we want!

When beginning to learn about cgroups (and namespaces), I found one of the most time-consuming things is to try and discover what PIDs belong to cgroup. What tools does Linux provide out-of-the-box, or at least a quick package manager download away?

Let’s see how many bash processes are running on the system:

$ ps u -C bash
USER         PID %CPU %MEM    VSZ   RSS TTY      STAT START   TIME COMMAND
btoll        798  0.0  0.0   9096  5960 pts/0    Ss   Oct04   0:00 -bash
btoll       1185  0.0  0.0   9228  6268 pts/2    Ss   00:09   0:00 -bash
root        1233  0.0  0.0   7340  4116 pts/1    S+   00:09   0:00 bash

Now, we’re get more granular information by also printing the inode number of the namespace the process belongs to:

$ for pid in $(pidof bash); do sudo ps -o pid,comm,pidns -p "$pid"; done
    PID COMMAND              PIDNS
   1185 bash            4026531836
    PID COMMAND              PIDNS
    798 bash            4026531836
    PID COMMAND              PIDNS
   1233 bash            4026532671

The last one looks like it may be the fella we’re interested in. Let’s compare it to the inode of the namespace that the init process belongs to:

$ sudo ls -l /proc/$(pidof sandbox-init)/ns/pid | awk '{print $NF}'
pid:[4026532671]

And, there it is. So, PID 1233 it is. Let’s see what cgroup it currently belongs to:

$ sudo cat /proc/1233/cgroup
0::/user.slice/user-1000.slice/session-1.scope

That’s not the cgroup that contains PID 1 of the pod in which the bash process is running. If the fork bomb were to be executed in that bash shell, it would bring down the whole virtual machine. Why? Because, it doesn’t have the limits needed to limit the damage. Observe:

$ cat /sys/fs/cgroup/user.slice/user-1000.slice/session-1.scope/{pids.max,memory.max}
max
max

The process isn’t contained at all in the amount of CPU and memory it can consume. This is a terrible, horrible, no-good rotten thing.

However, this teaches us a good lesson, children. When reparenting a process to PID 1, the process that was reparented is still a member of the cgroup in which it was created. To make it a member of a new cgroup, you must add it to the cgroup.procs file in the directory in the cgroups hierarchy:

$ /sys/fs/cgroup/node0.slice/node0-pod0.slice/node0_pod0.service echo 1233 | \
    sudo tee cgroup.procs

If the process were spawned by PID 1, then it would automatically be a member of the same cgroup as its parent. Manually adding it as we did above is not necessary.

To sum up, if you don’t add the new PID to the cgroup, then it will not have the memory and pid limits, and it will consume all of the resources on the virtual machine.

Let’s take another look in procfs to see what cgroup the bash process now belongs to:

$ sudo cat /proc/1176/cgroup
0::/node0.slice/node0-pod0.slice/node0_pod0.service

Nice! And, it’s resources limits:

$ cat /sys/fs/cgroup/node0.slice/node0-pod0.slice/node0_pod0.service/{pids.max,memory.max}
50
134217728

That’s what we expected. Now, we can safely detonate the fork bomb and know that the damage is contained and won’t affect any container neighbors or the virtual machine:

$ :(){ :|:& };:
$ cat /sys/fs/cgroup/node0.slice/node0-pod0.slice/node0_pod0.service/pids.current
50
$ cat /sys/fs/cgroup/node0.slice/node0-pod0.slice/node0_pod0.service/memory.current
12943360

What is the fork bomb doing? It is calling itself and piping the results to another invocation and backgrounding itself. This will keep happening recursively until the resources are completely exhausted, leaving the machine inoperable.

Let’s now look to an easier and simpler way to manage our cgroup hierarchies using systemd.

systemd

systemd is an init system. At one point, it was controversial in the Linux community, and my understanding is because it didn’t follow the Unix philosophy of small tools doing one thing well that are composable. There are some distros that have still not adopted it. I have no opinion on this myself, but I respect both sides of the argument.

Having said that, let’s look into how systemd can help us simplify cgroups and aid in our lab cluster (see On Building A Cluster From Linux Primitives).

systemd has units that are perfect for creating hierarchies of cgroups.

system.slice is a directory that can contain other slices or a one or more services. However, a service cannot contain a slice. A service can only be a leaf node.

$ sudo systemctl status 1233
● node0_pod0.service - [systemd-run] /usr/sbin/ip netns exec node0-pod0 unshare --fork --pid --mount-proc --cgroup --ipc --uts --time -- /mnt/shared/cluster/sa>
     Loaded: loaded (/run/systemd/transient/node0_pod0.service; transient)
  Transient: yes
     Active: active (running) since Sun 2026-10-04 20:32:39 UTC; 11min ago
 Invocation: f35d071b0fd446e49019ea754db4ebd1
   Main PID: 1171 (unshare)
      Tasks: 50 (limit: 50)
     Memory: 11.8M (max: 128M, available: 116.1M, peak: 14.6M)
        CPU: 122ms
     CGroup: /node0.slice/node0-pod0.slice/node0_pod0.service
             ├─1171 unshare --fork --pid --mount-proc --cgroup --ipc --uts --time -- /mnt/shared/cluster/sandbox-init
             ├─1172 /mnt/shared/cluster/sandbox-init
             ├─1233 bash
             ├─1432 bash
             ├─1434 bash
             ├─1435 bash
             ├─1436 bash
             ├─1437 bash
             ├─1438 bash

The following will create a new slice in cgroupfs kernel interface, i.e., /sys/fs/cgroup, which will be a directory with the name node0.slice.

systemctl set-property node0.slice MemoryMax=2G TasksMax=1000
systemd-run \
    --no-block \
    --slice=node0-pod0.slice \
    --unit=node0_pod0.service \
    -property MemoryMax=128M \
    --property TasksMax=50 \
    ip netns exec "$podns" unshare --fork --pid --mount-proc --cgroup --ipc --uts --time -- "$SOURCE_DIR/sandbox-init"

To stop a pod and have systemd send SIGTERM to the init in each service:

$ sudo systemctl stop node0-pod0.slice

To remove the systemd file, it’s necessary to run the following command after the pod has been stopped:

$ sudo systemctl reset-failed node0-pod0.service

To stop an entire node:

$ sudo systemctl stop node0.slice

You’ll need to do remove each systemd service file. Here’s a snippet that will find every pod in a node and remove it:

while read -r podns
do
    ! [[ "$podns" =~ $nodens ]] && continue
    # `reset-failed` will remove the service (after stopping it, it will be in a `failed` state).
    systemctl reset-failed "${podns/-/_}.service" # Replace the `-` with an `_` (systemd services CANNOT include a hyphen).
done < <(awk '$0 ~ /-/ {print $1}' <(ip netns list))

Note that this assumes that the network namespaces are in a specific format:

$ ip netns list
node1-pod1
node1-pod0
node0-pod0
node1 (id: 1)
node0 (id: 0)

The pods have a hyphen (-) in their name. The snippet above is taken from the shell script I used in the award-winning On Building A Cluster From Linux Primitives article.

References