Introduction
Last time we met, we were building a cluster from Linux primitives like a couple of old pals. It was fun, and we shared some big laughs.
However, there is a bit more to do. For instance, I lobbed a fork bomb into one of the pods, and it crashed the entire virtual machine. Why would an accumulation of running processes consume all of the resources on the virtual machine? Why doesn’t it affect just the container?
Recall that containers are very different from virtual machines. I have written on this topic before (see On Virtualization And Virtual Machines), so I won’t go into detail here, but the main difference that you should be aware of is that all of the containers running on a (virtual) machine share that machine’s kernel. This is different from many virtual machines running on a hypervisor, where each virtual machine has its own (guest) kernel.
That difference is huge though, and it is very important to know when thinking about security and, in our case, resource sharing and consumption. The latter is what the aforementioned fork bomb exposed: since all of the containers share the same kernel, that means that they are all fighting over resources such as CPU, memory and I/O. When there are no resource limits on any container, then it can continue to use resources and starve other containers on the same virtual machine. So, left to their own devices, containers don’t play well together.
This hogging of resources, like greedy billionaires hoovering up cash at the expense of everyone else, is known as the noisy neighbor problem, and it describes the disruption of quality of service (QoS) of all the other hosts on a node that share the same kernel as the container (the neighbor) that is consuming all of the resources. It makes everyone involved sad.
The solution to this is to constrain the available resources such as CPU and memory by limiting the amount of consumption by any process. The Linux kernel subsystem that allows for this is control groups, or cgroups. Conveniently, the processes are organized into hierarchical groups, which allows for easier control and maintainability.
And, tie to a neat bow on the fork bomb example, cgroups can easily solve this issue by setting the maximum number of processes that a container can spawn. Easy peasy.
In this article, we’re going to look at adding any process in all pods created in the cluster to one of the cgroups. We’ll show how to do it manually for learning purposes, and then we’ll look at using systemd units to manage a cgroup hierarchy that makes maintainability easy.
This article will assume that you’ve read the previous one. The shell scripts and commands will have been described there. See On Building A Cluster From Linux Primitives.
cgroups
There are currently two versions of cgroups, but we will only be talking about cgroups-v2. This version contains big improvements over cgroups-v1 (which is now deprecated), notably a single kernel interface and the ability to organize in hierarchical groups.
The kernel interface is a virtual filesystem called cgroupfs and is mounted at /sys/fs/cgroup. Its type is cgroup2.
$ mount | grep cgroup
cgroup2 on /sys/fs/cgroup type cgroup2 (rw,nosuid,nodev,noexec,relatime,nsdelegate,memory_recursiveprot)
If you list out the contents of the cgroupfs, you’ll see a mixture of both versions:
$ ls -F /sys/fs/cgroup/
cgroup.controllers cgroup.stat cpuset.mems.effective io.cost.model memory.numa_stat proc-fs-nfsd.mount/ sys-kernel-tracing.mount/
cgroup.max.depth cgroup.subtree_control cpu.stat io.cost.qos memory.pressure proc-sys-fs-binfmt_misc.mount/ system.slice/
cgroup.max.descendants cgroup.threads dev-hugepages.mount/ io.pressure memory.reclaim sys-fs-fuse-connections.mount/ user.slice/
cgroup.pressure cpu.pressure dev-mqueue.mount/ io.stat memory.stat sys-kernel-config.mount/
cgroup.procs cpuset.cpus.effective init.scope/ machine.slice/ misc.capacity sys-kernel-debug.mount/
Most of the files in this directory are part of cgroups-v1, and we’re not going to worry about it. Instead, let’s look very briefly at the following directories only:
machine.slice- This is a directory that contains cgroups for any
systemd-nspawncontainers that are in use. - For example, I have a container that calls
hugoto build my website. When running, it appears in this directory. - See these fantastic articles!
- This is a directory that contains cgroups for any
system.slice- This contains directories for every service that is running.
user.slice- This directory will contain services for every user on the system, delineated by
uid.
- This directory will contain services for every user on the system, delineated by
Creating a new cgroup is simply a matter of making a new directory in /sys/fs/cgroup.
The directories ending in
.slicearesystemdunits, and we’ll get to them in thesystemdsection.
$ sudo bash /mnt/shared/cluster/pod.sh --nodens node0 --pods 1 --property MemoryMax=128M --property TasksMax=50
Running as unit: node0_pod0.service
Let’s get the PID of the reaper (the init process).
$ sudo bash /mnt/shared/cluster/ctl.sh --get pids --process sandbox-init
NODE | NODE IP | POD | POD IP | PIDS | PROCESS
node0 | 10.0.0.101 | node0-pod0 | 172.16.0.100 | 1172,1 | sandbox-init
If you’re not using the ctl.sh script, you can find it like this:
$ pidof sandbox-init
1172
Note that
pidofonly returns one PID in this example, because there is only onesandbox-initprocess currently executing.pidofcan return multiple values.
$ sudo nsenter --target 1172 --net --pid --mount --cgroup --ipc --uts --time -- bash
Ok, the next step is the most important. You must put the PID of the bash process in the cgroup or it won’t have any of the limitations inherited limits or the ones that may have been explicitly defined on the process.
You may be asking why this step is necessary, and it’s a good question! After all this is running in the same pid namespace as the other containers in the pod. The difference is subtle: unlike the other processes in the pod, this one was not able to be reparented to PID 1 (sandbox-init) in the pod. The reason is because I need a pty to run the fork bomb.
This helps demonstrate another reason why it is crucial to reparent a process to PID 1 in the new namespace. Not only can PID 1 then reap all of its children, but it will also have the new process join the parent’s cgroup. Of course, this will have the new process inherit all of the limits, which is exactly what we want!
When beginning to learn about cgroups (and namespaces), I found one of the most time-consuming things is to try and discover what PIDs belong to cgroup. What tools does Linux provide out-of-the-box, or at least a quick package manager download away?
Let’s see how many bash processes are running on the system:
$ ps u -C bash
USER PID %CPU %MEM VSZ RSS TTY STAT START TIME COMMAND
btoll 798 0.0 0.0 9096 5960 pts/0 Ss Oct04 0:00 -bash
btoll 1185 0.0 0.0 9228 6268 pts/2 Ss 00:09 0:00 -bash
root 1233 0.0 0.0 7340 4116 pts/1 S+ 00:09 0:00 bash
Now, we’re get more granular information by also printing the inode number of the namespace the process belongs to:
$ for pid in $(pidof bash); do sudo ps -o pid,comm,pidns -p "$pid"; done
PID COMMAND PIDNS
1185 bash 4026531836
PID COMMAND PIDNS
798 bash 4026531836
PID COMMAND PIDNS
1233 bash 4026532671
The last one looks like it may be the fella we’re interested in. Let’s compare it to the inode of the namespace that the init process belongs to:
$ sudo ls -l /proc/$(pidof sandbox-init)/ns/pid | awk '{print $NF}'
pid:[4026532671]
And, there it is. So, PID 1233 it is. Let’s see what cgroup it currently belongs to:
$ sudo cat /proc/1233/cgroup
0::/user.slice/user-1000.slice/session-1.scope
That’s not the cgroup that contains PID 1 of the pod in which the bash process is running. If the fork bomb were to be executed in that bash shell, it would bring down the whole virtual machine. Why? Because, it doesn’t have the limits needed to limit the damage. Observe:
$ cat /sys/fs/cgroup/user.slice/user-1000.slice/session-1.scope/{pids.max,memory.max}
max
max
The process isn’t contained at all in the amount of CPU and memory it can consume. This is a terrible, horrible, no-good rotten thing.
However, this teaches us a good lesson, children. When reparenting a process to PID 1, the process that was reparented is still a member of the cgroup in which it was created. To make it a member of a new cgroup, you must add it to the cgroup.procs file in the directory in the cgroups hierarchy:
$ /sys/fs/cgroup/node0.slice/node0-pod0.slice/node0_pod0.service echo 1233 | \
sudo tee cgroup.procs
If the process were spawned by PID 1, then it would automatically be a member of the same cgroup as its parent. Manually adding it as we did above is not necessary.
To sum up, if you don’t add the new PID to the cgroup, then it will not have the memory and pid limits, and it will consume all of the resources on the virtual machine.
Let’s take another look in procfs to see what cgroup the bash process now belongs to:
$ sudo cat /proc/1176/cgroup
0::/node0.slice/node0-pod0.slice/node0_pod0.service
Nice! And, it’s resources limits:
$ cat /sys/fs/cgroup/node0.slice/node0-pod0.slice/node0_pod0.service/{pids.max,memory.max}
50
134217728
That’s what we expected. Now, we can safely detonate the fork bomb and know that the damage is contained and won’t affect any container neighbors or the virtual machine:
$ :(){ :|:& };:
$ cat /sys/fs/cgroup/node0.slice/node0-pod0.slice/node0_pod0.service/pids.current
50
$ cat /sys/fs/cgroup/node0.slice/node0-pod0.slice/node0_pod0.service/memory.current
12943360
What is the fork bomb doing? It is calling itself and piping the results to another invocation and backgrounding itself. This will keep happening recursively until the resources are completely exhausted, leaving the machine inoperable.
Let’s now look to an easier and simpler way to manage our cgroup hierarchies using systemd.
systemd
systemd is an init system. At one point, it was controversial in the Linux community, and my understanding is because it didn’t follow the Unix philosophy of small tools doing one thing well that are composable. There are some distros that have still not adopted it. I have no opinion on this myself, but I respect both sides of the argument.
Having said that, let’s look into how systemd can help us simplify cgroups and aid in our lab cluster (see On Building A Cluster From Linux Primitives).
systemd has units that are perfect for creating hierarchies of cgroups.
system.slice is a directory that can contain other slices or a one or more services. However, a service cannot contain a slice. A service can only be a leaf node.
$ sudo systemctl status 1233
● node0_pod0.service - [systemd-run] /usr/sbin/ip netns exec node0-pod0 unshare --fork --pid --mount-proc --cgroup --ipc --uts --time -- /mnt/shared/cluster/sa>
Loaded: loaded (/run/systemd/transient/node0_pod0.service; transient)
Transient: yes
Active: active (running) since Sun 2026-10-04 20:32:39 UTC; 11min ago
Invocation: f35d071b0fd446e49019ea754db4ebd1
Main PID: 1171 (unshare)
Tasks: 50 (limit: 50)
Memory: 11.8M (max: 128M, available: 116.1M, peak: 14.6M)
CPU: 122ms
CGroup: /node0.slice/node0-pod0.slice/node0_pod0.service
├─1171 unshare --fork --pid --mount-proc --cgroup --ipc --uts --time -- /mnt/shared/cluster/sandbox-init
├─1172 /mnt/shared/cluster/sandbox-init
├─1233 bash
├─1432 bash
├─1434 bash
├─1435 bash
├─1436 bash
├─1437 bash
├─1438 bash
The following will create a new slice in cgroupfs kernel interface, i.e., /sys/fs/cgroup, which will be a directory with the name node0.slice.
systemctl set-property node0.slice MemoryMax=2G TasksMax=1000
systemd-run \
--no-block \
--slice=node0-pod0.slice \
--unit=node0_pod0.service \
-property MemoryMax=128M \
--property TasksMax=50 \
ip netns exec "$podns" unshare --fork --pid --mount-proc --cgroup --ipc --uts --time -- "$SOURCE_DIR/sandbox-init"
To stop a pod and have systemd send SIGTERM to the init in each service:
$ sudo systemctl stop node0-pod0.slice
To remove the systemd file, it’s necessary to run the following command after the pod has been stopped:
$ sudo systemctl reset-failed node0-pod0.service
To stop an entire node:
$ sudo systemctl stop node0.slice
You’ll need to do remove each systemd service file. Here’s a snippet that will find every pod in a node and remove it:
while read -r podns
do
! [[ "$podns" =~ $nodens ]] && continue
# `reset-failed` will remove the service (after stopping it, it will be in a `failed` state).
systemctl reset-failed "${podns/-/_}.service" # Replace the `-` with an `_` (systemd services CANNOT include a hyphen).
done < <(awk '$0 ~ /-/ {print $1}' <(ip netns list))
Note that this assumes that the network namespaces are in a specific format:
$ ip netns list
node1-pod1
node1-pod0
node0-pod0
node1 (id: 1)
node0 (id: 0)
The pods have a hyphen (-) in their name. The snippet above is taken from the shell script I used in the award-winning On Building A Cluster From Linux Primitives article.