Report an Incident Become a Partner Careers Contact
Book a Demo
Threat Research

Escaping a Docker Container: An Attacker's Perspective

C CYBERSHIELD Threat Intelligence Team · Threat Intelligence, Wizard Cyber 5 February 2024 9 min read
A laptop on a desk displaying the Docker logo and wordmark

Docker is a cornerstone of modern deployment, and it is routinely mistaken for a security boundary. It is not one by default. A container shares the host's kernel, and what stops a process inside it reaching the host is a set of restrictions that can be — and frequently are — handed back at run time.

This is a write-up of what our team sees when a container is over-privileged: how that becomes host compromise, why it is so often invisible to the team running the workload, and what actually prevents it.

The boundary is thinner than people assume

A virtual machine has its own kernel. A container does not — it is a process on the host, wrapped in namespaces that limit what it can see and cgroups that limit what it can consume, with a restricted set of Linux capabilities.

Those capabilities are the interesting part. Linux splits root's historic all-or-nothing power into discrete privileges, and Docker grants a container a conservative subset by default. That default is reasonable. The problem is how often it is widened — by a developer chasing a permissions error, by a base image that assumes it, or by a CI pipeline that was easier to make work with --privileged than to debug.

Every capability added to a container is a piece of the host boundary handed back. Most of them are added to fix something that looked like a bug.

What we found

In a controlled lab, our team took a container granted CAP_SYS_PTRACE — the capability that permits debugging and inspection of other processes — and used it to reach the host.

We are not publishing the payload or the steps. The finding is what matters, and it is this: with that one capability and a shared process namespace, a process inside the container could act on a process belonging to the host, and code execution followed. No kernel exploit. No CVE. No vulnerable software anywhere in the chain. The container was configured to permit it, and the configuration was the whole attack.

That is the part worth internalising. A container escape is usually not a vulnerability you can patch. It is a permission somebody granted.

The capabilities that matter

If you audit one thing, audit these. Any of them in a container that does not have an explicit, documented reason for it should be treated as a finding:

  • --privileged — grants essentially everything and disables most isolation. There is almost never a good reason on a production workload, and it is by far the most common cause of a container escape.
  • CAP_SYS_PTRACE — process inspection and manipulation. The one above.
  • CAP_SYS_ADMIN — enormous, vaguely named, and frequently added because something needed one small part of it.
  • CAP_SYS_MODULE — load kernel modules. Game over for the host, immediately.
  • CAP_DAC_READ_SEARCH — bypass file read permission checks.

Two configuration choices belong in the same list, because both dissolve the boundary just as effectively:

  • The Docker socket mounted into a container. /var/run/docker.sock inside a container is root on the host, by design — anything that can talk to it can start a new container with the host filesystem mounted. This is extremely common in CI runners.
  • A shared host namespace — --pid=host, --net=host, --ipc=host. Each removes one of the walls the container model depends on.

Preventing it

Drop everything, add back what breaks

Start from --cap-drop=ALL and add only the capabilities the workload actually fails without. It is a short exercise and it is the single highest-value change available, because it converts "what did we leave open" into "what did we deliberately grant".

Run as a non-root user inside the container, and set --security-opt=no-new-privileges so a setuid binary cannot escalate within it.

Make it a pipeline check, not an audit

A container audited once drifts. The rule belongs where images and manifests are built: fail the build on privileged: true, on a mounted Docker socket, on a host namespace, and on any capability outside an agreed allow-list. Policy enforcement in the cluster admission path does the same job for anything that reaches Kubernetes.

Know what you are running

Most teams cannot answer "which of our containers are privileged" without going and looking. That question should have a dashboard, because the answer changes every time somebody ships.

Detecting it

Prevention is better, and prevention is never complete. What to watch for:

  • A container starting with elevated privileges — the run-time event itself, not the periodic audit. This is the cheapest high-value detection on the list.
  • Process activity crossing the container boundary: a process in a container interacting with one outside it, or a container process whose parent is not what the image should produce.
  • Unexpected outbound connections from a container. Most workloads talk to a small, knowable set of destinations. A shell calling out is rarely subtle when you have that baseline.
  • Developer tooling appearing at run time — a compiler, a debugger, a network utility inside a production container that was built without them.
  • Access to the Docker socket from any process that is not the daemon.

The broader point

Containers were built to solve a packaging problem, and they solved it. Isolation came along with them and is regularly mistaken for the primary feature.

The practical consequence is that container security is mostly configuration security. There is no patch for a capability somebody granted deliberately, and no scanner that flags a container image as vulnerable because of how it will later be run. The gap sits between the people who build images and the people who run them, and that is exactly where it tends to go unexamined.

Audit what you grant. Fail the build when something asks for more. And make sure somebody is watching the container run time, because the configuration that opens this door is usually added by someone who was only trying to make a deployment work.

Threat ResearchDockerContainersLinuxPrivilege EscalationDetection Engineering

Ready to see the agents work?

Book a demo of CYBERSHIELD AI against a real scenario.