Skip to content

Fix two defects in the runc launch path - #941

Merged
crosbymichael merged 1 commit into
apple:mainfrom
crosbymichael:cz-runc-fixes
Sep 29, 2026
Merged

crosbymichael merged 1 commit into
apple:mainfrom
crosbymichael:cz-runc-fixes

Conversation

@crosbymichael

Copy link
Copy Markdown
Contributor

runc kill exits 1 with "container not running" once the init process has exited, and RuncProcess.kill propagated that. CRI requires StopContainer to succeed on an already-stopped container, so every teardown after a normal exit failed: containerd could not reap the task and sandboxes could not be removed. RuncExecProcess.kill and RuncProcess.resize already return early on .exited; only the container init path omitted the guard.

Separately, an OCI runtime refuses a net.* sysctl unless the container's spec owns a network namespace. Pod containers never do — they share the VM's, which is how they reach the pod's address — so runc create failed with sysctl "net.ipv4.ip_local_port_range" not allowed in host network namespace and the container never started, while vmexec applied the same key by writing /proc/sys. That check guards a shared kernel; in a VM the guest kernel is the pod's, so a pod-level net.* sysctl is pod-scoped. ManagedContainer now applies those keys itself and hands the runtime only the remainder, which is ipc- or uts-namespaced and therefore accepted. net.ipv4.ip_local_port_range, tcp_syncookies, ping_group_range and ip_unprivileged_port_start are four of Kubernetes' six safe sysctls, allowed with no node configuration, so this affected pods that opted into nothing.

`runc kill` exits 1 with "container not running" once the init process has
exited, and RuncProcess.kill propagated that. CRI requires StopContainer to
succeed on an already-stopped container, so every teardown after a normal exit
failed: containerd could not reap the task and sandboxes could not be removed.
RuncExecProcess.kill and RuncProcess.resize already return early on .exited;
only the container init path omitted the guard.

Separately, an OCI runtime refuses a `net.*` sysctl unless the container's spec
owns a network namespace. Pod containers never do — they share the VM's, which
is how they reach the pod's address — so `runc create` failed with
`sysctl "net.ipv4.ip_local_port_range" not allowed in host network namespace`
and the container never started, while vmexec applied the same key by writing
/proc/sys. That check guards a shared kernel; in a VM the guest kernel is the
pod's, so a pod-level net.* sysctl is pod-scoped. ManagedContainer now applies
those keys itself and hands the runtime only the remainder, which is ipc- or
uts-namespaced and therefore accepted. net.ipv4.ip_local_port_range,
tcp_syncookies, ping_group_range and ip_unprivileged_port_start are four of
Kubernetes' six safe sysctls, allowed with no node configuration, so this
affected pods that opted into nothing.

Signed-off-by: michael_crosby <michael_crosby@apple.com>
@crosbymichael
crosbymichael merged commit 1fca8c5 into apple:main Sep 29, 2026
7 checks passed
@crosbymichael
crosbymichael deleted the cz-runc-fixes branch September 29, 2026 18:03
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants