Docker 常见问题排查

Docker 常见问题排查

1 Error response from daemon: service endpoint with name es already exists

启动容器时报「端点已存在」。

排查:先查看容器当前所在的网络。

docker inspect discuz-mysql

输出中 Networks 一节可看到加入的网络:

"Networks": { "bridge": { "IPAMConfig": null, "Links": null, "Aliases": null, "NetworkID": "", "EndpointID": "", ... } }

再查看该网络中有哪些容器:

docker network inspect bridge

解决:将容器从该网络中断开:

docker network disconnect -f bridge discuz-mysql

2 Got permission denied while trying to connect to the Docker daemon socket

现象:docker 命令报

Got permission denied while trying to connect to the Docker daemon socket at unix:///var/run/docker.sock

解决(二选一):

# 方案一:临时放开 socket 权限(不太安全) sudo chmod 666 /var/run/docker.sock # 方案二(推荐):把当前用户加入 docker 组后重新登录 sudo usermod -aG docker $USER

3 OCI runtime create failed: container with is already exists

Error response from daemon: OCI runtime create failed: container with id exists

一般是同名容器残留,确认后强制删除:

docker rm -f <container_name_or_id>

4 The container name “/xxx” is already in use by container

现象:docker run 报

docker: Error response from daemon: Conflict. The container name "..." is already in use

解决:删除或重命名同名容器:

docker rm -f <container_name> # 或直接改名字再启动 docker rename <旧名> <新名>

5 无法删除已经 stopped 的容器:device or resource busy

现象:docker rm -f a19 报

driver "overlay2" failed to remove root filesystem: remove .../merged: device or resource busy

解决:多发生在 host 文件系统仍被占用时,可尝试:

# 1. 确认无进程占用后重试 lsof +D <容器 data-root 对应目录> # 2. 重启 docker daemon 后重新删除 sudo systemctl restart docker docker rm -f a19

6 无法启动已经 stopped 的容器:failed to start containers

现象:docker start a19 报

Error response from daemon: container "...": already exists Error: failed to start containers: a19

解决:容器状态与 daemon 记录不一致,一般删掉重建,或在重启 daemon 后重试。

7 OCI runtime create failed:pivot_root invalid argument

现象:docker run --rm busybox uname -a 报

process_linux.go:107: jailing process inside rootfs caused "pivot_root invalid argument"

解决:在 /etc/init.d/docker 添加以下配置,然后重启 docker:

export DOCKER_RAMDISK=true

8 copying bootstrap data to pipe caused “write init-p: broken pipe”

现象:构建或运行时提示 write init-p: broken pipe。

解决:升级内核到支持 overlay 的版本:

sudo apt-get install --install-recommends linux-generic-lts-xenial reboot

9 cgroups: cannot find cgroup mount destination

现象:docker run hello-world 报

WARNING: IPv4 forwarding is disabled. Networking will not work. cgroups: cannot find cgroup mount destination: unknown.

原因:host 上同时运行了两个 dockerd 进程。

解决:kill 掉多余的 dockerd 进程即可。

10 在容器中使用 systemctl

容器内一般没有 init/systemd。可替换为 python 实现的 systemctl:

# 下载 docker-systemctl-replacement 的 systemctl.py 覆盖容器内的 systemctl

参考:https://github.com/gdraheim/docker-systemctl-replacement

11 network xxx has active endpoints

现象:docker network rm <id> 报

ERROR: network docker-compose-elasticsearch-kibana_default ... has active endpoints

原因:该网络内还有其他 active 容器。

排查:

docker network inspect docker-compose-elasticsearch-kibana_default

输出 Containers 节点。解决:将相关容器从网络中断开:

docker network disconnect -f docker-compose-elasticsearch-kibana_default elasticsearch3 elasticsearch docker network inspect docker-compose-elasticsearch-kibana_default # 再次检查为空 docker network rm docker-compose-elasticsearch-kibana_default

12 unable to find user root: no matching entries in passwd file

现象:docker daemon 重启后,进入容器报 unable to find user root: no matching entries in passwd file。 原因尚不完全清楚,有两个解决办法:

# 方案1:docker exec 时显式指定 -u 0 docker exec -it -u 0 <container_id> /bin/bash # 方案2:重启容器 docker stop <container_id> docker start <container_id>

参考:https://stackoverflow.com/questions/41636759

13 核对容器跨主机通信

容器默认在 bridge 网络内,跨主机需要 overlay 网络或通过 host 端口转发,详见《容器跨主机相互访问》一文。

阅读 — · 全站 —
🎸 我的歌单 0 首