AKS容器exec命令短时间后执行失败问题排查求助
问题描述
在Azure Kubernetes Service(AKS)上部署了微软官方教程中的应用,该应用会启动两个Pod。Pod刚实例化完成时可以通过kubectl exec进入容器,但不久后该命令执行失败,无法定位根因。
执行的命令
dtaylor@DESKTOP-98A8U05:~/development/terraform/aqua_azure_terraform$ source envvars.sh dtaylor@DESKTOP-98A8U05:~/development/terraform/aqua_azure_terraform$ kubectl get pods NAME READY STATUS RESTARTS AGE azure-vote-back-65c595548d-2bmvx 0/1 ContainerCreating 0 33s azure-vote-front-d99b7676c-hxdwt 0/1 ContainerCreating 0 33s dtaylor@DESKTOP-98A8U05:~/development/terraform/aqua_azure_terraform$ watch kubectl get pods dtaylor@DESKTOP-98A8U05:~/development/terraform/aqua_azure_terraform$ kubectl get service azure-vote-front NAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGE azure-vote-front LoadBalancer 10.0.68.181 20.248.234.249 80:30027/TCP 2m13s dtaylor@DESKTOP-98A8U05:~/development/terraform/aqua_azure_terraform$ kubectl exec --stdin --tty azure-vote-front-d99b7676c-hxdwt-- /bin/bash kubectl exec [POD] [COMMAND] is DEPRECATED and will be removed in a future version. Use kubectl exec [POD] -- [COMMAND] instead. Error from server (NotFound): pods "azure-vote-front-d99b7676c-hxdwt--" not found dtaylor@DESKTOP-98A8U05:~/development/terraform/aqua_azure_terraform$ kubectl exec --stdin --tty azure-vote-front-d99b7676c-hxdwt -- /bin/bash root@azure-vote-front-d99b7676c-hxdwt:/app# exit dtaylor@DESKTOP-98A8U05:~/development/terraform/aqua_azure_terraform$ kubectl exec --stdin --tty azure-vote-front-d99b7676c-hxdwt -- /bin/bash root@azure-vote-front-d99b7676c-hxdwt:/app# exit dtaylor@DESKTOP-98A8U05:~/development/terraform/aqua_azure_terraform$ kubectl exec --stdin --tty azure-vote-front-d99b7676c-hxdwt -- /bin/bash root@azure-vote-front-d99b7676c-hxdwt:/app# dtaylor@DESKTOP-98A8U05:~/development/terraform/aqua_azure_terraform$ kubectl exec --stdin --tty azure-vote-front-d99b7676c-hxdwt -- /bin/bash error: Internal error occurred: error executing command in container: failed to exec in container: failed to start exec "54fb73a731803d4efb811970ede9dc9ef886aba616ece4f4a319a2bce5d863f6": OCI runtime exec failed: /usr/bin/runc did not terminate successfully: exit status 137: : unknown dtaylor@DESKTOP-98A8U05:~/development/terraform/aqua_azure_terraform$
Pod日志
dtaylor@DESKTOP-98A8U05:~/development/terraform/aqua_azure_terraform$ kubectl logs azure-vote-front-d99b7676c-hxdwt Checking for script in /app/prestart.sh Running script /app/prestart.sh Running inside /app/prestart.sh, you could add migrations to this file, e.g.: #! /usr/bin/env bash # Let the DB start sleep 10; # Run migrations alembic upgrade head /usr/lib/python2.7/dist-packages/supervisor/options.py:298: UserWarning: Supervisord is running as root and it is searching for its configuration file in default locations (including its current working directory); you probably want to specify a "-c" argument specifying an absolute path to a configuration file for improved security. 'Supervisord is running as root and it is searching ' 2023-08-10 00:33:44,775 CRIT Supervisor running as root (no user in config file) 2023-08-10 00:33:44,775 INFO Included extra file "/etc/supervisor/conf.d/supervisord.conf" during parsing 2023-08-10 00:33:44,782 INFO RPC interface 'supervisor' initialized 2023-08-10 00:33:44,783 CRIT Server 'unix_http_server' running without any HTTP authentication checking 2023-08-10 00:33:44,783 INFO supervisord started with pid 7 2023-08-10 00:33:45,850 INFO spawned: 'nginx' with pid 10 2023-08-10 00:33:45,852 INFO spawned: 'uwsgi' with pid 11 [uWSGI] getting INI configuration from /app/uwsgi.ini [uWSGI] getting INI configuration from /etc/uwsgi/uwsgi.ini *** Starting uWSGI 2.0.15 (64bit) on [Thu Aug 10 00:33:45 2023] *** compiled with version: 6.3.0 20170516 on 14 December 2017 18:49:24 os: Linux-5.15.0-1042-azure #49-Ubuntu SMP Tue Jul 11 17:28:46 UTC 2023 nodename: azure-vote-front-d99b7676c-hxdwt machine: x86_64 clock source: unix pcre jit disabled detected number of CPU cores: 2 current working directory: /app detected binary path: /usr/local/bin/uwsgi your memory page size is 4096 bytes detected max file descriptor number: 1048576 lock engine: pthread robust mutexes thunder lock: disabled (you can enable it with --thunder-lock) uwsgi socket 0 bound to UNIX address /tmp/uwsgi.sock fd 3 uWSGI running as root, you can use --uid/--gid/--chroot options *** WARNING: you are running uWSGI as root !!! (use the --uid flag) *** Python version: 3.6.3 (default, Dec 12 2017, 16:37:57) [GCC 6.3.0 20170516] *** Python threads support is disabled. You can enable it with --enable-threads *** Python main interpreter initialized at 0x55f1741170f0 your server socket listen backlog is limited to 100 connections your mercy for graceful operations on workers is 60 seconds mapped 1237056 bytes (1208 KB) for 16 cores *** Operational MODE: preforking *** WSGI app 0 (mountpoint='') ready in 1 seconds on interpreter 0x55f1741170f0 pid: 11 (default app) 2023-08-10 00:33:46,865 INFO success: nginx entered RUNNING state, process has stayed up for > than 1 seconds (startsecs) 2023-08-10 00:33:46,866 INFO success: uwsgi entered RUNNING state, process has stayed up for > than 1 seconds (startsecs) *** uWSGI is running in multiple interpreter mode *** spawned uWSGI master process (pid: 11) spawned uWSGI worker 1 (pid: 14, cores: 1) spawned uWSGI worker 2 (pid: 15, cores: 1) [pid: 15|app: 0|req: 1/1] 10.224.0.4 () {32 vars in 442 bytes} [Thu Aug 10 00:49:10 2023] GET / => generated 950 bytes in 16 msecs (HTTP/1.1 200) 2 headers in 80 bytes (1 switches on core 0) 10.224.0.4 - - [10/Aug/2023:00:49:10 +0000] "GET / HTTP/1.1" 200 950 "-" "Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/51.0.2704.103 Safari/537.36" "-"
解决方案
错误里的exit status 137是容器被系统OOM Killer(内存不足终止器)杀死的典型标志,结合你“刚启动能exec,不久后失败”的现象,核心问题大概率是Pod内存资源不够。
排查步骤
- 检查Pod资源配置:打开该应用的Deployment YAML文件,看是否设置了
resources.limits.memory和resources.requests.memory。如果没配置或配置值太低,容器运行中内存占用超过节点分配的阈值,就会被强制终止。 - 查看节点内存状态:执行
kubectl describe node <节点名称>,检查节点的剩余内存,确认是不是节点本身内存不够用了。 - 查看Pod事件记录:执行
kubectl describe pod azure-vote-front-d99b7676c-hxdwt,在Events部分找有没有OOMKilled相关记录,这能直接坐实内存不足的问题。
解决方法
- 调整Pod资源限制:修改Deployment的YAML,提高内存的request和limit值,比如:
执行resources: requests: memory: "256Mi" cpu: "100m" limits: memory: "512Mi" cpu: "500m"kubectl apply -f <deployment文件>更新配置,Pod会重新调度并获得足够内存。 - 排查应用内存泄漏:如果调大资源后问题还存在,就需要检查应用本身有没有内存泄漏。可以用
kubectl top pod持续监控Pod的内存占用趋势,看是不是内存一直在增长。
内容的提问来源于stack exchange,提问作者davetayl
相关产品推荐
相关产品推荐

