uWSGI进程重启后挂起并占用100%CPU问题求助
问题概述
在新部署的服务器上使用uWSGI 2.0.28-1时,进程重启阶段出现挂起,CPU占用率拉满至100%。另一台配置相似的服务器未出现该问题,目前通过设置skip-atexit选项强制终止进程作为临时解决方案,寻求更优处理方式。
GDB调用栈分析
挂起进程的调用栈如下:
(gdb) bt #0 0x00007f8642c3a069 in SpinLock::lock (this=0x55a9104d9724) at src/spinlock.h:27 #1 std::lock_guard<SpinLock>::lock_guard (__m=..., this=<synthetic pointer>) at /opt/rh/devtoolset-10/root/usr/include/c++/10/bits/std_mutex.h:159 #2 HeapProfiler::HandleFree (ptr=0x7f863571fd30, this=0x55a9104d9720) at src/heap.h:116 #3 (anonymous namespace)::WrappedFree (ctx=0x7f8642d120b0 <(anonymous namespace)::g_base_allocators+80>, ptr=0x7f863571fd30) at src/malloc_patch.cc:95 #4 0x00007f864bd2781b in PyUnicode_FromFormatV () from target:/usr/lib/libpython3.9.so.1.0 #5 0x00007f864bd2ba90 in PyErr_Format () from target:/usr/lib/libpython3.9.so.1.0 #6 0x00007f864bd208db in _PyObject_GenericGetAttrWithDict () from target:/usr/lib/libpython3.9.so.1.0 #7 0x00007f864bd350e2 in PyImport_ImportModuleLevelObject () from target:/usr/lib/libpython3.9.so.1.0 #8 0x00007f864bddfa99 in PyImport_ImportModuleLevel () from target:/usr/lib/libpython3.9.so.1.0 #9 0x00007f864bd580fe in PyImport_Import () from target:/usr/lib/libpython3.9.so.1.0 #10 0x00007f864bddfa3d in PyImport_ImportModule () from target:/usr/lib/libpython3.9.so.1.0 #11 0x00007f864bfc0ada in get_uwsgi_pydict () from target:/usr/lib/uwsgi/python_plugin.so #12 0x00007f864bfbd460 in uwsgi_python_atexit () from target:/usr/lib/uwsgi/python_plugin.so #13 0x000055a90e2c3811 in uwsgi_plugins_atexit () #14 0x00007f864e3d8697 in __run_exit_handlers () from target:/usr/lib/libc.so.6 #15 0x00007f864e3d883e in exit () from target:/usr/lib/libc.so.6 #16 0x000055a90e276650 in uwsgi_exit () #17 0x000055a90e2c3867 in end_me () #18 0x000055a90e2c73aa in uwsgi_ignition () #19 0x000055a90e2cbdca in uwsgi_worker_run () #20 0x000055a90e2cc360 in uwsgi_run ()
从栈中可以明确:死锁发生在HeapProfiler的自旋锁(SpinLock) 阶段。进程退出时,uWSGI Python插件的uwsgi_python_atexit触发模块导入操作,期间调用内存释放函数WrappedFree,进而触发HeapProfiler的HandleFree,此时自旋锁因资源竞争陷入死循环,导致CPU占用拉满。
相关uWSGI配置
[uwsgi] master = true single-interpreter = true disable-logging = false log-4xx = true log-5xx = true master-fifo = /var/tmp/service.fifo heartbeat = 10 uid = tech guid = tech thunder-lock = true enable-threads = true vacuum = true workers = 160 plugin = python socket = /opt/service/service.sock stats = 127.0.0.1:4402 chmod-socket = 666 chdir = /opt/service venv = /opt/service/venv module = application.wsgi:application logto = /var/log/service/service_uwsgi.log pidfile2 = /var/tmp/service.pid procname = Service worker procname-master = Service master need-app = true hook-pre-app = exec:sudo chown tech:tech /var/log/service/*.log listen = 2048 harakiri = 120 buffer-size = 16384 max-request = 1000 max-worker-lifetime = 1200 max-worker-lifetime-delta = 60 reload-on-rss=1024
配置中enable-threads = true和高worker数量(160)可能加剧退出阶段的资源竞争,thunder-lock开启的全局锁也可能与自旋锁产生冲突。
解决方案建议
1. 升级uWSGI版本(最优方案)
uWSGI 2.0.28属于较旧版本,后续稳定版本(如2.0.20+的更新分支或最新LTS版本)已修复堆分析器在atexit阶段的死锁问题。升级到最新兼容版本可以从根源上解决该问题。
2. 禁用堆分析器
如果暂时无法升级,可在uWSGI配置中添加以下参数直接关闭堆分析功能,避免触发锁逻辑:
disable-heap-profiler = true
3. 调整运行时配置
- 降低worker数量:160个worker可能超出服务器CPU核心的合理负载范围,可根据CPU核心数调整(如核心数2或4),减少退出时的资源竞争。
- 添加退出缓冲时间:配置
reload-mercy = 30(单位秒),给进程退出阶段留足清理时间,避免因资源释放过快导致锁冲突。
4. 排查服务器环境差异
对比两台服务器的以下环境参数,找出可能的差异点:
- 内核版本与调度策略
- libc、Python 3.9的小版本
- GCC版本(调用栈中涉及devtoolset-10)
- 系统性能优化参数(如CPU频率、锁调度相关内核参数)
临时方案说明
skip-atexit确实可以跳过插件的退出清理逻辑,避免触发死锁,但可能会导致少量资源泄漏(不过进程退出后操作系统会自动回收这些资源,影响有限)。如果应用不依赖插件的优雅清理逻辑(如数据库连接池关闭),该方案可以短期使用,但不建议长期依赖。
内容的提问来源于stack exchange,提问作者Nikolay Andreev

