Airflow Helm Chart 8.7.0 Web Server重启:Worker退出并收到Signal 15
Airflow Web Server Pod 频繁重启问题排查与解决
问题现象
使用Airflow Helm Chart 8.7.0部署后,Web Server Pod出现频繁重启,日志显示Worker进程退出,主进程收到信号15(Received signal: 15)并关闭Gunicorn服务,关键日志如下:
[2023-07-10 05:42:16,665] {providers_manager.py:235} INFO - Optional provider feature disabled when importing 'airflow.providers.google.leveldb.hooks.leveldb.LevelDBHook' from 'apache-airflow-providers-google' package [2023-07-10 05:42:18 +0000] [113] [INFO] Starting gunicorn 20.1.0 [2023-07-10 05:42:19,182] {providers_manager.py:235} INFO - Optional provider feature disabled when importing 'airflow.providers.google.leveldb.hooks.leveldb.LevelDBHook' from 'apache-airflow-providers-google' package [2023-07-10 05:42:19 +0000] [113] [INFO] Listening at: http://0.0.0.0:8080 (113) [2023-07-10 05:42:19 +0000] [113] [INFO] Using worker: sync [2023-07-10 05:42:19 +0000] [222] [INFO] Booting worker with pid: 222 [2023-07-10 05:42:19 +0000] [223] [INFO] Booting worker with pid: 223 [2023-07-10 05:42:19 +0000] [224] [INFO] Booting worker with pid: 224 [2023-07-10 05:42:19 +0000] [225] [INFO] Booting worker with pid: 225 [2023-07-10 05:42:49 +0000] [223] [INFO] Worker exiting (pid: 223) [2023-07-10 05:42:49 +0000] [225] [INFO] Worker exiting (pid: 225) [2023-07-10 05:42:49 +0000] [224] [INFO] Worker exiting (pid: 224) [2023-07-10 05:42:49 +0000] [222] [INFO] Worker exiting (pid: 222) [2023-07-10 05:42:49 +0000] [113] [INFO] Handling signal: term ____________ _____________ ____ |__( )_________ __/__ /________ __ ____ /| |_ /__ ___/_ /_ __ /_ __ \_ | /| / / ___ ___ | / _ / _ __/ _ / / /_/ /_ |/ |/ / _/_/ |_/_/ /_/ /_/ /_/ \____/____/|__/ Running the Gunicorn Server with: Workers: 4 sync Host: 0.0.0.0:8080 Timeout: 120 Logfiles: - - Access Logformat: ================================================================= [2023-07-10 05:42:49,382] {webserver_command.py:432} INFO - Received signal: 15. Closing gunicorn. [2023-07-10 05:42:51 +0000] [113] [INFO] Shutting down: Master
排查方向与解决方案
1. 检查Kubernetes Pod资源限制与状态
- 执行
kubectl describe pod <web-server-pod-name>,查看Events字段:- 若存在
OOMKilled事件,说明Pod内存不足,需在Helm的values.yaml中调高webserver.resources的内存限制,例如:webserver: resources: requests: memory: "2Gi" cpu: "1" limits: memory: "4Gi" cpu: "2" - 若存在Readiness/Liveness探针失败事件,调整探针配置:
webserver: livenessProbe: initialDelaySeconds: 120 timeoutSeconds: 30 readinessProbe: initialDelaySeconds: 60 timeoutSeconds: 30
- 若存在
2. 调整Gunicorn配置
- 日志中Gunicorn超时时间为120秒,可延长超时时间避免Worker因处理长请求退出,在
values.yaml中添加:webserver: gunicornArgs: ["--timeout=300", "--workers=4"] - 若Worker数量过多导致资源竞争,可适当减少
workers数值,例如调整为2。
3. 排查数据库连接与Airflow内部状态
- 查看Web Server Pod的完整日志,确认是否存在元数据库连接失败的错误(如PostgreSQL/MySQL连接超时),若有需检查数据库服务是否正常,或调整Airflow数据库连接参数。
4. 验证Helm Chart版本兼容性
- Airflow Helm Chart 8.7.0对应的Airflow版本为2.5.x,若存在版本兼容性问题,可尝试升级到稳定版Helm Chart(如9.x系列对应Airflow 2.6+),或回退到之前可正常运行的Chart版本。
内容的提问来源于stack exchange,提问作者Indrajeet singh thakur
相关产品推荐
相关产品推荐

