You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

生产环境虚拟环境中Celery守护进程异常:WorkerLostError求助

Hey there, let's tackle this Celery WorkerLostError (SIGKILL) issue you're facing in production. First off, a SIGKILL signal almost always means the system is killing your Celery worker processes—most commonly because they're consuming too much memory (triggering the OOM Killer), or hitting OS-enforced resource limits. Let's break down targeted fixes and production best practices tailored to your setup:


1. Diagnose & Fix the Immediate SIGKILL Issue

Check if OOM Killer is to blame

Run this command to search system logs for OOM Killer activity:

grep -i oom /var/log/syslog

If you see entries mentioning Celery workers, that confirms memory exhaustion is killing your processes.

Adjust Celery Worker Resource Limits

  • Reduce concurrency: Your current --concurrency=4 might be too high for your production server's available memory. Start with --concurrency=2 (or even 1 for small servers) in your CELERYD_OPTS to lower memory footprint.
  • Add soft time limits: You already have a hard time limit (--time-limit=300), but adding a soft limit gives workers a chance to gracefully exit before being force-killed:
    CELERYD_OPTS="--time-limit=300 --soft-time-limit=240 --concurrency=2"
    
  • Monitor task memory leaks: Some tasks might hold onto large objects or unclosed connections. Use celery inspect memory (while workers are running) to check per-worker memory usage, or add lightweight monitoring with the psutil library inside task code.

2. Fix Daemon Configuration Mistakes

Correct Path & Permission Issues

  • Fix CELERYD_CHDIR: Your path is missing a leading /—use the full absolute path:
    CELERYD_CHDIR="/home/myuser/project/myproj"
    
  • Fix directory permissions: You set /var/log/celery/ and /var/run/celery/ to root:root, but your Celery daemon runs as myuser—this will cause permission denied errors (which can indirectly trigger unexpected crashes). Run these commands to fix ownership:
    sudo chown -R myuser:myuser /var/log/celery/
    sudo chown -R myuser:myuser /var/run/celery/
    
  • Verify virtual environment CELERY_BIN: Ensure CELERY_BIN points directly to your venv's Celery executable, e.g.:
    CELERY_BIN="/home/myuser/project/myproj/venv/bin/celery"
    
    Using the system-wide Celery can cause dependency conflicts with your Django project.

3. Optimize Django & Celery Production Settings

Replace SQLite as Result Backend

SQLite is not designed for production Celery workloads—it uses single-file locking that can cause worker hangs and memory bloat. Switch to Redis or PostgreSQL instead:

# settings.py
# Option 1: Redis (recommended for speed)
CELERY_RESULT_BACKEND = 'redis://localhost:6379/0'

# Option 2: PostgreSQL (matches Django's DB if you use it)
# CELERY_RESULT_BACKEND = 'db+postgresql://db_user:db_pass@localhost/db_name'

Ensure Proper Celery-Django Integration

Double-check your Celery app initialization (in myproj/__init__.py) to make sure it loads Django settings correctly:

import os
from celery import Celery

# Set Django settings module for Celery
os.environ.setdefault('DJANGO_SETTINGS_MODULE', 'myproj.settings')

app = Celery('myproj')
# Load config from Django settings, prefix with CELERY_
app.config_from_object('django.conf:settings', namespace='CELERY')
# Auto-discover tasks from all Django apps
app.autodiscover_tasks()

4. Production Daemon Best Practices

The old init.d script is less reliable than systemd. Create a /etc/systemd/system/celery.service file with these contents:

[Unit]
Description=Celery Worker Service
After=network.target rabbitmq-server.target

[Service]
Type=forking
User=myuser
Group=myuser
WorkingDirectory=/home/myuser/project/myproj
# Point to your virtual environment's PATH
Environment="PATH=/home/myuser/project/myproj/venv/bin"
# Start command with optimized options
ExecStart=/home/myuser/project/myproj/venv/bin/celery multi start worker1 \
  -A myproj \
  --time-limit=300 --soft-time-limit=240 \
  --concurrency=2 \
  --logfile=/var/log/celery/%n%I.log \
  --pidfile=/var/run/celery/%n.pid
# Graceful stop
ExecStop=/home/myuser/project/myproj/venv/bin/celery multi stopwait worker1 \
  --pidfile=/var/run/celery/%n.pid
# Restart with fresh config
ExecReload=/home/myuser/project/myproj/venv/bin/celery multi restart worker1 \
  -A myproj \
  --time-limit=300 --soft-time-limit=240 \
  --concurrency=2 \
  --logfile=/var/log/celery/%n%I.log \
  --pidfile=/var/run/celery/%n.pid

[Install]
WantedBy=multi-user.target

Then enable and start the service:

sudo systemctl daemon-reload
sudo systemctl enable celery.service
sudo systemctl start celery.service

Monitor Worker Health

  • Use celery status or celery inspect active to verify workers are running and connected to the broker.
  • Regularly check Celery logs (/var/log/celery/worker1.log) for task-specific errors that might cause crashes.

内容的提问来源于stack exchange,提问作者Krish V

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:44:12