Django集成Celery执行Whisper转录时Worker异常退出(Signal 11)求助
问题:Django+Celery执行Whisper转录任务触发SIGSEGV错误
开发带时间戳音频转录的Django应用,用户点击按钮触发服务器执行转录脚本。最初在视图直接运行transcribe.py会导致Django服务器终止,改用Celery+Redis后台执行后,脚本在Django Shell中正常,但视图触发时Celery Worker抛出signal 11 (SIGSEGV)异常,尝试音频分片、调整Worker内存池无效。
错误日志
[tasks] . transcribeApp.tasks.run_transcription [2024-11-25 03:26:04,500: INFO/MainProcess] Connected to redis://localhost:6379/0 [2024-11-25 03:26:04,514: INFO/MainProcess] mingle: searching for neighbors [2024-11-25 03:26:05,520: INFO/MainProcess] mingle: all alone [2024-11-25 03:26:05,544: INFO/MainProcess] celery@user.local ready. [2024-11-25 03:26:16,253: INFO/MainProcess] Task searchApp.tasks.run_transcription[c684bdfa-ec21-4b4e-9542-0ca1f7729682] received [2024-11-25 03:26:16,255: INFO/ForkPoolWorker-15] Starting transcription process. [2024-11-25 03:26:16,509: WARNING/ForkPoolWorker-15] /Users/user/Desktop/project/django_app/django_venv/lib/python3.12/site-packages/whisper/__init__.py:150: FutureWarning: You are using `torch.load` with `weights_only=False` (the current default value), which uses the default pickle module implicitly. It is possible to construct malicious pickle data which will execute arbitrary code during unpickling (See https://github.com/pytorch/pytorch/blob/main/SECURITY.md#untrusted-models for more details). In a future release, the default value for `weights_only` will be flipped to `True`. This limits the functions that could be executed during unpickling. Arbitrary objects will no longer be allowed to be loaded via this mode unless they are explicitly allowlisted by the user via `torch.serialization.add_safe_globals`. We recommend you start setting `weights_only=True` for any use case where you don't have full control of the loaded file. Please open an issue on GitHub for any issues related to this experimental feature. checkpoint = torch.load(fp, map_location=device) [2024-11-25 03:26:16,670: ERROR/MainProcess] Process 'ForkPoolWorker-15' pid:38956 exited with 'signal 11 (SIGSEGV)' [2024-11-25 03:26:16,683: ERROR/MainProcess] Task handler raised error: WorkerLostError('Worker exited prematurely: signal 11 (SIGSEGV) Job: 0.') Traceback (most recent call last): File "/Users/user/Desktop/project/django_app/django_venv/lib/python3.12/site-packages/billiard/pool.py", line 1265, in mark_as_worker_lost raise WorkerLostError( billiard.einfo.ExceptionWithTraceback: """ Traceback (most recent call last): File "/Users/user/Desktop/project/django_app/django_venv/lib/python3.12/site-packages/billiard/pool.py", line 1265, in mark_as_worker_lost raise WorkerLostError( billiard.exceptions.WorkerLostError: Worker exited prematurely: signal 11 (SIGSEGV) Job: 0. """
相关代码实现
Views.py
# Views.py from . import tasks from django.shortcuts import render from django.http import HttpResponse, JsonResponse def trainVideos(request): try: tasks.run_transcription.delay() return JsonResponse({"status": "success", "message": "Transcription has started check back later."}) except Exception as e: JsonResponse({"status": "error", "message": str(e)})
transcribe.py
# transcribe.py import whisper_timestamped as whisper import os def transcribeTexts(model_id, filePath): result = [] fileNames = os.listdir() model = whisper.load_model(model_id) for files in fileNames: audioPath = filePath + "/" + files audio = whisper.load_audio(audioPath) result.append(model.transcribe(audio, language="en")) return result model_id = "tiny" audioFilePath = "path/to/audio" transcribeTexts(model_id, audioFilePath)
Celery配置
celery.py
from __future__ import absolute_import, unicode_literals import os from celery import Celery os.environ.setdefault('DJANGO_SETTINGS_MODULE', 'main_app.settings') app = Celery('main_app') app.config_from_object('django.conf:settings', namespace='CELERY') app.autodiscover_tasks() def debug_tasks(self): print(f"Request: {self.request!r}")
tasks.py
from __future__ import absolute_import, unicode_literals from . import transcribe from celery import shared_task @shared_task def run_transcription(): transcribe.transcribe() return "Transcription Completed..."
settings.py相关配置
CELERY_BROKER_URL = 'redis://localhost:6379/0' CELERY_BROKER_CONNECTION_RETRY_ON_STARTUP = True
环境依赖版本
Package Version -------------------- ----------- amqp 5.3.1 asgiref 3.8.1 billiard 4.2.1 celery 5.4.0 certifi 2024.8.30 charset-normalizer 3.3.2 click 8.1.7 click-didyoumean 0.3.1 click-plugins 1.1.1 click-repl 0.3.0 Cython 3.0.11 Django 5.1.2 django-widget-tweaks 1.5.0 dtw-python 1.5.3 faiss-cpu 1.9.0 ffmpeg 1.4 filelock 3.16.1 fsspec 2024.9.0 huggingface-hub 0.25.2 idna 3.10 Jinja2 3.1.4 kombu 5.4.2 lfs 0.2 llvmlite 0.43.0 MarkupSafe 3.0.1 more-itertools 10.5.0 mpmath 1.3.0 msgpack 1.1.0 networkx 3.3 numba 0.60.0 numpy 2.0.2 packaging 24.1 panda 0.3.1 pillow 10.4.0 pip 24.3.1 prompt_toolkit 3.0.48 pydub 0.25.1 python-dateutil 2.9.0.post0 PyYAML 6.0.2 redis 5.2.0 regex 2024.9.11 requests 2.32.3 safetensors 0.4.5 scipy 1.14.1 semantic-version 2.10.0 setuptools 75.1.0 setuptools-rust 1.10.2 six 1.16.0 sqlparse 0.5.1 sympy 1.13.3 tiktoken 0.8.0 tokenizers 0.20.1 torch 2.4.1 torchaudio 2.4.1 torchvision 0.19.1 tqdm 4.66.5 transformers 4.45.2 txtai 7.4.0 typing_extensions 4.12.2 tzdata 2024.2 urllib3 2.2.3 vine 5.1.0 wcwidth 0.2.13 whisper-timestamped 1.15.4
原因排查与解决方案
核心原因分析
SIGSEGV是内存访问错误,结合日志中出错时机(加载Whisper模型后),大概率是以下问题:
- Torch与Celery Fork模式冲突:Celery默认用ForkPoolWorker,Fork进程会继承父进程的Torch内存上下文,导致内存访问冲突。
- numpy 2.0兼容性问题:numpy 2.0对旧版科学计算库兼容性差,whisper-timestamped依赖的底层库未适配,引发内存错误。
- 路径与调用逻辑错误:
os.listdir()默认读取当前目录而非目标音频目录,任务调用函数名不匹配,可能触发异常连锁反应。
具体解决方案
1. 切换Celery Worker为Spawn模式
Fork模式会继承父进程内存空间,Spawn模式创建全新进程,避免Torch上下文冲突。修改Celery启动命令:
celery -A main_app worker --loglevel=info --pool=solo
或在settings.py中配置全局生效:
CELERY_WORKER_POOL = 'solo'
生产环境可替换为gevent或eventlet,需提前安装对应依赖。
2. 降级numpy到1.x版本
numpy 2.0的API变动可能导致whisper-timestamped底层崩溃,执行降级:
pip install numpy==1.26.4
3. 修复转录脚本逻辑
修改transcribe.py,修正路径读取错误并确保模型加载安全:
import whisper_timestamped as whisper import os def transcribeTexts(model_id, filePath): result = [] # 读取目标音频目录下的文件,而非当前目录 fileNames = os.listdir(filePath) # 在函数内部加载模型,避免进程共享内存冲突 model = whisper.load_model(model_id) for filename in fileNames: audioPath = os.path.join(filePath, filename) # 过滤非文件类型(如子目录) if os.path.isfile(audioPath): audio = whisper.load_audio(audioPath) result.append(model.transcribe(audio, language="en")) return result
同时修复tasks.py中的调用错误:
from __future__ import absolute_import, unicode_literals from . import transcribe from celery import shared_task @shared_task def run_transcription(): model_id = "tiny" audioFilePath = "path/to/audio" transcribe.transcribeTexts(model_id, audioFilePath) return "Transcription Completed..."
4. 优化Torch权重加载(可选)
消除日志警告并提升安全性,在模型加载前添加:
import torch # 设置weights_only=True避免pickle安全风险 torch.load.__defaults__ = (None, False, True, 'cpu')
内容的提问来源于stack exchange,提问作者Mendax
相关产品推荐
相关产品推荐

