Celery任务中face_recognition触发CUDA初始化错误问题
问题:Celery执行face_recognition任务触发CUDA初始化错误(错误码3)
环境信息
- face_recognition 1.3.0
- face-recognition-models 0.3.0
- Python 3.8.16
- Ubuntu 22.04.2 LTS
- Django 4.1 + Celery + Redis(虚拟环境运行)
异常现象
- 直接调用face_recognition代码、直接执行
find_faces_task任务(无论是否启用CUDA),均可正常识别5张人脸; - 通过Celery的
delay方法执行时:- 启用CUDA:在
face_recognition.face_locations环节触发cudaGetDevice初始化错误(错误码3); - 不启用CUDA:在
face_recognition.face_encodings环节触发相同的CUDA初始化错误。
- 启用CUDA:在
任务代码
@app.task(bind=True, base=utils.BaseTaskWithRetry) def find_faces_task(self, document_id, use_cuda=settings.USE_CUDA): logger.debug("find_faces_task in tasks START") try: temp_file = None from memorabilia.models import TaskStatus, Document args = "document_id=%s, use_cuda=%s" % (document_id, use_cuda) ts = TaskStatus(document_id_id=document_id, task_id=self.request.id, task_name='find_faces_task', task_args=args, task_status=TaskStatus.PENDING) ts.save() import time time_start = time.time() from memorabilia.models import Document from biometric_identification.models import Face if len(Face.objects.filter(document_id=document_id)) != 0: # This document has already been scanned, so need to remove it and rescan # Have to manually delete each object per django docs to insure the # model delete method is run to update the metadata. logger.debug("Document %s has already been scanned" % document_id) faces = Face.objects.filter(document_id=document_id) for face in faces: face.delete() logger.debug("Deleted face=%s" % face.tag_value.value) document = Document.objects.get(document_id=document_id) image_file = document.get_default_image_file(settings.DEFAULT_DISPLAY_IMAGE) image_path = image_file.path logger.debug("document_id=%s, image_path=%s" % (document_id, image_path)) time_start_looking = time.time() temp_file = open(image_path, 'rb') temp_image = Image.open(temp_file) logger.debug("temp_image.mode=%s" % temp_image.mode) width, height = temp_image.size image = face_recognition.load_image_file(temp_file) # Get the coordinates of each face if use_cuda: # With CUDA installed logger.debug("Using CUDA for face recognition") face_locations = face_recognition.face_locations(image, model="cnn", number_of_times_to_upsample=0) else: # without CUDA installed logger.debug("NOT using CUDA for face recognition") #face_locations = face_recognition.face_locations(image, number_of_times_to_upsample=2) face_locations = face_recognition.face_locations(image, model="hog", number_of_times_to_upsample=2) if len(face_locations) == 0: ts.task_status = TaskStatus.WARNING ts.comment = "Found %s faces" % len(face_locations) else: time_find_faces = time.time() # Get the face encodings for each face in the picture face_encodings = face_recognition.face_encodings(image, known_face_locations=face_locations) logger.debug("Found %s face locations and %s encodings" % (len(face_locations), len(face_encodings))) time_face_encodings = time.time() # Save the faces found in the database for location, encoding in zip(face_locations, face_encodings): # Create the new Face object and load in the document, encoding, and location of a face found # Locations seem to be of the form (y,x) from memorabilia.models import MetaData, MetaDataValue tag_type_people = MetaDataValue.objects.filter(metadata_id=MetaData.objects.filter(name='Tag_types')[0].metadata_id, value='People')[0] tag_value_unknown = MetaDataValue.objects.filter(metadata_id=MetaData.objects.filter(name='Unknown')[0].metadata_id, value='Unknown')[0] new_face = Face(document=document, face_encoding=numpy_to_json(encoding), face_location=location, image_size={'width': width, "height":height}, tag_type=tag_type_people, tag_value=tag_value_unknown) # save the newly found Face object new_face.save() logger.debug("Saved new_face %s" % new_face.face_file) time_end = time.time() logger.debug("total time = {}".format(time_end - time_start)) logger.debug("time to find faces = {}".format(time_find_faces - time_start_looking)) logger.debug("time to find encodings = {}".format(time_face_encodings - time_find_faces)) ts.task_status = TaskStatus.SUCCESS ts.comment = "Found %s faces" % len(face_encodings) return document_id except Exception as e: logger.exception("Hit an exception in find_faces_task %s" % str(e)) ts.task_status = TaskStatus.ERROR ts.comment = "An exception while finding faces: %s" % repr(e) ts.save(update_fields=['task_status', 'comment']) raise Exception("Hit an exception in find_faces_task for document_id=%s and use_cuda=%s" %(document_id, use_cuda)) from e finally: logger.debug("Finally clause in find-faces_task") logger.debug("temp_file=%s" % temp_file) logger.debug("temp_image=%s" % temp_image) if temp_file: temp_file.close() logger.debug("closed temp_file=%s" % temp_file) if temp_image: temp_image.close() logger.debug("closed temp_image=%s" % temp_image) ts.save(update_fields=['task_status', 'comment']) logger.debug("find_faces_task END")
原因分析
- Celery worker环境变量缺失:直接运行任务时继承了当前shell的CUDA环境变量,但Celery worker启动时未加载这些变量,导致dlib无法找到CUDA库路径;
- 多进程CUDA初始化冲突:Celery默认使用多进程池,dlib的CUDA初始化在多进程共享上下文时会出现资源竞争,触发初始化错误;
- GPU访问权限问题:Celery worker运行的用户(如服务用户)没有GPU设备的访问权限,即使CUDA安装正常也无法调用;
- face_recognition底层逻辑:即使指定使用HOG模型,face_recognition的
face_encodings仍可能尝试初始化CUDA上下文,导致错误。
解决方案
1. 确保Celery worker加载CUDA环境变量
- 启动worker前显式导出CUDA变量:
export CUDA_HOME=/usr/local/cuda export LD_LIBRARY_PATH=$CUDA_HOME/lib64:$LD_LIBRARY_PATH celery -A your_django_project worker -l info - 若用systemd管理Celery服务,在
.service文件的[Service]段添加:Environment="CUDA_HOME=/usr/local/cuda" Environment="LD_LIBRARY_PATH=/usr/local/cuda/lib64:$LD_LIBRARY_PATH"
2. 强制禁用CUDA(快速验证)
- 在任务代码开头添加环境变量设置,强制dlib使用CPU:
import os os.environ['CUDA_VISIBLE_DEVICES'] = '-1' - 或启动worker时添加该变量:
CUDA_VISIBLE_DEVICES=-1 celery -A your_django_project worker -l info
3. 调整Celery worker进程模式
- 用单进程模式启动worker,避免多进程CUDA冲突:
celery -A your_django_project worker -l info --pool=solo - 若需要多进程,可尝试降低并发数(如
--concurrency=2),并确保每个进程独立初始化CUDA上下文。
4. 验证worker用户的GPU权限
- 切换到Celery运行的用户,执行
nvidia-smi和官方CUDA测试程序,确认能正常访问GPU; - 检查GPU设备文件权限:
若权限不足,可将worker用户添加到ls -l /dev/nvidia*video组:sudo usermod -aG video celery_user
内容的提问来源于stack exchange,提问作者user1045680
相关产品推荐
相关产品推荐

