Django执行迁移时如何跳过apps.ready函数中的爬虫启动代码
解决方案
首选方案:校验manage.py执行命令(匹配你推荐的思路)
你只需要在apps.ready函数中添加命令行参数校验逻辑即可,完全不需要修改Django源码,也不会受版本升级影响,修改后的ready函数代码如下:
import sys def ready(self): set_default_database_items() if environment == "prod": # 仅当执行runserver命令时启动爬虫线程 if len(sys.argv) > 1 and sys.argv[1] == "runserver": from .threading.scraping import TimerScrape from .threading.keep_alive import KeepAliveThread TimerScrape() KeepAliveThread(1) KeepAliveThread(2)
Django执行任何manage.py子命令时,命令名称都会作为sys.argv的第二个参数存在,上面的逻辑会直接过滤掉makemigrations、migrate、shell等所有非runserver的命令,从根源上避免爬虫线程误启动。
如果是生产环境用uWSGI、Gunicorn等WSGI服务器启动而非runserver,可以再加一层环境变量校验,比如启动服务时主动设置自定义环境变量:
# 生产环境启动服务时执行 export RUN_CRAWLER=1 gunicorn your_project.wsgi:application
对应修改ready逻辑:
import sys, os def ready(self): set_default_database_items() if environment == "prod": # 本地开发runserver,或生产环境设置了启动变量时才启动爬虫 run_crawler = (len(sys.argv) > 1 and sys.argv[1] == "runserver") or os.getenv("RUN_CRAWLER") == "1" if run_crawler: from .threading.scraping import TimerScrape from .threading.keep_alive import KeepAliveThread TimerScrape() KeepAliveThread(1) KeepAliveThread(2)
该方案兼容开发、生产全场景,不需要修改原有爬虫业务逻辑。
替代方案:使用Django启动信号
如果不想在apps.ready里加命令判断,也可以监听Django的wsgi_application_started信号,这个信号只会在WSGI应用初始化完成后触发,不会在执行migrate等管理命令时触发,实现代码如下:
首先在你的app目录下新增signals.py文件:
from django.core.signals import wsgi_application_started from django.dispatch import receiver @receiver(wsgi_application_started) def start_crawler_threads(sender, **kwargs): if environment == "prod": from .threading.scraping import TimerScrape from .threading.keep_alive import KeepAliveThread TimerScrape() KeepAliveThread(1) KeepAliveThread(2)
然后在apps.py的ready函数里注册信号即可:
def ready(self): set_default_database_items() # 注册启动信号 import your_app_name.signals
该方案的优势是完全不需要判断命令类型,信号本身的触发时机就匹配你的需求。
优化建议
你现有代码里的60秒启动延迟可以保留,作为服务初始化缓冲时间,和上面的判断逻辑叠加使用稳妥性更高。
内容的提问来源于stack exchange,提问作者Jimmy Fort
相关产品推荐
相关产品推荐

