Flask应用清理数据库批次时后台线程意外静默终止问题排查
Flask后台清理线程静默终止问题排查
问题描述
开发Flask应用时,设置了后台线程定期清理数据库中的过期批次:线程在收到用户请求时启动,10分钟无活动则停止。但线程会意外静默终止,无任何错误提示,日志仅停留在"Starting background task",无法查明终止原因。
代码
import threading from time import sleep from datetime import datetime, timedelta from flask import Flask, request from sqlalchemy import text app = Flask(__name__) last_request_time = None background_task_running = False task_lock = threading.Lock() # 更新最后请求时间 def update_last_request_time(): global last_request_time last_request_time = datetime.utcnow() # 启动后台任务(确保单例运行) def start_background_task(): global background_task_running with task_lock: if not background_task_running: try: print("Starting background task") background_thread = threading.Thread(target=check_for_inactive_batches) background_thread.daemon = True # 随应用停止而终止 background_thread.start() background_task_running = True except Exception as e: print(f"Error Starting background task: {e}") # 停止后台任务 def stop_background_task(): global background_task_running with task_lock: if background_task_running: print("Stopping background task") background_task_running = False # 后台清理任务逻辑 def check_for_inactive_batches(): global background_task_running engine = db_connection() # 数据库连接初始化 while background_task_running: if last_request_time and (datetime.utcnow() - last_request_time).total_seconds() > 600: stop_background_task() print("No requests for 10 minutes, stopping background task.") break current_time = datetime.utcnow() with engine.connect() as conn: try: delete_query = text(""" DELETE FROM TB_BATCHES WHERE LastHeartbeat < :threshold_time AND Status = 'scanning' OR Status = 'uploading' """) threshold_time = current_time - timedelta(seconds=70) conn.execute(delete_query, {'threshold_time': threshold_time}) conn.commit() print(f"Checked and deleted inactive batches at {current_time}") except Exception as e: print(f"Error deleting inactive batches: {e}") sleep(60) # 每分钟检查一次 # 请求前钩子:追踪用户活动并启动后台任务 @app.before_request def track_user_activity(): if request.endpoint == 'get_batches': return # 跳过该路由的清理逻辑 update_last_request_time() start_background_task()
环境
- Python版本: 3.11.3
- Flask版本: 2.3.2
- WSGI服务器: waitress
已尝试的解决方法
- 日志排查:添加
print语句,但仅输出"Starting background task"后无后续日志 - 守护线程设置:将线程设为守护线程,未解决问题
- SQLAlchemy错误捕获:在数据库查询块添加try/except,未捕获到错误
- 资源检查:问题发生时无明显内存/CPU使用率飙升
预期行为
后台线程持续运行,10分钟无请求后自动停止;收到新请求时重新启动,不会中途静默终止。
实际行为
线程无任何错误或日志输出就停止,最后一条日志为"Starting background task",数据库清理查询未执行。
问题根源及修复方案
数据库连接未捕获异常
check_for_inactive_batches中engine = db_connection()无异常捕获,一旦数据库连接失败(比如配置错误、网络中断),线程会直接崩溃且无日志。
修复:def check_for_inactive_batches(): global background_task_running try: engine = db_connection() except Exception as e: print(f"Database connection failed: {e}") with task_lock: background_task_running = False return # 原有循环逻辑...SQL查询逻辑错误
DELETE语句中AND与OR优先级问题:原语句会误删所有Status='uploading'的批次(无论心跳时间),若触发数据库约束异常,可能导致线程崩溃。
修复:给OR条件加括号:delete_query = text(""" DELETE FROM TB_BATCHES WHERE LastHeartbeat < :threshold_time AND (Status = 'scanning' OR Status = 'uploading') """)全局变量线程不安全
last_request_time的读写未加锁,多请求并发时可能导致线程读取到过时值,甚至引发数据不一致。
修复:添加锁保护全局变量:request_lock = threading.Lock() def update_last_request_time(): global last_request_time with request_lock: last_request_time = datetime.utcnow() # 在check_for_inactive_batches中读取时加锁 with request_lock: if last_request_time and (datetime.utcnow() - last_request_time).total_seconds() > 600: stop_background_task() print("No requests for 10 minutes, stopping background task.") break线程状态更新不完整
若线程启动后因异常退出,需手动将background_task_running设为False,否则下次请求无法重新启动线程。替换print为Flask日志
WSGI环境中print输出可能无法被正确捕获,改用Flask内置日志系统:import logging app.logger.setLevel(logging.INFO) # 替换所有print语句 app.logger.info("Starting background task") app.logger.error(f"Database connection failed: {e}")
内容的提问来源于stack exchange,提问作者Bilal Zafar
相关产品推荐
相关产品推荐

