在IIS 10上部署Flask应用时,子进程执行机器学习模型训练脚本无法等待完成的求助
首先,你的问题核心原因是IIS FastCGI的超时限制。开发环境用Flask自带的Werkzeug服务器没有请求超时限制,所以能完整执行8-9分钟的脚本;但部署到IIS后,FastCGI默认的请求超时(通常是90秒左右)远小于你的脚本执行时间,当超时后IIS会强制终止FastCGI进程,连带你的模型训练子进程也被杀死,这就是为什么日志显示脚本只执行了10%就终止。
下面是分步骤的解决方案:
一、紧急修复:调整IIS FastCGI超时设置
这是最快让现有代码工作的方法,先解决超时问题:
- 打开IIS管理器,找到你的站点,进入处理程序映射
- 找到对应Python Flask的FastCGI映射条目(通常是
Python FastCGI),右键选择编辑FastCGI设置 - 在弹出的窗口中,修改以下两个参数:
- 活动超时:设置为至少600秒(10分钟,比你的脚本执行时间长一点)
- 请求超时:同样设置为600秒以上
- 保存设置后,重启IIS站点和FastCGI进程
注意:即使调大超时,长时间占用HTTP连接也不是生产环境的最佳实践,用户需要一直等待响应,体验不好,所以更推荐下面的异步任务方案。
二、生产环境最佳实践:使用异步任务队列
把模型训练任务放到后台执行,客户端通过轮询获取状态,这样既不会占用HTTP连接,也能避免IIS超时问题。这里推荐用RQ(Redis Queue),比Celery更轻量,适合Python场景:
步骤1:安装依赖
pip install rq redis
确保你的Windows服务器上安装了Redis服务(可以下载Windows版本的Redis并启动)
步骤2:重构Flask代码
把模型训练逻辑封装成任务,放到队列中:
import os import time import glob import subprocess import pandas as pd from flask import Flask, request, jsonify from werkzeug.utils import secure_filename from datetime import datetime import logging from rq import Queue from redis import Redis # 初始化Redis连接和RQ队列 redis_conn = Redis(host='localhost', port=6379) q = Queue(connection=redis_conn) ALLOWED_EXTENSIONS = {'csv', 'xlsx'} app = Flask(__name__) app.config['UPLOAD_FOLDER'] = r"C:\inetpub\wwwroot\iAssist_IT_support\New_IT_support_datasets" currentDateTime = datetime.now() filenames = None # 日志配置保持不变... logger = logging.getLogger(__name__) app.logger.setLevel(logging.DEBUG) formatter = logging.Formatter('%(asctime)s:%(name)s:%(message)s') file_handler = logging.FileHandler('model-creation-status.log') file_handler.setFormatter(formatter) app.logger.addHandler(file_handler) def allowed_file(filename): return '.' in filename and filename.rsplit('.', 1)[1].lower() in ALLOWED_EXTENSIONS # 模型训练任务函数,放到后台执行 def train_model_task(filename): app.logger.debug("model script starts to run") # 注意路径用原始字符串或者转义 subprocess.run(r"python C:\.....\IT_support_chatbot-master\Python_files\main.py", shell=True) app.logger.debug("script ran successfully") # 可以把结果保存到Redis或者文件,供客户端查询 return f"Model created successfully for file {filename}" # 其他路由保持不变... @app.route('/file_upload') def home(): return jsonify("Hello, This is a file-upload API, To send the file, use http://13.213.81.139/file_upload/send_file") @app.route('/file_upload/status1', methods=['POST']) def upload_file(): app.logger.debug("/file_upload/status1 is execution") if 'file' not in request.files: app.logger.debug("No file part in the request") response = jsonify({'message': 'No file part in the request'}) response.status_code = 400 return response file = request.files['file'] if file.filename == '': app.logger.debug("No file selected for uploading") response = jsonify({'message': 'No file selected for uploading'}) response.status_code = 400 return response if file and allowed_file(file.filename): filename = secure_filename(file.filename) file.save(os.path.join(app.config['UPLOAD_FOLDER'], filename)) app.logger.debug("Spreadsheet received successfully") response = jsonify({'message': 'Spreadsheet uploaded successfully'}) response.status_code = 201 return response else: app.logger.debug("Allowed file types are csv or xlsx") response = jsonify({'message': 'Allowed file types are csv or xlsx'}) response.status_code = 400 return response @app.route('/file_upload/status2', methods=['POST']) def status1(): global filenames app.logger.debug("file_upload/status2 route is executed") if request.method == 'POST' and request.get_json(): filenames = request.get_json()['data'] app.logger.debug(filenames) folderpath = glob.glob(r'C:\inetpub\wwwroot\iAssist_IT_support\New_IT_support_datasets\*.csv') latest_file = max(folderpath, key=os.path.getctime) time.sleep(3) if filenames in latest_file: df1 = pd.read_csv(r"C:\inetpub\wwwroot\iAssist_IT_support\New_IT_support_datasets\" + filenames, names=["errors", "solutions"]) df1 = df1.drop(0) df2 = pd.read_csv(r"C:\inetpub\wwwroot\iAssist_IT_support\existing_tickets.csv", names=["errors", "solutions"]) combined_csv = pd.concat([df2, df1]) combined_csv.to_csv(r"C:\inetpub\wwwroot\iAssist_IT_support\new_tickets-chatdataset.csv", index=False, encoding='utf-8-sig') time.sleep(2) return jsonify('New data merged with existing datasets') @app.route('/file_upload/status3', methods=['POST']) def status2(): app.logger.debug("file_upload/status3 route is executed") if request.method == 'POST' and request.get_json(): message = request.get_json()['data'] app.logger.debug(message) return jsonify("New model training is in progress don't upload new file") @app.route('/file_upload/status4', methods=['POST']) def model_creation(): app.logger.debug("file_upload/status4 route is executed") if request.method == 'POST' and request.get_json(): message = request.get_json()['data'] app.logger.debug(message) app.logger.debug(currentDateTime) # 将任务加入队列,获取任务ID task = q.enqueue(train_model_task, filenames) # 返回任务ID给客户端,用于轮询状态 return jsonify({ "task_id": task.get_id(), "message": "Model training started in background. Use task_id to check status." }) # 添加查询任务状态的路由 @app.route('/file_upload/task_status/<task_id>', methods=['GET']) def get_task_status(task_id): task = q.fetch_job(task_id) if not task: return jsonify({"status": "error", "message": "Task not found"}), 404 if task.is_finished: return jsonify({ "status": "completed", "result": task.result }) elif task.is_failed: return jsonify({ "status": "failed", "error": str(task.exc_info) }) else: return jsonify({"status": "in_progress"}) if __name__ == "__main__": app.run()
步骤3:启动RQ Worker
在服务器上打开命令行,进入你的Flask应用目录,运行:
rq worker --with-scheduler
这个worker会监听队列,执行后台的模型训练任务。
步骤4:客户端流程调整
- 客户端调用
/file_upload/status4获取任务ID - 定期调用
/file_upload/task_status/<task_id>查询任务状态 - 当状态为
completed时,获取最终结果
三、备选方案:让子进程脱离父进程(不推荐)
如果暂时不想用任务队列,可以修改subprocess调用,让子进程独立于Flask的FastCGI进程,这样即使IIS终止FastCGI进程,子进程还能继续运行。但缺点是无法直接跟踪任务状态,需要自己实现状态记录:
修改model_run函数中的subprocess调用:
def model_run(): app.logger.debug("model script starts to run") # Windows下使用CREATE_NEW_PROCESS_GROUP让子进程独立 subprocess.Popen( r"python C:\.....\IT_support_chatbot-master\Python_files\main.py", shell=True, creationflags=subprocess.CREATE_NEW_PROCESS_GROUP, stdout=open('model_log.txt', 'w'), stderr=open('model_error.txt', 'w') ) app.logger.debug("script started in background")
然后你需要在模型脚本的末尾写入一个状态文件(比如model_completed.txt),客户端通过检查这个文件来判断任务是否完成。但这种方式不够可靠,适合临时测试。
内容的提问来源于stack exchange,提问作者Kavin Devarajan

