You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在IIS 10上部署Flask应用时,子进程执行机器学习模型训练脚本无法等待完成的求助

问题分析与解决方案

首先,你的问题核心原因是IIS FastCGI的超时限制。开发环境用Flask自带的Werkzeug服务器没有请求超时限制,所以能完整执行8-9分钟的脚本;但部署到IIS后,FastCGI默认的请求超时(通常是90秒左右)远小于你的脚本执行时间,当超时后IIS会强制终止FastCGI进程,连带你的模型训练子进程也被杀死,这就是为什么日志显示脚本只执行了10%就终止。

下面是分步骤的解决方案:

一、紧急修复:调整IIS FastCGI超时设置

这是最快让现有代码工作的方法,先解决超时问题:

  1. 打开IIS管理器,找到你的站点,进入处理程序映射
  2. 找到对应Python Flask的FastCGI映射条目(通常是Python FastCGI),右键选择编辑FastCGI设置
  3. 在弹出的窗口中,修改以下两个参数:
    • 活动超时:设置为至少600秒(10分钟,比你的脚本执行时间长一点)
    • 请求超时:同样设置为600秒以上
  4. 保存设置后,重启IIS站点和FastCGI进程

注意:即使调大超时,长时间占用HTTP连接也不是生产环境的最佳实践,用户需要一直等待响应,体验不好,所以更推荐下面的异步任务方案。

二、生产环境最佳实践:使用异步任务队列

把模型训练任务放到后台执行,客户端通过轮询获取状态,这样既不会占用HTTP连接,也能避免IIS超时问题。这里推荐用RQ(Redis Queue),比Celery更轻量,适合Python场景:

步骤1:安装依赖

pip install rq redis

确保你的Windows服务器上安装了Redis服务(可以下载Windows版本的Redis并启动)

步骤2:重构Flask代码

把模型训练逻辑封装成任务,放到队列中:

import os
import time
import glob
import subprocess
import pandas as pd
from flask import Flask, request, jsonify
from werkzeug.utils import secure_filename
from datetime import datetime
import logging
from rq import Queue
from redis import Redis

# 初始化Redis连接和RQ队列
redis_conn = Redis(host='localhost', port=6379)
q = Queue(connection=redis_conn)

ALLOWED_EXTENSIONS = {'csv', 'xlsx'}

app = Flask(__name__)
app.config['UPLOAD_FOLDER'] = r"C:\inetpub\wwwroot\iAssist_IT_support\New_IT_support_datasets"
currentDateTime = datetime.now()
filenames = None

# 日志配置保持不变...
logger = logging.getLogger(__name__)
app.logger.setLevel(logging.DEBUG)
formatter = logging.Formatter('%(asctime)s:%(name)s:%(message)s')
file_handler = logging.FileHandler('model-creation-status.log')
file_handler.setFormatter(formatter)
app.logger.addHandler(file_handler)

def allowed_file(filename):
    return '.' in filename and filename.rsplit('.', 1)[1].lower() in ALLOWED_EXTENSIONS

# 模型训练任务函数,放到后台执行
def train_model_task(filename):
    app.logger.debug("model script starts to run")
    # 注意路径用原始字符串或者转义
    subprocess.run(r"python C:\.....\IT_support_chatbot-master\Python_files\main.py", shell=True)
    app.logger.debug("script ran successfully")
    # 可以把结果保存到Redis或者文件,供客户端查询
    return f"Model created successfully for file {filename}"

# 其他路由保持不变...
@app.route('/file_upload')
def home():
    return jsonify("Hello, This is a file-upload API, To send the file, use http://13.213.81.139/file_upload/send_file")

@app.route('/file_upload/status1', methods=['POST'])
def upload_file():
    app.logger.debug("/file_upload/status1 is execution")
    if 'file' not in request.files:
        app.logger.debug("No file part in the request")
        response = jsonify({'message': 'No file part in the request'})
        response.status_code = 400
        return response
    file = request.files['file']
    if file.filename == '':
        app.logger.debug("No file selected for uploading")
        response = jsonify({'message': 'No file selected for uploading'})
        response.status_code = 400
        return response
    if file and allowed_file(file.filename):
        filename = secure_filename(file.filename)
        file.save(os.path.join(app.config['UPLOAD_FOLDER'], filename))
        app.logger.debug("Spreadsheet received successfully")
        response = jsonify({'message': 'Spreadsheet uploaded successfully'})
        response.status_code = 201
        return response
    else:
        app.logger.debug("Allowed file types are csv or xlsx")
        response = jsonify({'message': 'Allowed file types are csv or xlsx'})
        response.status_code = 400
        return response

@app.route('/file_upload/status2', methods=['POST'])
def status1():
    global filenames
    app.logger.debug("file_upload/status2 route is executed")
    if request.method == 'POST' and request.get_json():
        filenames = request.get_json()['data']
        app.logger.debug(filenames)
        folderpath = glob.glob(r'C:\inetpub\wwwroot\iAssist_IT_support\New_IT_support_datasets\*.csv')
        latest_file = max(folderpath, key=os.path.getctime)
        time.sleep(3)
        if filenames in latest_file:
            df1 = pd.read_csv(r"C:\inetpub\wwwroot\iAssist_IT_support\New_IT_support_datasets\" + filenames, names=["errors", "solutions"])
            df1 = df1.drop(0)
            df2 = pd.read_csv(r"C:\inetpub\wwwroot\iAssist_IT_support\existing_tickets.csv", names=["errors", "solutions"])
            combined_csv = pd.concat([df2, df1])
            combined_csv.to_csv(r"C:\inetpub\wwwroot\iAssist_IT_support\new_tickets-chatdataset.csv", index=False, encoding='utf-8-sig')
            time.sleep(2)
            return jsonify('New data merged with existing datasets')

@app.route('/file_upload/status3', methods=['POST'])
def status2():
    app.logger.debug("file_upload/status3 route is executed")
    if request.method == 'POST' and request.get_json():
        message = request.get_json()['data']
        app.logger.debug(message)
        return jsonify("New model training is in progress don't upload new file")

@app.route('/file_upload/status4', methods=['POST'])
def model_creation():
    app.logger.debug("file_upload/status4 route is executed")
    if request.method == 'POST' and request.get_json():
        message = request.get_json()['data']
        app.logger.debug(message)
        app.logger.debug(currentDateTime)
        
        # 将任务加入队列,获取任务ID
        task = q.enqueue(train_model_task, filenames)
        
        # 返回任务ID给客户端,用于轮询状态
        return jsonify({
            "task_id": task.get_id(),
            "message": "Model training started in background. Use task_id to check status."
        })

# 添加查询任务状态的路由
@app.route('/file_upload/task_status/<task_id>', methods=['GET'])
def get_task_status(task_id):
    task = q.fetch_job(task_id)
    if not task:
        return jsonify({"status": "error", "message": "Task not found"}), 404
    
    if task.is_finished:
        return jsonify({
            "status": "completed",
            "result": task.result
        })
    elif task.is_failed:
        return jsonify({
            "status": "failed",
            "error": str(task.exc_info)
        })
    else:
        return jsonify({"status": "in_progress"})

if __name__ == "__main__":
    app.run()

步骤3:启动RQ Worker

在服务器上打开命令行,进入你的Flask应用目录,运行:

rq worker --with-scheduler

这个worker会监听队列,执行后台的模型训练任务。

步骤4:客户端流程调整

  1. 客户端调用/file_upload/status4获取任务ID
  2. 定期调用/file_upload/task_status/<task_id>查询任务状态
  3. 当状态为completed时,获取最终结果

三、备选方案:让子进程脱离父进程(不推荐)

如果暂时不想用任务队列,可以修改subprocess调用,让子进程独立于Flask的FastCGI进程,这样即使IIS终止FastCGI进程,子进程还能继续运行。但缺点是无法直接跟踪任务状态,需要自己实现状态记录:

修改model_run函数中的subprocess调用:

def model_run():
    app.logger.debug("model script starts to run")
    # Windows下使用CREATE_NEW_PROCESS_GROUP让子进程独立
    subprocess.Popen(
        r"python C:\.....\IT_support_chatbot-master\Python_files\main.py",
        shell=True,
        creationflags=subprocess.CREATE_NEW_PROCESS_GROUP,
        stdout=open('model_log.txt', 'w'),
        stderr=open('model_error.txt', 'w')
    )
    app.logger.debug("script started in background")

然后你需要在模型脚本的末尾写入一个状态文件(比如model_completed.txt),客户端通过检查这个文件来判断任务是否完成。但这种方式不够可靠,适合临时测试。


内容的提问来源于stack exchange,提问作者Kavin Devarajan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 13:32:30