You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure部署Python3.9+Flask Excel处理应用遇HTTP503错误求助

Azure App Service部署Flask Excel处理应用HTTP 503错误排查与修复

问题背景

基于Python 3.9 + Flask开发的Excel文件处理应用,本地运行完全正常,但部署至Azure App Service B1计划(1vCPU、1.75GB内存、100ACU)后出现HTTP 503错误。已尝试内存中处理文件(不落地存储),本地可行但Azure端仍失败,怀疑问题与Excel处理逻辑、权限配置或服务启动配置有关。

应用包含以下核心文件:

  • app.py:Flask主应用
  • scripts/misfunciones.py:Excel数据处理逻辑
  • templates/index.html:前端上传页面

原始代码

app.py

from flask import Flask, render_template, request, send_file
import os
from scripts.misfunciones import process_data

app = Flask(__name__)

app.config['UPLOAD_FOLDER'] = os.path.join(os.path.dirname(__file__), 'uploads')

@app.route('/', methods=['GET', 'POST'])
def index():
    if request.method == 'POST':
        # Guardar los archivos subidos
        data_file = request.files['data_file']
        parametro_file = request.files['parametro_file']
        unidades_file = request.files['unidades_file']
        
        if not os.path.exists(app.config['UPLOAD_FOLDER']):
            os.makedirs(app.config['UPLOAD_FOLDER'])

        data_path = os.path.join(os.path.dirname(__file__), 'uploads', 'DATA.xlsx')
        parametro_path = os.path.join(os.path.dirname(__file__), 'uploads', '240424RF_Parametro.xlsx')
        unidades_path = os.path.join(os.path.dirname(__file__), 'uploads', '240306RF_Unidades.xlsx')
        output_path = os.path.join(os.path.dirname(__file__), 'uploads', 'DATA_FIN.xlsx')

        data_file.save(data_path)
        parametro_file.save(parametro_path)
        unidades_file.save(unidades_path)
        
        process_data(data_path, parametro_path, unidades_path, output_path)
        
        return send_file(output_path, as_attachment=True)

    return render_template('index.html')

if __name__ == '__main__':
    app.run(debug=True)

scripts/misfunciones.py

import pandas as pd
import openpyxl
import re
chars_dict = {
    '\n':'','*':'','°':'','(':'',')':'','+':'',
    '/':'','.':'','_':'',' ':'',',':'','\'':'',
    'á':'a','é':'e','í':'i','ó':'o','ú':'u','-':''
}

def var_lower(df):
    for i in df.columns[1:]:
        df[str(i)] = df[str(i)].str.lower()
        
def remove_pattern(df, pattern):
    for column in df.columns:
        df[column] = df[column].replace(pattern, '', regex=True)
    return df

def process_data(data_path, parametro_path, unidades_path, output_path):
    DATA = pd.read_excel(data_path, engine='openpyxl')
    Parametro = pd.read_excel(parametro_path, engine='openpyxl')
    Unidades = pd.read_excel(unidades_path, engine='openpyxl')

    DATA_CORR = DATA.copy()
    Parametro_CORR = Parametro.copy()
    Unidades_CORR = Unidades.copy()

    DATA_CORR = remove_pattern(DATA_CORR, 'αφ')
    DATA_CORR = remove_pattern(DATA_CORR, 'αδ')
    DATA_CORR = remove_pattern(DATA_CORR, 'φ')

    var_lower(DATA_CORR)
    var_lower(Parametro_CORR)
    var_lower(Unidades_CORR)
    for char, replacement in chars_dict.items():
        escaped_char = re.escape(char)
        DATA_CORR = DATA_CORR.replace(escaped_char, replacement, regex=True)
        Parametro_CORR = Parametro_CORR.replace(escaped_char, replacement, regex=True)
        Unidades_CORR = Unidades_CORR.replace(escaped_char, replacement, regex=True)

    Unificar_Unidad = []
    cantidad = 0
    for i in DATA_CORR['ID-UNIDADES']:
        cantidad += 1
        for j in range(len(Unidades_CORR)):
            for k in Unidades_CORR.columns[1:]:
                if i == Unidades_CORR[k][j]:
                    Unificar_Unidad.append(j + 1)
                    break
            if cantidad == len(Unificar_Unidad):
                break
        if cantidad != len(Unificar_Unidad):
            Unificar_Unidad.append(9999)
    DATA['num_unid'] = Unificar_Unidad

    Unificar_Param = []
    cantidad = 0
    for i in DATA_CORR['ID-PARANETROS']:
        cantidad += 1
        for j in range(len(Parametro_CORR)):
            for k in Parametro_CORR.columns[1:]:
                if i == Parametro_CORR[k][j]:
                    Unificar_Param.append(j + 1)
                    break
            if cantidad == len(Unificar_Param):
                break
        if cantidad != len(Unificar_Param):
            Unificar_Param.append(9999)
    DATA['num_param'] = Unificar_Param

    DATA.to_excel(output_path, index=False)
    return output_path

templates/index.html

<!DOCTYPE html>
<html lang="en">
<head>
    <meta charset="UTF-8">
    <meta name="viewport" content="width=device-width, initial-scale=1.0">
    <title>Subir Archivos</title>
</head>
<body>
    <h1>Subir Archivos para Procesar</h1>
    <form method="POST" enctype="multipart/form-data">
        <label for="data_file">Archivo DATA:</label>
        <input type="file" name="data_file" id="data_file" required><br><br>
        <label for="parametro_file">Archivo Parametro:</label>
        <input type="file" name="parametro_file" id="parametro_file" required><br><br>
        <label for="unidades_file">Archivo Unidades:</label>
        <input type="file" name="unidades_file" id="unidades_file" required><br><br>
        <input type="submit" value="Procesar">
    </form>
</body>
</html>

问题排查与修复步骤

1. 修复服务启动配置(核心原因)

Azure App Service默认不会自动识别Flask应用,需指定WSGI服务器启动命令:

  • 新增requirements.txt文件,包含所有依赖:
    flask==2.3.3
    pandas==2.1.4
    openpyxl==3.1.2
    gunicorn==21.2.0
    
  • 在Azure门户的App Service配置中,设置启动命令为:
    gunicorn --bind=0.0.0.0 --timeout 600 app:app
    
    (--timeout 600避免大文件处理超时)

2. 修正文件存储权限问题

Azure App Service仅允许/tmp目录可写,原代码中应用根目录下的uploads目录无写入权限:
修改app.py中的UPLOAD_FOLDER配置:

app.config['UPLOAD_FOLDER'] = os.path.join('/tmp', 'uploads')

3. 优化内存处理(避免内存溢出)

B1计划内存有限,原代码三重循环+全量加载Excel易导致内存不足:

  • 修改misfunciones.py中的匹配逻辑,用pandas的map替代三重循环,减少资源消耗:
    # 替换原Unificar_Unidad的三重循环
    unidad_map = {}
    for j in range(len(Unidades_CORR)):
        for k in Unidades_CORR.columns[1:]:
            unidad_val = Unidades_CORR[k][j]
            if unidad_val not in unidad_map:
                unidad_map[unidad_val] = j + 1
    DATA['num_unid'] = DATA_CORR['ID-UNIDADES'].map(unidad_map).fillna(9999).astype(int)
    
  • 改用内存流处理文件,完全避免落地存储:
    修改app.py,直接用BytesIO读取上传文件,处理后返回:
    from io import BytesIO
    
    @app.route('/', methods=['GET', 'POST'])
    def index():
        if request.method == 'POST':
            # 直接读取文件到内存流
            data_io = BytesIO(request.files['data_file'].read())
            parametro_io = BytesIO(request.files['parametro_file'].read())
            unidades_io = BytesIO(request.files['unidades_file'].read())
            
            # 修改process_data函数,接受IO对象而非文件路径
            output_io = process_data_io(data_io, parametro_io, unidades_io)
            
            return send_file(output_io, as_attachment=True, download_name='DATA_FIN.xlsx')
        return render_template('index.html')
    
    对应的process_data_io函数修改为:
    def process_data_io(data_io, parametro_io, unidades_io):
        DATA = pd.read_excel(data_io, engine='openpyxl')
        Parametro = pd.read_excel(parametro_io, engine='openpyxl')
        Unidades = pd.read_excel(unidades_io, engine='openpyxl')
        
        # 原处理逻辑不变...
        
        # 输出到内存流
        output_io = BytesIO()
        DATA.to_excel(output_io, index=False)
        output_io.seek(0)
        return output_io
    

4. 启用Azure日志排查

在Azure门户开启应用日志和Web服务器日志,查看具体错误信息(503通常是应用崩溃或启动失败,日志会显示具体异常,如依赖缺失、内存不足等)。


内容的提问来源于stack exchange,提问作者Rayan Renato Figueroa Asencios

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 02:59:52