Azure部署Python3.9+Flask Excel处理应用遇HTTP503错误求助
Azure App Service部署Flask Excel处理应用HTTP 503错误排查与修复
问题背景
基于Python 3.9 + Flask开发的Excel文件处理应用,本地运行完全正常,但部署至Azure App Service B1计划(1vCPU、1.75GB内存、100ACU)后出现HTTP 503错误。已尝试内存中处理文件(不落地存储),本地可行但Azure端仍失败,怀疑问题与Excel处理逻辑、权限配置或服务启动配置有关。
应用包含以下核心文件:
- app.py:Flask主应用
- scripts/misfunciones.py:Excel数据处理逻辑
- templates/index.html:前端上传页面
原始代码
app.py
from flask import Flask, render_template, request, send_file import os from scripts.misfunciones import process_data app = Flask(__name__) app.config['UPLOAD_FOLDER'] = os.path.join(os.path.dirname(__file__), 'uploads') @app.route('/', methods=['GET', 'POST']) def index(): if request.method == 'POST': # Guardar los archivos subidos data_file = request.files['data_file'] parametro_file = request.files['parametro_file'] unidades_file = request.files['unidades_file'] if not os.path.exists(app.config['UPLOAD_FOLDER']): os.makedirs(app.config['UPLOAD_FOLDER']) data_path = os.path.join(os.path.dirname(__file__), 'uploads', 'DATA.xlsx') parametro_path = os.path.join(os.path.dirname(__file__), 'uploads', '240424RF_Parametro.xlsx') unidades_path = os.path.join(os.path.dirname(__file__), 'uploads', '240306RF_Unidades.xlsx') output_path = os.path.join(os.path.dirname(__file__), 'uploads', 'DATA_FIN.xlsx') data_file.save(data_path) parametro_file.save(parametro_path) unidades_file.save(unidades_path) process_data(data_path, parametro_path, unidades_path, output_path) return send_file(output_path, as_attachment=True) return render_template('index.html') if __name__ == '__main__': app.run(debug=True)
scripts/misfunciones.py
import pandas as pd import openpyxl import re chars_dict = { '\n':'','*':'','°':'','(':'',')':'','+':'', '/':'','.':'','_':'',' ':'',',':'','\'':'', 'á':'a','é':'e','í':'i','ó':'o','ú':'u','-':'' } def var_lower(df): for i in df.columns[1:]: df[str(i)] = df[str(i)].str.lower() def remove_pattern(df, pattern): for column in df.columns: df[column] = df[column].replace(pattern, '', regex=True) return df def process_data(data_path, parametro_path, unidades_path, output_path): DATA = pd.read_excel(data_path, engine='openpyxl') Parametro = pd.read_excel(parametro_path, engine='openpyxl') Unidades = pd.read_excel(unidades_path, engine='openpyxl') DATA_CORR = DATA.copy() Parametro_CORR = Parametro.copy() Unidades_CORR = Unidades.copy() DATA_CORR = remove_pattern(DATA_CORR, 'αφ') DATA_CORR = remove_pattern(DATA_CORR, 'αδ') DATA_CORR = remove_pattern(DATA_CORR, 'φ') var_lower(DATA_CORR) var_lower(Parametro_CORR) var_lower(Unidades_CORR) for char, replacement in chars_dict.items(): escaped_char = re.escape(char) DATA_CORR = DATA_CORR.replace(escaped_char, replacement, regex=True) Parametro_CORR = Parametro_CORR.replace(escaped_char, replacement, regex=True) Unidades_CORR = Unidades_CORR.replace(escaped_char, replacement, regex=True) Unificar_Unidad = [] cantidad = 0 for i in DATA_CORR['ID-UNIDADES']: cantidad += 1 for j in range(len(Unidades_CORR)): for k in Unidades_CORR.columns[1:]: if i == Unidades_CORR[k][j]: Unificar_Unidad.append(j + 1) break if cantidad == len(Unificar_Unidad): break if cantidad != len(Unificar_Unidad): Unificar_Unidad.append(9999) DATA['num_unid'] = Unificar_Unidad Unificar_Param = [] cantidad = 0 for i in DATA_CORR['ID-PARANETROS']: cantidad += 1 for j in range(len(Parametro_CORR)): for k in Parametro_CORR.columns[1:]: if i == Parametro_CORR[k][j]: Unificar_Param.append(j + 1) break if cantidad == len(Unificar_Param): break if cantidad != len(Unificar_Param): Unificar_Param.append(9999) DATA['num_param'] = Unificar_Param DATA.to_excel(output_path, index=False) return output_path
templates/index.html
<!DOCTYPE html> <html lang="en"> <head> <meta charset="UTF-8"> <meta name="viewport" content="width=device-width, initial-scale=1.0"> <title>Subir Archivos</title> </head> <body> <h1>Subir Archivos para Procesar</h1> <form method="POST" enctype="multipart/form-data"> <label for="data_file">Archivo DATA:</label> <input type="file" name="data_file" id="data_file" required><br><br> <label for="parametro_file">Archivo Parametro:</label> <input type="file" name="parametro_file" id="parametro_file" required><br><br> <label for="unidades_file">Archivo Unidades:</label> <input type="file" name="unidades_file" id="unidades_file" required><br><br> <input type="submit" value="Procesar"> </form> </body> </html>
问题排查与修复步骤
1. 修复服务启动配置(核心原因)
Azure App Service默认不会自动识别Flask应用,需指定WSGI服务器启动命令:
- 新增
requirements.txt文件,包含所有依赖:flask==2.3.3 pandas==2.1.4 openpyxl==3.1.2 gunicorn==21.2.0 - 在Azure门户的App Service配置中,设置启动命令为:
(gunicorn --bind=0.0.0.0 --timeout 600 app:app--timeout 600避免大文件处理超时)
2. 修正文件存储权限问题
Azure App Service仅允许/tmp目录可写,原代码中应用根目录下的uploads目录无写入权限:
修改app.py中的UPLOAD_FOLDER配置:
app.config['UPLOAD_FOLDER'] = os.path.join('/tmp', 'uploads')
3. 优化内存处理(避免内存溢出)
B1计划内存有限,原代码三重循环+全量加载Excel易导致内存不足:
- 修改
misfunciones.py中的匹配逻辑,用pandas的map替代三重循环,减少资源消耗:# 替换原Unificar_Unidad的三重循环 unidad_map = {} for j in range(len(Unidades_CORR)): for k in Unidades_CORR.columns[1:]: unidad_val = Unidades_CORR[k][j] if unidad_val not in unidad_map: unidad_map[unidad_val] = j + 1 DATA['num_unid'] = DATA_CORR['ID-UNIDADES'].map(unidad_map).fillna(9999).astype(int) - 改用内存流处理文件,完全避免落地存储:
修改app.py,直接用BytesIO读取上传文件,处理后返回:
对应的from io import BytesIO @app.route('/', methods=['GET', 'POST']) def index(): if request.method == 'POST': # 直接读取文件到内存流 data_io = BytesIO(request.files['data_file'].read()) parametro_io = BytesIO(request.files['parametro_file'].read()) unidades_io = BytesIO(request.files['unidades_file'].read()) # 修改process_data函数,接受IO对象而非文件路径 output_io = process_data_io(data_io, parametro_io, unidades_io) return send_file(output_io, as_attachment=True, download_name='DATA_FIN.xlsx') return render_template('index.html')process_data_io函数修改为:def process_data_io(data_io, parametro_io, unidades_io): DATA = pd.read_excel(data_io, engine='openpyxl') Parametro = pd.read_excel(parametro_io, engine='openpyxl') Unidades = pd.read_excel(unidades_io, engine='openpyxl') # 原处理逻辑不变... # 输出到内存流 output_io = BytesIO() DATA.to_excel(output_io, index=False) output_io.seek(0) return output_io
4. 启用Azure日志排查
在Azure门户开启应用日志和Web服务器日志,查看具体错误信息(503通常是应用崩溃或启动失败,日志会显示具体异常,如依赖缺失、内存不足等)。
内容的提问来源于stack exchange,提问作者Rayan Renato Figueroa Asencios
相关产品推荐
相关产品推荐

