Flask发票识别API在Postman中无法接收原始数据的问题排查
Flask调用Azure Document Intelligence识别Subway发票报错的解决方案
我用Flask搭了个调用Azure Document Intelligence的发票识别API,Postman测试时大部分发票都正常,但特定的Subway发票图片会出问题,报两个错:
"Incorrect format, please input the right format to import"
"TypeError: cannot use a string pattern on a bytes-like object"
问题原因分析
- 特殊格式/编码问题:Subway发票图片可能采用非标准图片编码,或带有额外元数据,转base64后无法被Azure API识别。
- 冗余文件处理逻辑:当前代码先保存上传文件到本地再读取转base64,这个过程可能损坏特殊格式的文件。
- API调用方式缺陷:base64编码传递文件存在兼容风险,直接传递字节流的可靠性更高。
解决方案步骤
1. 移除冗余的临时文件保存逻辑
直接读取上传的文件对象字节流,跳过本地保存步骤,避免文件损坏。
2. 优化Azure API调用方式
使用文件字节流直接传入begin_analyze_document,替代base64编码方式。
3. 添加文件格式校验
提前校验上传文件的MIME类型,确保是Azure支持的格式(JPG/PNG/PDF等)。
修改后的代码
utils.py(新增格式校验函数)
import os import base64 from urllib.parse import urlparse from azure.core.credentials import AzureKeyCredential from azure.ai.documentintelligence import DocumentIntelligenceClient from magic import from_buffer # Windows安装:pip install python-magic-bin;Linux/Mac安装:pip install python-magic def get_client(): endpoint = "your-azure-endpoint" api_key = "your-azure-api-key" client = DocumentIntelligenceClient(endpoint=endpoint,credential=AzureKeyCredential(api_key)) return client def is_file_or_url(input_string): if os.path.isfile(input_string): return 'file' elif urlparse(input_string).scheme in ['http', 'https']: return 'url' else: return 'unknown' def is_supported_file(file_obj): # 重置文件指针到开头 file_obj.seek(0) mime_type = from_buffer(file_obj.read(), mime=True) # Azure支持的文档格式:JPG/PNG/PDF等 supported_types = ['image/jpeg', 'image/png', 'application/pdf'] # 重置指针供后续读取 file_obj.seek(0) return mime_type in supported_types
app.py(简化文件处理逻辑)
import os from flask import Flask, request, jsonify from azure.ai.documentintelligence.models import AnalyzeDocumentRequest from utils import get_client, is_supported_file app = Flask(__name__) @app.route('/extract_invoice', methods=['POST']) def extract_invoice(): # 获取上传文件 file = request.files.get('file') if not file: return jsonify({"error": "未上传文件"}), 400 # 校验文件格式 if not is_supported_file(file): return jsonify({"error": "不支持的文件格式,请上传JPG/PNG/PDF格式文件"}), 400 model_id = 'prebuilt-invoice' document_ai_client = get_client() try: # 直接传递文件字节流给Azure API,无需转base64 poller = document_ai_client.begin_analyze_document( model_id, file, # Flask的FileStorage对象支持直接传入API locale="en-US", ) result = poller.result() except Exception as e: return jsonify({"error": str(e)}), 500 # 提取发票详情(原有逻辑保留) invoice_details = [] for document in result.documents: document_fields = document['fields'] fields = document_fields.keys() invoice_detail = {} for field in fields: if field == 'Items': items_list = [] items = document_fields[field] for item in items['valueArray']: item_fields = item['valueObject'] item_dict = {} for item_field in item_fields.keys(): value = item_fields[item_field].get('content', '') item_dict[item_field] = value items_list.append(item_dict) invoice_detail[field] = items_list else: value = document_fields[field].get('content', '') invoice_detail[field] = value invoice_details.append(invoice_detail) return jsonify(invoice_details) if __name__ == '__main__': app.run(debug=True)
额外排查建议
- 检查发票图片本身:用图片工具打开确认是否能正常显示,尝试用画图工具重新保存为标准JPG/PNG格式后再测试。
- 查看Azure日志:登录Azure门户,进入Document Intelligence资源的「日志」面板,查看详细错误信息定位问题。
- 调整地区参数:如果Subway发票是其他地区版本,尝试将
locale改为对应地区(比如en-GB)。
内容的提问来源于stack exchange,提问作者JaS
相关产品推荐
相关产品推荐

