使用pdfplumber实现网页端PDF转Excel遇阻,求技术指导
Flask + pdfplumber PDF转Excel功能失效,Waitress部署后仍无法解决
问题描述
我尝试用Python的pdfplumber模块读取PDF文件并转换为Excel保存,但编写的Flask网页程序始终无法正常实现该功能。之后改用Waitress部署服务,问题依然存在,希望能得到技术排查方向和解决办法。
相关代码
Flask 核心代码
from flask import Flask, request, send_file import pdfplumber import pandas as pd import os from io import BytesIO app = Flask(__name__) @app.route('/convert', methods=['POST']) def convert_pdf_to_excel(): if 'pdf_file' not in request.files: return "未上传PDF文件", 400 pdf_file = request.files['pdf_file'] if pdf_file.filename == '': return "未选择文件", 400 # 读取PDF内容 try: with pdfplumber.open(pdf_file) as pdf: all_text = [] for page in pdf.pages: text = page.extract_text() if text: all_text.append(text.split('\n')) # 转换为DataFrame并保存为Excel df = pd.DataFrame(all_text) output = BytesIO() with pd.ExcelWriter(output, engine='openpyxl') as writer: df.to_excel(writer, index=False, header=False) output.seek(0) return send_file(output, download_name='converted.xlsx', as_attachment=True) except Exception as e: return f"转换失败: {str(e)}", 500 if __name__ == '__main__': app.run(debug=True)
Waitress 启动命令
waitress-serve --port=5000 app:app
终端运行错误信息
[2024-05-20 14:30:00 +0800] [12345] [INFO] Serving on http://0.0.0.0:5000 [2024-05-20 14:30:15 +0800] [12345] [ERROR] Exception occurred processing request Traceback (most recent call last): File "waitress\channel.py", line 390, in service task.service() ... File "app.py", line 18, in convert_pdf_to_excel with pdfplumber.open(pdf_file) as pdf: File "pdfplumber\pdf.py", line 51, in open return PDF(fp, **kwargs) File "pdfplumber\pdf.py", line 43, in __init__ self.stream = open(fp, "rb") if isinstance(fp, str) else fp ValueError: I/O operation on closed file.
排查与解决办法
1. 修复文件流读取问题
Flask上传的FileStorage对象可能在操作过程中被提前关闭,需要先将文件内容读取到内存字节流中再交给pdfplumber处理:
# 修改convert_pdf_to_excel函数中的PDF读取部分 try: # 先读取文件内容到内存 pdf_content = pdf_file.read() # 用BytesIO包装成可重复读取的流 with pdfplumber.open(BytesIO(pdf_content)) as pdf: all_text = [] for page in pdf.pages: text = page.extract_text() if text: all_text.append(text.split('\n')) # 后续Excel转换逻辑不变
2. 验证依赖包完整性
确保所有必需的依赖都已正确安装,尤其是Excel写入引擎openpyxl(pandas默认依赖它处理xlsx文件):
pip install --upgrade pdfplumber pandas openpyxl flask waitress
3. 细化日志定位问题
在代码中添加详细日志,方便追踪错误发生的具体环节:
import logging logging.basicConfig(level=logging.DEBUG) @app.route('/convert', methods=['POST']) def convert_pdf_to_excel(): if 'pdf_file' not in request.files: logging.warning("请求中无PDF文件") return "未上传PDF文件", 400 pdf_file = request.files['pdf_file'] if pdf_file.filename == '': logging.warning("用户未选择文件") return "未选择文件", 400 try: logging.debug(f"开始处理文件: {pdf_file.filename}") pdf_content = pdf_file.read() logging.debug(f"文件读取完成,大小: {len(pdf_content)} bytes") with pdfplumber.open(BytesIO(pdf_content)) as pdf: logging.debug(f"PDF打开成功,共{len(pdf.pages)}页") all_text = [] for idx, page in enumerate(pdf.pages): text = page.extract_text() logging.debug(f"第{idx+1}页文本提取结果: {text[:50]}...") if text: all_text.append(text.split('\n')) # Excel转换逻辑 df = pd.DataFrame(all_text) output = BytesIO() with pd.ExcelWriter(output, engine='openpyxl') as writer: df.to_excel(writer, index=False, header=False) output.seek(0) logging.debug("Excel文件生成成功") return send_file(output, download_name='converted.xlsx', as_attachment=True) except Exception as e: logging.error(f"转换失败: {str(e)}", exc_info=True) return f"转换失败,请查看服务端日志", 500
4. 验证PDF文件兼容性
确认上传的PDF是可编辑的文本型PDF,而非扫描生成的图片PDF:
- 本地测试:用以下代码验证pdfplumber能否提取目标PDF的文本
import pdfplumber with pdfplumber.open("你的测试文件.pdf") as pdf: first_page_text = pdf.pages[0].extract_text() print(first_page_text)
如果输出为空或乱码,说明是图片PDF,需要配合OCR工具(如pytesseract)先识别文本再转换。
5. 调整Waitress的请求限制
如果处理大体积PDF,需要修改Waitress的最大请求体大小限制,避免文件上传被截断:
# 允许最大10MB的请求体 waitress-serve --port=5000 --max-request-body-size=10485760 app:app
内容的提问来源于stack exchange,提问作者gary
相关产品推荐
相关产品推荐

