You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pdfplumber实现网页端PDF转Excel遇阻,求技术指导

Flask + pdfplumber PDF转Excel功能失效,Waitress部署后仍无法解决

问题描述

我尝试用Python的pdfplumber模块读取PDF文件并转换为Excel保存,但编写的Flask网页程序始终无法正常实现该功能。之后改用Waitress部署服务,问题依然存在,希望能得到技术排查方向和解决办法。

相关代码

Flask 核心代码

from flask import Flask, request, send_file
import pdfplumber
import pandas as pd
import os
from io import BytesIO

app = Flask(__name__)

@app.route('/convert', methods=['POST'])
def convert_pdf_to_excel():
    if 'pdf_file' not in request.files:
        return "未上传PDF文件", 400
    pdf_file = request.files['pdf_file']
    if pdf_file.filename == '':
        return "未选择文件", 400
    
    # 读取PDF内容
    try:
        with pdfplumber.open(pdf_file) as pdf:
            all_text = []
            for page in pdf.pages:
                text = page.extract_text()
                if text:
                    all_text.append(text.split('\n'))
            # 转换为DataFrame并保存为Excel
            df = pd.DataFrame(all_text)
            output = BytesIO()
            with pd.ExcelWriter(output, engine='openpyxl') as writer:
                df.to_excel(writer, index=False, header=False)
            output.seek(0)
            return send_file(output, download_name='converted.xlsx', as_attachment=True)
    except Exception as e:
        return f"转换失败: {str(e)}", 500

if __name__ == '__main__':
    app.run(debug=True)

Waitress 启动命令

waitress-serve --port=5000 app:app

终端运行错误信息

[2024-05-20 14:30:00 +0800] [12345] [INFO] Serving on http://0.0.0.0:5000
[2024-05-20 14:30:15 +0800] [12345] [ERROR] Exception occurred processing request
Traceback (most recent call last):
  File "waitress\channel.py", line 390, in service
    task.service()
  ...
  File "app.py", line 18, in convert_pdf_to_excel
    with pdfplumber.open(pdf_file) as pdf:
  File "pdfplumber\pdf.py", line 51, in open
    return PDF(fp, **kwargs)
  File "pdfplumber\pdf.py", line 43, in __init__
    self.stream = open(fp, "rb") if isinstance(fp, str) else fp
ValueError: I/O operation on closed file.

排查与解决办法

1. 修复文件流读取问题

Flask上传的FileStorage对象可能在操作过程中被提前关闭,需要先将文件内容读取到内存字节流中再交给pdfplumber处理:

# 修改convert_pdf_to_excel函数中的PDF读取部分
try:
    # 先读取文件内容到内存
    pdf_content = pdf_file.read()
    # 用BytesIO包装成可重复读取的流
    with pdfplumber.open(BytesIO(pdf_content)) as pdf:
        all_text = []
        for page in pdf.pages:
            text = page.extract_text()
            if text:
                all_text.append(text.split('\n'))
    # 后续Excel转换逻辑不变

2. 验证依赖包完整性

确保所有必需的依赖都已正确安装,尤其是Excel写入引擎openpyxl(pandas默认依赖它处理xlsx文件):

pip install --upgrade pdfplumber pandas openpyxl flask waitress

3. 细化日志定位问题

在代码中添加详细日志,方便追踪错误发生的具体环节:

import logging
logging.basicConfig(level=logging.DEBUG)

@app.route('/convert', methods=['POST'])
def convert_pdf_to_excel():
    if 'pdf_file' not in request.files:
        logging.warning("请求中无PDF文件")
        return "未上传PDF文件", 400
    pdf_file = request.files['pdf_file']
    if pdf_file.filename == '':
        logging.warning("用户未选择文件")
        return "未选择文件", 400
    
    try:
        logging.debug(f"开始处理文件: {pdf_file.filename}")
        pdf_content = pdf_file.read()
        logging.debug(f"文件读取完成,大小: {len(pdf_content)} bytes")
        with pdfplumber.open(BytesIO(pdf_content)) as pdf:
            logging.debug(f"PDF打开成功,共{len(pdf.pages)}页")
            all_text = []
            for idx, page in enumerate(pdf.pages):
                text = page.extract_text()
                logging.debug(f"第{idx+1}页文本提取结果: {text[:50]}...")
                if text:
                    all_text.append(text.split('\n'))
        # Excel转换逻辑
        df = pd.DataFrame(all_text)
        output = BytesIO()
        with pd.ExcelWriter(output, engine='openpyxl') as writer:
            df.to_excel(writer, index=False, header=False)
        output.seek(0)
        logging.debug("Excel文件生成成功")
        return send_file(output, download_name='converted.xlsx', as_attachment=True)
    except Exception as e:
        logging.error(f"转换失败: {str(e)}", exc_info=True)
        return f"转换失败,请查看服务端日志", 500

4. 验证PDF文件兼容性

确认上传的PDF是可编辑的文本型PDF,而非扫描生成的图片PDF:

  • 本地测试:用以下代码验证pdfplumber能否提取目标PDF的文本
import pdfplumber
with pdfplumber.open("你的测试文件.pdf") as pdf:
    first_page_text = pdf.pages[0].extract_text()
    print(first_page_text)

如果输出为空或乱码,说明是图片PDF,需要配合OCR工具(如pytesseract)先识别文本再转换。

5. 调整Waitress的请求限制

如果处理大体积PDF,需要修改Waitress的最大请求体大小限制,避免文件上传被截断:

# 允许最大10MB的请求体
waitress-serve --port=5000 --max-request-body-size=10485760 app:app

内容的提问来源于stack exchange,提问作者gary

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 17:23:12