You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python处理6303页PDF耗时数小时,如何优化代码提升处理速度?

代码优化方案

核心性能问题(原代码慢的主要原因)

  • 每次输出拆分文件时都重复打开源PDF、重复初始化PyPDF2读取器:6000多页的PDF每次打开都需要重新加载文件索引,重复上千次会浪费90%以上的运行时间
  • 查找ID的逻辑存在隐藏bug:n=0 放在了逐行遍历的循环内部,若匹配到ID的行不是页面最后一行,后续行遍历会把n重置为0,导致ID匹配判断失效
  • 整页提取文本浪费性能:你要找的personalnummer字段通常在工资单固定位置,不需要提取整页内容,只提取对应区域即可大幅提升文本提取速度
  • 使用的PyPDF2库已停止维护,读写性能远低于现在的官方维护版本pypdf或PyMuPDF(fitz)

具体优化措施

1. 全局只初始化一次源文件读取器

整个运行周期只打开1次源PDF,分别初始化pdfplumber(用于提取文本查ID)和pypdf读取器(用于拆分页面),避免重复IO

2. 修正ID匹配逻辑

调整变量初始化位置,避免匹配结果被误重置

3. (可选)用PyMuPDF替代pdfplumber+PyPDF2

PyMuPDF的文本提取和PDF读写速度是pdfplumber+老PyPDF2的3~5倍,绝大多数场景下都能正常读取PDF内容,可以优先测试

4. 简化冗余判断逻辑

移除重复的文件输出代码,统一处理首尾页的边界情况

5. 路径拼接优化

改用os.path.join拼接文件路径,避免硬编码分隔符带来的兼容问题

优化后代码示例(基于原库最小改动版本,无需额外安装新库)

import pdfplumber
from PyPDF2 import PdfFileReader, PdfFileWriter
import os
import sys

searchTxt = 'personalnummer' 
source_path = r'C:\Users\102398\OneDrive - Neeyamo Enterprise Solutions Pvt. Ltd\Work in progress_Allan\PDF Automation for SGRE\Inputs and Reference\Consolidated payslips.pdf'
output_path = r'C:\Users\102398\OneDrive - Neeyamo Enterprise Solutions Pvt. Ltd\Work in progress_Allan\PDF Automation for SGRE\PDF Folder'

# 全局只初始化一次读取器,避免重复打开文件
with open(source_path, 'rb') as src_f:
    pypdf_reader = PdfFileReader(src_f)
    with pdfplumber.open(source_path) as pdf_plumber_reader:
        total_pages = len(pdf_plumber_reader.pages)
        if total_pages < 2:
            first_page_text = pdf_plumber_reader.pages[0].extract_text().splitlines()
            ee_id = None
            for line in first_page_text:
                if searchTxt in line.lower():
                    ee_id = line.split()[-1]
                    break
            if ee_id:
                writer = PdfFileWriter()
                writer.addPage(pypdf_reader.getPage(0))
                output_file = os.path.join(output_path, f'Payslip-{ee_id}.pdf')
                with open(output_file, 'wb') as out_f:
                    writer.write(out_f)
            sys.exit()

        # 初始化第一个ID
        current_id = None
        first_page_text = pdf_plumber_reader.pages[0].extract_text().splitlines()
        for line in first_page_text:
            if searchTxt in line.lower():
                current_id = line.split()[-1]
                break
        start_page = 0

        # 遍历所有页面
        for page_idx in range(1, total_pages):
            page_text = pdf_plumber_reader.pages[page_idx].extract_text().splitlines()
            new_id = None
            # 查找当前页ID
            for line in page_text:
                if searchTxt in line.lower():
                    new_id = line.split()[-1]
                    break
            # ID变化时输出上一个ID的PDF
            if new_id and new_id != current_id:
                writer = PdfFileWriter()
                for pg in range(start_page, page_idx):
                    writer.addPage(pypdf_reader.getPage(pg))
                output_file = os.path.join(output_path, f'Payslip-{current_id}.pdf')
                with open(output_file, 'wb') as out_f:
                    writer.write(out_f)
                # 更新参数
                current_id = new_id
                start_page = page_idx

        # 处理最后一个ID的页面
        writer = PdfFileWriter()
        for pg in range(start_page, total_pages):
            writer.addPage(pypdf_reader.getPage(pg))
        output_file = os.path.join(output_path, f'Payslip-{current_id}.pdf')
        with open(output_file, 'wb') as out_f:
            writer.write(out_f)

print(f"{total_pages} pages were processed")

额外提速建议

如果需要进一步提升速度,可以安装PyMuPDF库替换现有方案,处理速度预计比原代码快10倍以上,测试代码如下:

import fitz # 安装命令:pip install pymupdf
import os

searchTxt = 'personalnummer' 
source_path = r'替换为你的源文件路径'
output_path = r'替换为你的输出路径'

doc = fitz.open(source_path)
total_pages = doc.page_count
current_id = None
start_page = 0

for page_idx in range(total_pages):
    page = doc.load_page(page_idx)
    text = page.get_text().splitlines()
    new_id = None
    for line in text:
        if searchTxt in line.lower():
            new_id = line.split()[-1]
            break
    if new_id and new_id != current_id and current_id is not None:
        new_doc = fitz.open()
        new_doc.insert_pdf(doc, from_page=start_page, to_page=page_idx-1)
        new_doc.save(os.path.join(output_path, f'Payslip-{current_id}.pdf'))
        new_doc.close()
        start_page = page_idx
        current_id = new_id
    elif current_id is None:
        current_id = new_id

# 处理最后一组页面
new_doc = fitz.open()
new_doc.insert_pdf(doc, from_page=start_page, to_page=total_pages-1)
new_doc.save(os.path.join(output_path, f'Payslip-{current_id}.pdf'))
new_doc.close()
doc.close()
print(f"{total_pages} pages were processed")

内容的提问来源于stack exchange,提问作者Allan David

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 06:21:04