使用Python处理6303页PDF耗时数小时,如何优化代码提升处理速度?
代码优化方案
核心性能问题(原代码慢的主要原因)
- 每次输出拆分文件时都重复打开源PDF、重复初始化PyPDF2读取器:6000多页的PDF每次打开都需要重新加载文件索引,重复上千次会浪费90%以上的运行时间
- 查找ID的逻辑存在隐藏bug:
n=0放在了逐行遍历的循环内部,若匹配到ID的行不是页面最后一行,后续行遍历会把n重置为0,导致ID匹配判断失效 - 整页提取文本浪费性能:你要找的
personalnummer字段通常在工资单固定位置,不需要提取整页内容,只提取对应区域即可大幅提升文本提取速度 - 使用的PyPDF2库已停止维护,读写性能远低于现在的官方维护版本pypdf或PyMuPDF(fitz)
具体优化措施
1. 全局只初始化一次源文件读取器
整个运行周期只打开1次源PDF,分别初始化pdfplumber(用于提取文本查ID)和pypdf读取器(用于拆分页面),避免重复IO
2. 修正ID匹配逻辑
调整变量初始化位置,避免匹配结果被误重置
3. (可选)用PyMuPDF替代pdfplumber+PyPDF2
PyMuPDF的文本提取和PDF读写速度是pdfplumber+老PyPDF2的3~5倍,绝大多数场景下都能正常读取PDF内容,可以优先测试
4. 简化冗余判断逻辑
移除重复的文件输出代码,统一处理首尾页的边界情况
5. 路径拼接优化
改用os.path.join拼接文件路径,避免硬编码分隔符带来的兼容问题
优化后代码示例(基于原库最小改动版本,无需额外安装新库)
import pdfplumber from PyPDF2 import PdfFileReader, PdfFileWriter import os import sys searchTxt = 'personalnummer' source_path = r'C:\Users\102398\OneDrive - Neeyamo Enterprise Solutions Pvt. Ltd\Work in progress_Allan\PDF Automation for SGRE\Inputs and Reference\Consolidated payslips.pdf' output_path = r'C:\Users\102398\OneDrive - Neeyamo Enterprise Solutions Pvt. Ltd\Work in progress_Allan\PDF Automation for SGRE\PDF Folder' # 全局只初始化一次读取器,避免重复打开文件 with open(source_path, 'rb') as src_f: pypdf_reader = PdfFileReader(src_f) with pdfplumber.open(source_path) as pdf_plumber_reader: total_pages = len(pdf_plumber_reader.pages) if total_pages < 2: first_page_text = pdf_plumber_reader.pages[0].extract_text().splitlines() ee_id = None for line in first_page_text: if searchTxt in line.lower(): ee_id = line.split()[-1] break if ee_id: writer = PdfFileWriter() writer.addPage(pypdf_reader.getPage(0)) output_file = os.path.join(output_path, f'Payslip-{ee_id}.pdf') with open(output_file, 'wb') as out_f: writer.write(out_f) sys.exit() # 初始化第一个ID current_id = None first_page_text = pdf_plumber_reader.pages[0].extract_text().splitlines() for line in first_page_text: if searchTxt in line.lower(): current_id = line.split()[-1] break start_page = 0 # 遍历所有页面 for page_idx in range(1, total_pages): page_text = pdf_plumber_reader.pages[page_idx].extract_text().splitlines() new_id = None # 查找当前页ID for line in page_text: if searchTxt in line.lower(): new_id = line.split()[-1] break # ID变化时输出上一个ID的PDF if new_id and new_id != current_id: writer = PdfFileWriter() for pg in range(start_page, page_idx): writer.addPage(pypdf_reader.getPage(pg)) output_file = os.path.join(output_path, f'Payslip-{current_id}.pdf') with open(output_file, 'wb') as out_f: writer.write(out_f) # 更新参数 current_id = new_id start_page = page_idx # 处理最后一个ID的页面 writer = PdfFileWriter() for pg in range(start_page, total_pages): writer.addPage(pypdf_reader.getPage(pg)) output_file = os.path.join(output_path, f'Payslip-{current_id}.pdf') with open(output_file, 'wb') as out_f: writer.write(out_f) print(f"{total_pages} pages were processed")
额外提速建议
如果需要进一步提升速度,可以安装PyMuPDF库替换现有方案,处理速度预计比原代码快10倍以上,测试代码如下:
import fitz # 安装命令:pip install pymupdf import os searchTxt = 'personalnummer' source_path = r'替换为你的源文件路径' output_path = r'替换为你的输出路径' doc = fitz.open(source_path) total_pages = doc.page_count current_id = None start_page = 0 for page_idx in range(total_pages): page = doc.load_page(page_idx) text = page.get_text().splitlines() new_id = None for line in text: if searchTxt in line.lower(): new_id = line.split()[-1] break if new_id and new_id != current_id and current_id is not None: new_doc = fitz.open() new_doc.insert_pdf(doc, from_page=start_page, to_page=page_idx-1) new_doc.save(os.path.join(output_path, f'Payslip-{current_id}.pdf')) new_doc.close() start_page = page_idx current_id = new_id elif current_id is None: current_id = new_id # 处理最后一组页面 new_doc = fitz.open() new_doc.insert_pdf(doc, from_page=start_page, to_page=total_pages-1) new_doc.save(os.path.join(output_path, f'Payslip-{current_id}.pdf')) new_doc.close() doc.close() print(f"{total_pages} pages were processed")
内容的提问来源于stack exchange,提问作者Allan David
相关产品推荐
相关产品推荐

