You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

大尺寸PDF文本对比Python代码报错排查请求(附代码及错误)

大尺寸PDF处理失败:TypeError: %d format: a real number is required, not bytes

我编写了一段用于提取PDF文件文本并对比信息的Python代码,该代码可正常处理小尺寸PDF,但在处理大尺寸PDF时执行失败并弹出各类错误信息,最后出现的错误为:TypeError: %d format: a real number is required, not bytes。以下是完整代码:

import pdfminer
import pandas as pd
from time import sleep
from tqdm import tqdm
from itertools import chain
import slate


# List of pdf files to process
pdf_files = ['file1.pdf', 'file2.pdf']

# Create a list to store the text from each PDF
pdf1_text = []
pdf2_text = []

# Iterate through each pdf file
for pdf_file in tqdm(pdf_files):
    # Open the pdf file
    with open(pdf_file, 'rb') as pdf_now:
        # Extract text using slate
        text = slate.PDF(pdf_now)
        text = text[0].split('\n')
        if pdf_file == pdf_files[0]:    
            pdf1_text.append(text)
        else:
            pdf2_text.append(text)

    sleep(20)

pdf1_text = list(chain.from_iterable(pdf1_text))
pdf2_text = list(chain.from_iterable(pdf2_text))

differences = set(pdf1_text).symmetric_difference(pdf2_text)

## Create a new dataframe to hold the differences
differences_df = pd.DataFrame(columns=['pdf1_text', 'pdf2_text'])

# Iterate through the differences and add them to the dataframe
for difference in differences:
    # Create a new row in the dataframe with the difference from pdf1 and pdf2
    differences_df = differences_df.append({'pdf1_text': difference if difference in pdf1_text else '',
                                            'pdf2_text': difference if difference in pdf2_text else ''}, ignore_index=True)

# Write the dataframe to an excel sheet
differences_df = differences_df.applymap(lambda x: x.encode('unicode_escape').decode('utf-8') if isinstance(x, str) else x)

differences_df.to_excel('differences.xlsx', index=False, engine='openpyxl')


import openpyxl

import re

# Load the Excel file into a dataframe
df = pd.read_excel("differences.xlsx")

# Create a condition to check the number of words in each cell
for column in ["pdf1_text", "pdf2_text"]:
    df[f"{column}_word_count"] = df[column].str.split().str.len()
    condition = df[f"{column}_word_count"] < 10
    # Drop the rows that meet the condition
    df = df[~condition]

for column in ["pdf1_text", "pdf2_text"]:
    df = df.drop(f"{column}_word_count", axis=1)


# Save the modified dataframe to a new Excel file
df.to_excel("differences.xlsx", index=False)

问题根源分析

  • slate库的局限性:slate依赖老旧的pdfminer分支,且长期未维护,对大PDF的解析支持不足。你遇到的%d格式错误,是因为slate内部处理大文件时,误将字节流(bytes)传入了需要数字的格式化逻辑中,触发类型不匹配异常。
  • 内存过载:大PDF的文本量极大,你通过嵌套列表存储所有文本行再展开的方式,会占用大量内存,容易引发内存溢出,进而导致各类奇怪的类型错误。
  • 低效操作叠加:循环使用DataFrame.append添加行、不必要的sleep(20)等操作,会放大大文件处理时的性能问题,甚至触发隐性错误。

修复方案

1. 替换slate,改用pdfminer.six直接解析

使用官方维护的pdfminer.six库(需先安装:pip install pdfminer.six),直接控制文本提取流程,避免黑箱错误:

from pdfminer.high_level import extract_pages
from pdfminer.layout import LTTextContainer

def extract_pdf_text(pdf_path):
    text_lines = []
    # 逐页解析,降低内存占用
    for page_layout in extract_pages(pdf_path):
        for element in page_layout:
            if isinstance(element, LTTextContainer):
                text_lines.extend(element.get_text().split('\n'))
    return text_lines

2. 优化核心逻辑

替换原代码中slate的使用部分,同时移除无效操作:

for pdf_file in tqdm(pdf_files):
    # 直接获取拆分后的文本行,无需嵌套存储
    text_lines = extract_pdf_text(pdf_file)
    if pdf_file == pdf_files[0]:    
        pdf1_text = text_lines
    else:
        pdf2_text = text_lines
    # 移除不必要的sleep

3. 优化DataFrame构建

避免循环append,直接用字典列表批量创建DataFrame:

diff_rows = []
for diff in differences:
    diff_rows.append({
        'pdf1_text': diff if diff in pdf1_text else '',
        'pdf2_text': diff if diff in pdf2_text else ''
    })
differences_df = pd.DataFrame(diff_rows)

4. 简化特殊字符处理

去掉冗余的编码转换逻辑,避免引发额外的编码错误:

differences_df = differences_df.applymap(lambda x: x if isinstance(x, str) else '')

内容的提问来源于stack exchange,提问作者hexapod

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 04:55:23