You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何仅37条数据存入Python字典?PyPDF2读取PDF数据丢失问题

问题:PyPDF2提取PDF数据存入字典时丢失大量条目

运行以下Python脚本时,已确认代码遍历了所有数据值,但并非所有数据都成功存入字典lineData。脚本通过PyPDF2读取PDF文件,提取匹配Jan \d{2,}格式的数据并存入字典,实际应有80+数据对,但最终字典仅包含37条数据。

file = open('path', 'rb')
readFile = PyPDF2.PdfFileReader(file)

lineData = {}

totalPages = readFile.numPages

for i in range(totalPages):
    pageObj = readFile.getPage(i)
    pageText = pageObj.extractText
    newTrans = re.compile(r'Jan \d{2,}')
    for line in pageText(pageObj).split('\n'):
        if newTrans.match(line):
            newValue = re.split(r'Jan \d{2,}', line)
            newValueStr = ' '.join(newValue)
            newKey = newTrans.findall(line)
            newKeyStr = ' '.join(newKey)
            print(newKeyStr + newValueStr)
            lineData[newKeyStr] = newValueStr
print(len(lineData))

核心原因:字典键唯一性导致的覆盖

Python字典的键是唯一的,当多个条目匹配到相同的newKeyStr时,后续条目会直接覆盖之前的条目,这是条目数量远少于预期的根本原因。比如PDF中若存在两行Jan 01 早餐和Jan 01 午餐,Jan 01作为键最终只会保留午餐,早餐被覆盖。

其他潜在问题

  • 文本提取调用错误:原代码中pageText = pageObj.extractText是赋值方法对象,后续pageText(pageObj)的调用逻辑错误,正确方式是直接调用pageObj.extractText()。
  • 正则匹配范围受限:newTrans.match(line)仅匹配行首内容,若Jan \d{2,}出现在行中间,该行数据会被忽略,应改用search检查整行。
  • 正则重复编译:循环内重复编译正则表达式,既浪费资源也无必要。

解决方案

1. 多值存储解决键重复问题

将字典的值改为列表,相同键下可保存多个条目:

import re
import PyPDF2

file = open('path', 'rb')
readFile = PyPDF2.PdfFileReader(file)

lineData = {}  # 键为日期字符串,值为对应内容的列表

totalPages = readFile.numPages
newTrans = re.compile(r'Jan \d{2,}')  # 正则编译移至循环外,提升效率

for i in range(totalPages):
    pageObj = readFile.getPage(i)
    pageText = pageObj.extractText()  # 修复文本提取调用
    for line in pageText.split('\n'):
        match = newTrans.search(line)  # 用search匹配整行内容
        if match:
            newKeyStr = match.group()
            newValueStr = line.replace(newKeyStr, '').strip()
            # 处理多值存储
            if newKeyStr in lineData:
                lineData[newKeyStr].append(newValueStr)
            else:
                lineData[newKeyStr] = [newValueStr]
            print(f"{newKeyStr} {newValueStr}")

# 输出统计信息
print(f"总条目数:{sum(len(v) for v in lineData.values())}")
print(f"不同日期数:{len(lineData)}")

2. 关键优化点说明

  • 修复extractText()的调用错误,确保正确提取页面文本。
  • 用search替代match,扩大正则匹配范围至整行。
  • 正则编译移至循环外,减少不必要的性能开销。
  • 改用列表存储多值,避免相同键的条目被覆盖。

内容的提问来源于stack exchange,提问作者freezeboi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 14:55:13