为何仅37条数据存入Python字典?PyPDF2读取PDF数据丢失问题
问题:PyPDF2提取PDF数据存入字典时丢失大量条目
运行以下Python脚本时,已确认代码遍历了所有数据值,但并非所有数据都成功存入字典lineData。脚本通过PyPDF2读取PDF文件,提取匹配Jan \d{2,}格式的数据并存入字典,实际应有80+数据对,但最终字典仅包含37条数据。
file = open('path', 'rb') readFile = PyPDF2.PdfFileReader(file) lineData = {} totalPages = readFile.numPages for i in range(totalPages): pageObj = readFile.getPage(i) pageText = pageObj.extractText newTrans = re.compile(r'Jan \d{2,}') for line in pageText(pageObj).split('\n'): if newTrans.match(line): newValue = re.split(r'Jan \d{2,}', line) newValueStr = ' '.join(newValue) newKey = newTrans.findall(line) newKeyStr = ' '.join(newKey) print(newKeyStr + newValueStr) lineData[newKeyStr] = newValueStr print(len(lineData))
核心原因:字典键唯一性导致的覆盖
Python字典的键是唯一的,当多个条目匹配到相同的newKeyStr时,后续条目会直接覆盖之前的条目,这是条目数量远少于预期的根本原因。比如PDF中若存在两行Jan 01 早餐和Jan 01 午餐,Jan 01作为键最终只会保留午餐,早餐被覆盖。
其他潜在问题
- 文本提取调用错误:原代码中
pageText = pageObj.extractText是赋值方法对象,后续pageText(pageObj)的调用逻辑错误,正确方式是直接调用pageObj.extractText()。 - 正则匹配范围受限:
newTrans.match(line)仅匹配行首内容,若Jan \d{2,}出现在行中间,该行数据会被忽略,应改用search检查整行。 - 正则重复编译:循环内重复编译正则表达式,既浪费资源也无必要。
解决方案
1. 多值存储解决键重复问题
将字典的值改为列表,相同键下可保存多个条目:
import re import PyPDF2 file = open('path', 'rb') readFile = PyPDF2.PdfFileReader(file) lineData = {} # 键为日期字符串,值为对应内容的列表 totalPages = readFile.numPages newTrans = re.compile(r'Jan \d{2,}') # 正则编译移至循环外,提升效率 for i in range(totalPages): pageObj = readFile.getPage(i) pageText = pageObj.extractText() # 修复文本提取调用 for line in pageText.split('\n'): match = newTrans.search(line) # 用search匹配整行内容 if match: newKeyStr = match.group() newValueStr = line.replace(newKeyStr, '').strip() # 处理多值存储 if newKeyStr in lineData: lineData[newKeyStr].append(newValueStr) else: lineData[newKeyStr] = [newValueStr] print(f"{newKeyStr} {newValueStr}") # 输出统计信息 print(f"总条目数:{sum(len(v) for v in lineData.values())}") print(f"不同日期数:{len(lineData)}")
2. 关键优化点说明
- 修复
extractText()的调用错误,确保正确提取页面文本。 - 用
search替代match,扩大正则匹配范围至整行。 - 正则编译移至循环外,减少不必要的性能开销。
- 改用列表存储多值,避免相同键的条目被覆盖。
内容的提问来源于stack exchange,提问作者freezeboi
相关产品推荐
相关产品推荐

