如何用PyPDF4查找字符串所在行号并提取指定PDF文本内容?
用PyPDF4提取PDF中指定行的日期和分类信息
问题背景
我用PyPDF4读取PDF文件,提取后的文本如下:
Abrechnung30.11.2022
0,00+
Kontostand/Rechnungsabschlussam30.11.2022
672,06H
Rechnungsnummer:2022-11-3020:53:31.468209
01.12.2022
01.12.2022
Barausz.Debit.KFK
需要完成以下操作:
- 读取PDF文件
- 找到包含
Rechnungsnummer的行号,定位到下一行以及包含Barausz.的行,提取日期和分类信息
目前写的代码只能返回字符索引,无法获取行号,代码如下:
import PyPDF4 import re with open('../../Desktop/Konto_202212.pdf', 'rb') as pdfFile: reader = PyPDF4.PdfFileReader(pdfFile) page1 = reader.getPage(1) text = page1.extractText() a=text.find('Rechnungsnummer') print(a)
解决方案
因为提取的文本是带\n的长字符串,核心思路是把文本按换行符分割成行列表,这样就能通过索引直接对应行号,或者用正则直接匹配目标内容。
方法1:分割文本为行列表(获取行号+内容)
import PyPDF4 with open('../../Desktop/Konto_202212.pdf', 'rb') as pdfFile: reader = PyPDF4.PdfFileReader(pdfFile) page1 = reader.getPage(1) text = page1.extractText() # 按换行分割并清理空行、首尾空格 lines = [line.strip() for line in text.split('\n') if line.strip()] # 查找Rechnungsnummer所在的行索引(即行号,从0开始) target_idx = None for idx, line in enumerate(lines): if 'Rechnungsnummer' in line: target_idx = idx print(f"Rechnungsnummer所在行号:{idx+1}") # 若要从1开始计数就加1 break if target_idx is not None: # 获取下一行的日期 if target_idx + 1 < len(lines): next_line_date = lines[target_idx + 1] print(f"Rechnungsnummer下一行的日期:{next_line_date}") # 查找包含Barausz.的行内容 for line in lines: if 'Barausz.' in line: category_info = line print(f"分类信息:{category_info}")
方法2:正则表达式直接提取(无需行号)
如果不需要行号,只关心目标内容,用正则匹配更高效:
import PyPDF4 import re with open('../../Desktop/Konto_202212.pdf', 'rb') as pdfFile: reader = PyPDF4.PdfFileReader(pdfFile) page1 = reader.getPage(1) text = page1.extractText() # 匹配Rechnungsnummer行的下一行日期 date_match = re.search(r'Rechnungsnummer:.+\n(.+)', text, re.DOTALL) if date_match: extracted_date = date_match.group(1).strip() print(f"提取的日期:{extracted_date}") # 匹配包含Barausz.的分类信息 category_match = re.search(r'(Barausz\..+)', text) if category_match: extracted_category = category_match.group(1).strip() print(f"提取的分类信息:{extracted_category}")
两种方法对比
- 行列表分割法:直观清晰,能明确获取行号,适合需要处理多行关联逻辑的场景
- 正则匹配法:代码更简洁,执行效率更高,适合直接提取目标内容的场景
内容的提问来源于stack exchange,提问作者Kevin
相关产品推荐
相关产品推荐

