You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用PyPDF4查找字符串所在行号并提取指定PDF文本内容?

用PyPDF4提取PDF中指定行的日期和分类信息

问题背景

我用PyPDF4读取PDF文件,提取后的文本如下:

Abrechnung30.11.2022
0,00+
Kontostand/Rechnungsabschlussam30.11.2022
672,06H
Rechnungsnummer:2022-11-3020:53:31.468209
01.12.2022
01.12.2022
Barausz.Debit.KFK

需要完成以下操作:

  1. 读取PDF文件
  2. 找到包含Rechnungsnummer的行号,定位到下一行以及包含Barausz.的行,提取日期和分类信息

目前写的代码只能返回字符索引,无法获取行号,代码如下:

import PyPDF4
import re


with open('../../Desktop/Konto_202212.pdf', 'rb') as pdfFile:
    reader = PyPDF4.PdfFileReader(pdfFile)
    page1 = reader.getPage(1)
    text = page1.extractText()

    a=text.find('Rechnungsnummer')
    print(a)

解决方案

因为提取的文本是带\n的长字符串,核心思路是把文本按换行符分割成行列表,这样就能通过索引直接对应行号,或者用正则直接匹配目标内容。

方法1:分割文本为行列表(获取行号+内容)

import PyPDF4

with open('../../Desktop/Konto_202212.pdf', 'rb') as pdfFile:
    reader = PyPDF4.PdfFileReader(pdfFile)
    page1 = reader.getPage(1)
    text = page1.extractText()
    # 按换行分割并清理空行、首尾空格
    lines = [line.strip() for line in text.split('\n') if line.strip()]
    
    # 查找Rechnungsnummer所在的行索引(即行号,从0开始)
    target_idx = None
    for idx, line in enumerate(lines):
        if 'Rechnungsnummer' in line:
            target_idx = idx
            print(f"Rechnungsnummer所在行号:{idx+1}") # 若要从1开始计数就加1
            break
    
    if target_idx is not None:
        # 获取下一行的日期
        if target_idx + 1 < len(lines):
            next_line_date = lines[target_idx + 1]
            print(f"Rechnungsnummer下一行的日期:{next_line_date}")
        
        # 查找包含Barausz.的行内容
        for line in lines:
            if 'Barausz.' in line:
                category_info = line
                print(f"分类信息:{category_info}")

方法2:正则表达式直接提取(无需行号)

如果不需要行号,只关心目标内容,用正则匹配更高效:

import PyPDF4
import re

with open('../../Desktop/Konto_202212.pdf', 'rb') as pdfFile:
    reader = PyPDF4.PdfFileReader(pdfFile)
    page1 = reader.getPage(1)
    text = page1.extractText()
    
    # 匹配Rechnungsnummer行的下一行日期
    date_match = re.search(r'Rechnungsnummer:.+\n(.+)', text, re.DOTALL)
    if date_match:
        extracted_date = date_match.group(1).strip()
        print(f"提取的日期:{extracted_date}")
    
    # 匹配包含Barausz.的分类信息
    category_match = re.search(r'(Barausz\..+)', text)
    if category_match:
        extracted_category = category_match.group(1).strip()
        print(f"提取的分类信息:{extracted_category}")

两种方法对比

  • 行列表分割法:直观清晰,能明确获取行号,适合需要处理多行关联逻辑的场景
  • 正则匹配法:代码更简洁,执行效率更高,适合直接提取目标内容的场景

内容的提问来源于stack exchange,提问作者Kevin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 21:41:10