You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python+Regex提取PDF银行账单数据,正则匹配返回None问题求助

问题根因

  • 现有正则仅定义了2个捕获组,代码中尝试调用group(3)、group(4)、group(5)时找不到对应分组,自然返回None
  • 正则没有匹配交易行后续的金额、余额、借贷标识字段的规则,无法拆分出需要的5个字段

修正方案

调整正则表达式,匹配每笔交易行的完整结构:每一行以日期开头,之后是任意字符的描述,接着是带逗号的交易金额、带逗号的余额,最后是括号包裹的Cr/Dr标识。

修正后的完整代码如下:

import re
from collections import namedtuple

BS_Kotak = namedtuple('BS_Kotak', 'Date Description Transaction_Amount Balance Balance_Dr_Cr')
line_items = []
# 适配交易行结构的正则
Pattern_BankTransactions_1 = re.compile(r'^(\d{2}-[A-Za-z]{3}-\d{4})(.*?)\s+([\d,]+\.\d{2})\s+([\d,]+\.\d{2})\((Cr|Dr)\)$')

for line in s.split('\n'):
    match_res = Pattern_BankTransactions_1.search(line.strip())
    if match_res:
        date = match_res.group(1)
        desc = match_res.group(2).strip()
        trans_amt = match_res.group(3)
        bal_amt = match_res.group(4)
        bal_dr_cr = f"({match_res.group(5)})"
        line_items.append(BS_Kotak(date, desc, trans_amt, bal_amt, bal_dr_cr))
        # 打印验证,和预期输出完全一致
        print(date)
        print(desc)
        print(trans_amt)
        print(bal_amt)
        print(bal_dr_cr)

补充说明

你之前PDF提取代码的路径问题虽然已经解决,额外提醒下:原代码extract_pdf('pdf_file')传的是固定字符串pdf_file,实际应该传拼接了文件夹路径的文件变量,避免后续其他月份账单出现路径错误:

import os
folder_path_pdf_file ='/home/sameer/PycharmProjects/pythonProject/Kotak Statements'
for pdf_file in os.listdir(folder_path_pdf_file):
    full_path = os.path.join(folder_path_pdf_file, pdf_file)
    Kotak_BT= extract_pdf(full_path)

内容的提问来源于stack exchange,提问作者Sam G

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 06:36:02