使用Python+Regex提取PDF银行账单数据,正则匹配返回None问题求助
问题根因
- 现有正则仅定义了2个捕获组,代码中尝试调用
group(3)、group(4)、group(5)时找不到对应分组,自然返回None - 正则没有匹配交易行后续的金额、余额、借贷标识字段的规则,无法拆分出需要的5个字段
修正方案
调整正则表达式,匹配每笔交易行的完整结构:每一行以日期开头,之后是任意字符的描述,接着是带逗号的交易金额、带逗号的余额,最后是括号包裹的Cr/Dr标识。
修正后的完整代码如下:
import re from collections import namedtuple BS_Kotak = namedtuple('BS_Kotak', 'Date Description Transaction_Amount Balance Balance_Dr_Cr') line_items = [] # 适配交易行结构的正则 Pattern_BankTransactions_1 = re.compile(r'^(\d{2}-[A-Za-z]{3}-\d{4})(.*?)\s+([\d,]+\.\d{2})\s+([\d,]+\.\d{2})\((Cr|Dr)\)$') for line in s.split('\n'): match_res = Pattern_BankTransactions_1.search(line.strip()) if match_res: date = match_res.group(1) desc = match_res.group(2).strip() trans_amt = match_res.group(3) bal_amt = match_res.group(4) bal_dr_cr = f"({match_res.group(5)})" line_items.append(BS_Kotak(date, desc, trans_amt, bal_amt, bal_dr_cr)) # 打印验证,和预期输出完全一致 print(date) print(desc) print(trans_amt) print(bal_amt) print(bal_dr_cr)
补充说明
你之前PDF提取代码的路径问题虽然已经解决,额外提醒下:原代码extract_pdf('pdf_file')传的是固定字符串pdf_file,实际应该传拼接了文件夹路径的文件变量,避免后续其他月份账单出现路径错误:
import os folder_path_pdf_file ='/home/sameer/PycharmProjects/pythonProject/Kotak Statements' for pdf_file in os.listdir(folder_path_pdf_file): full_path = os.path.join(folder_path_pdf_file, pdf_file) Kotak_BT= extract_pdf(full_path)
内容的提问来源于stack exchange,提问作者Sam G
相关产品推荐
相关产品推荐

