You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python正则匹配交易记录关联附加信息及清理冗余内容

解决方案

1. 关联附加信息到前一条交易记录

直接遍历PDF提取的文本行,区分交易记录与附加信息,将附加信息绑定到最近的交易条目即可:

  • 先定义匹配交易记录的正则表达式(需根据你实际的交易格式调整,比如假设交易以日期开头)
  • 初始化列表存储最终交易数据
  • 逐行处理文本:
    • 若当前行匹配交易正则,创建新的交易字典并加入列表
    • 若不匹配且列表不为空,将当前行内容追加到最后一条交易的附加信息字段中

代码示例:

import re

# 自定义交易匹配正则(根据你的实际交易格式修改)
transaction_regex = re.compile(r'^\d{2}/\d{2}/\d{4}.*')

# 模拟PDF提取的文本行
extracted_lines = [
    "01/05/2024 永辉超市 ¥128.5",
    "备注:购买粮油、日用品",
    "03/05/2024 地铁出行 ¥12",
    "发票号:SH20240503007"
]

transactions = []
for line in extracted_lines:
    line = line.strip()
    if not line:
        continue
    if transaction_regex.match(line):
        # 拆分交易字段(按需调整拆分规则)
        parts = re.split(r'\s{2,}', line)
        transactions.append({
            "日期": parts[0],
            "交易描述": parts[1],
            "金额": parts[2],
            "附加信息": ""
        })
    elif transactions:
        # 将附加信息追加到前一条交易
        transactions[-1]["附加信息"] += f" {line}"

# 导出为CSV
import csv
with open('transaction_records.csv', 'w', newline='', encoding='utf-8') as f:
    writer = csv.DictWriter(f, fieldnames=["日期", "交易描述", "金额", "附加信息"])
    writer.writeheader()
    writer.writerows(transactions)

2. 过滤超20字符的冗余长单词

可以通过两种方式实现,按需选择:

方法1:正则替换直接清除

import re

def remove_long_words(text, max_len=20):
    # 匹配长度超过max_len的连续字母/数字串,替换为空
    return re.sub(rf'\b\w{{{max_len+1},}}\b', '', text)

# 示例使用
raw_text = "文本中包含超长冗余串abcdefghijklmnopqrstuvwxyz12345"
cleaned_text = remove_long_words(raw_text)
print(cleaned_text)  # 输出:文本中包含超长冗余串

方法2:拆分过滤后重组

def remove_long_words(text, max_len=20):
    words = text.split()
    filtered = [word for word in words if len(word) <= max_len]
    return ' '.join(filtered)

注意:如果冗余长串包含特殊字符(如连字符、标点),需调整正则匹配规则,比如将\w替换为[a-zA-Z0-9-],适配你的实际冗余文本格式。


内容的提问来源于stack exchange,提问作者Madwolf

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 20:09:22