You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中从多格式交易短信提取指定信息的最优方案与模块选择

解决方案:针对多格式银行交易短信的信息抽取

核心思路分析

你的场景核心特点是同一发送方短信格式统一,但不同发送方格式差异极大,所以没必要强行写通用正则(覆盖所有格式的正则不仅难写,还容易出现匹配错误)。最优路径是:按发送方分组,为每组构建适配的抽取规则,同时用自动化工具降低重复劳动量。具体可分为三步:

  1. 按发送方(number字段)对短信归类,同一组的短信复用同一套规则
  2. 为每组生成对应的信息抽取规则(优先用工具自动化生成,减少手动编写成本)
  3. 批量将规则应用到对应分组的短信,提取目标字段

推荐的Python模块及实操方案

1. pandas + re:快速落地的规则式方案(适合发送方数量少、格式固定的场景)

如果你的发送方数量不多,且每个发送方的短信格式非常固定,这是最直接高效的选择。用pandas分组后,为每组编写针对性的正则表达式即可。

示例代码片段:

import pandas as pd
import re

# 把样本数据转为DataFrame
sms_data = [
    {"message": "*boi star sandesh* rs 20 has been debited to your account xx2136 from pos-paytm.com on 08-11-2014.available balance 275.00.", "number": "boiind"},
    {"message": "your a/c xxxxx388847 debited inr 7,500.00 on 12/08/16 -transferred to mr. rajendra kurmi . a/c balance inr 1,314.45", "number": "amcbssbi"},
    # 其他短信数据...
]
df = pd.DataFrame(sms_data)

# 为不同发送方定义抽取函数
def extract_boiind_content(text):
    amount_match = re.search(r'rs (\d+)', text)
    balance_match = re.search(r'available balance (\d+\.\d+)', text)
    date_match = re.search(r'on (\d{2}-\d{2}-\d{4})', text)
    trans_type = "debited" if "debited" in text else "credited"
    
    return pd.Series(
        [amount_match.group(1) if amount_match else None,
         balance_match.group(1) if balance_match else None,
         date_match.group(1) if date_match else None,
         trans_type],
        index=["transaction_amount", "remaining_balance", "transaction_date", "transaction_type"]
    )

def extract_amcbssbi_content(text):
    amount_match = re.search(r'inr ([\d,\.]+)', text)
    balance_match = re.search(r'a/c balance inr ([\d,\.]+)', text)
    date_match = re.search(r'on (\d{2}/\d{2}/\d{2})', text)
    trans_type = "debited" if "debited" in text else "credited"
    
    return pd.Series(
        [amount_match.group(1) if amount_match else None,
         balance_match.group(1) if balance_match else None,
         date_match.group(1) if date_match else None,
         trans_type],
        index=["transaction_amount", "remaining_balance", "transaction_date", "transaction_type"]
    )

# 按分组应用抽取函数
boiind_result = df[df["number"] == "boiind"].apply(lambda x: extract_boiind_content(x["message"]), axis=1)
amcbssbi_result = df[df["number"] == "amcbssbi"].apply(lambda x: extract_amcbssbi_content(x["message"]), axis=1)

# 合并最终结果
final_result = pd.concat([df, pd.concat([boiind_result, amcbssbi_result])], axis=1)
print(final_result)

2. spaCy + 自定义NER:智能适配新格式的方案(适合发送方多、未来会新增格式的场景)

如果你的短信发送方数量多,或者未来会有新的发送方加入,手动写正则会非常繁琐。这时可以用spaCy训练自定义命名实体识别(NER)模型,让模型自动识别交易金额、余额、日期、交易类型这些实体。

示例代码片段:

import spacy
from spacy.training import Example
from spacy.util import minibatch, compounding

# 加载基础英文模型
nlp = spacy.load("en_core_web_sm")
# 添加自定义实体标签
ner = nlp.get_pipe("ner")
ner.add_label("TRANSACTION_AMOUNT")
ner.add_label("REMAINING_BALANCE")
ner.add_label("TRANSACTION_DATE")
ner.add_label("TRANSACTION_TYPE")

# 标注好的训练样本(示例,实际需要标注更多不同格式的短信)
TRAIN_DATA = [
    (
        "*boi star sandesh* rs 20 has been debited to your account xx2136 from pos-paytm.com on 08-11-2014.available balance 275.00.",
        {"entities": [(17, 19, "TRANSACTION_AMOUNT"), (103, 108, "REMAINING_BALANCE"), (75, 85, "TRANSACTION_DATE"), (25, 32, "TRANSACTION_TYPE")]}
    ),
    (
        "your a/c xxxxx388847 debited inr 7,500.00 on 12/08/16 -transferred to mr. rajendra kurmi . a/c balance inr 1,314.45",
        {"entities": [(26, 35, "TRANSACTION_AMOUNT"), (83, 92, "REMAINING_BALANCE"), (39, 47, "TRANSACTION_DATE"), (16, 23, "TRANSACTION_TYPE")]}
    )
]

# 训练模型
other_pipes = [pipe for pipe in nlp.pipe_names if pipe != "ner"]
with nlp.disable_pipes(*other_pipes):
    optimizer = nlp.begin_training()
    for iteration in range(10):
        losses = {}
        # 分批处理训练数据
        batches = minibatch(TRAIN_DATA, size=compounding(4.0, 32.0, 1.001))
        for batch in batches:
            for text, annotations in batch:
                doc = nlp.make_doc(text)
                example = Example.from_dict(doc, annotations)
                nlp.update([example], sgd=optimizer, losses=losses)
        print(f"Iteration {iteration + 1}, Loss: {losses['ner']:.4f}")

# 使用训练好的模型抽取信息
test_text = "your a/c no. xxxxxxxx1152 is debited for rs. 10,000.00 on 11-08-16 and a/c xxxxxxx847 credited (imps ref no 622421331357)"
doc = nlp(test_text)
print("提取结果:")
for ent in doc.ents:
    print(f"{ent.label_}: {ent.text}")

3. pyparsing:复杂结构的模板匹配方案(适合格式规则复杂的场景)

如果短信的格式有非常复杂的嵌套或逻辑结构,正则表达式难以处理,可以用pyparsing定义语法规则来匹配文本。比如可以定义金额的匹配规则(支持rs/inr前缀,逗号分隔的千位符、小数点)、日期的匹配规则(支持dd-mm-yyyy/dd/mm/yy等格式)。

方案选择总结

  • 发送方少、格式固定:优先选pandas+re,快速落地,开发成本低
  • 发送方多、需适配新格式:选spaCy自定义NER,一次训练后可自动适配同类格式的新短信
  • 格式结构复杂:考虑pyparsing,用语法规则精准匹配文本

另外,处理时要注意边界情况:比如金额的不同写法(rs 20、rs.10,000.00)、日期的不同格式(08-11-2014、12/08/16)、交易类型的多样化表述(debited/credited、online payment/imps transfer)。

内容的提问来源于stack exchange,提问作者Rahul

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:33:31