You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python Pandas DataFrame中提取含后续n字符的所有cash子串

Pandas提取DataFrame中包含cash及对应金额的子串

直接用正则表达式配合Pandas的字符串方法就能解决,步骤如下:

  1. 编写正则表达式匹配目标格式:
    正则r'cash\s*:?\s*\d+'可以精准覆盖两种格式:

    • cash 数字(如cash 15906810)
    • cash : 数字(如cash : 11234)
      其中\s*匹配任意数量空格,:?表示冒号可选,\d+匹配一串数字(金额)。
  2. 应用到DataFrame生成新列:
    使用str.findall提取所有匹配项,再用str.join将匹配项拼接成空格分隔的字符串。

完整代码:

import pandas as pd

sample = pd.DataFrame({'LongString': ["I am trying to find out how much cash 15906810 and this needs to be consistent cash :  69105060", "other words that are wrong cash : 11234 and more words cash 1526"]})

# 生成cash_string列
sample['cash_string'] = sample['LongString'].str.findall(r'cash\s*:?\s*\d+').str.join(' ')

# 查看结果
print(sample)

运行后得到的结果和你期望的sample_resolved完全一致。

内容的提问来源于stack exchange,提问作者alphonsethe3

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 12:45:43