如何在Python Pandas DataFrame中提取含后续n字符的所有cash子串
Pandas提取DataFrame中包含cash及对应金额的子串
直接用正则表达式配合Pandas的字符串方法就能解决,步骤如下:
编写正则表达式匹配目标格式:
正则r'cash\s*:?\s*\d+'可以精准覆盖两种格式:cash 数字(如cash 15906810)cash : 数字(如cash : 11234)
其中\s*匹配任意数量空格,:?表示冒号可选,\d+匹配一串数字(金额)。
应用到DataFrame生成新列:
使用str.findall提取所有匹配项,再用str.join将匹配项拼接成空格分隔的字符串。
完整代码:
import pandas as pd sample = pd.DataFrame({'LongString': ["I am trying to find out how much cash 15906810 and this needs to be consistent cash : 69105060", "other words that are wrong cash : 11234 and more words cash 1526"]}) # 生成cash_string列 sample['cash_string'] = sample['LongString'].str.findall(r'cash\s*:?\s*\d+').str.join(' ') # 查看结果 print(sample)
运行后得到的结果和你期望的sample_resolved完全一致。
内容的提问来源于stack exchange,提问作者alphonsethe3
相关产品推荐
相关产品推荐

