如何用Python正则表达式提取DataFrame中指定开头的文本单词
解决Pandas提取指定开头单词的问题
- 核心方案:利用Pandas字符串方法结合正则表达式,批量提取所有符合条件的单词并拼接成新字符串
- 正则逻辑:用
\b确保匹配完整单词,[TG]\w+匹配以T/G(实际场景全大写)开头的单词内容
代码示例
import pandas as pd # 构造示例DataFrame df = pd.DataFrame({ 'columnA': ['Hello world the sky is blue and the grass is green'] }) # 针对示例小写文本提取以t/g开头的单词 df['new_column'] = df['columnA'].str.findall(r'\b[tg]\w+').str.join(' ') # 实际全大写场景的正则(直接替换为大写匹配) # df['new_column'] = df['columnA'].str.findall(r'\b[TG]\w+').str.join(' ')
执行后df['new_column']的结果为the the grass green,完全符合需求。
常见问题说明
如果之前仅能获取首次匹配,大概率是误用了单次匹配方法(比如str.search()),而str.findall()会返回当前文本中所有符合正则的结果,再通过str.join(' ')将匹配到的单词列表转为空格分隔的字符串,实现批量高效处理。
扩展:不区分大小写匹配
如果需要兼容大小写混合的场景,可以引入正则忽略大小写标志:
import re df['new_column'] = df['columnA'].str.findall(r'\b[tg]\w+', flags=re.IGNORECASE).str.join(' ')
内容的提问来源于stack exchange,提问作者Lukemorgan75
相关产品推荐
相关产品推荐

