You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python正则表达式提取DataFrame中指定开头的文本单词

解决Pandas提取指定开头单词的问题
  • 核心方案:利用Pandas字符串方法结合正则表达式,批量提取所有符合条件的单词并拼接成新字符串
  • 正则逻辑:用\b确保匹配完整单词,[TG]\w+匹配以T/G(实际场景全大写)开头的单词内容

代码示例

import pandas as pd

# 构造示例DataFrame
df = pd.DataFrame({
    'columnA': ['Hello world the sky is blue and the grass is green']
})

# 针对示例小写文本提取以t/g开头的单词
df['new_column'] = df['columnA'].str.findall(r'\b[tg]\w+').str.join(' ')

# 实际全大写场景的正则(直接替换为大写匹配)
# df['new_column'] = df['columnA'].str.findall(r'\b[TG]\w+').str.join(' ')

执行后df['new_column']的结果为the the grass green,完全符合需求。

常见问题说明

如果之前仅能获取首次匹配,大概率是误用了单次匹配方法(比如str.search()),而str.findall()会返回当前文本中所有符合正则的结果,再通过str.join(' ')将匹配到的单词列表转为空格分隔的字符串,实现批量高效处理。

扩展:不区分大小写匹配

如果需要兼容大小写混合的场景,可以引入正则忽略大小写标志:

import re
df['new_column'] = df['columnA'].str.findall(r'\b[tg]\w+', flags=re.IGNORECASE).str.join(' ')

内容的提问来源于stack exchange,提问作者Lukemorgan75

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 11:15:01