You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何不使用正则表达式,用Pandas提取DataFrame文本中匹配短语列表的内容?

解决方案:无需正则表达式实现匹配短语列

可以实现,直接利用Python字符串的in成员判断和pandas的apply方法就能完成,完全不需要正则表达式。

步骤代码示例:

import pandas as pd

# 构造示例DataFrame
df = pd.DataFrame({
    'text': [
        'my name is abc',
        'xyz is a fruit',
        'abc likes per'
    ]
})

# 目标短语列表
phrase_list = ['abc', 'fruit', 'likes per']

# 定义匹配函数:遍历短语列表,收集存在于当前text中的短语
def match_phrases(text):
    return [phrase for phrase in phrase_list if phrase in text]

# 生成terms列
df['terms'] = df['text'].apply(match_phrases)

# 查看结果
print(df)

运行结果:

text               terms
0  my name is abc              ['abc']
1  xyz is a fruit            ['fruit']
2  abc likes per  ['abc', 'likes per']

说明:

  • 核心逻辑是通过列表推导式遍历短语列表,用phrase in text判断短语是否完整出现在当前文本字符串中,这是Python原生的字符串成员检查,完全不依赖正则。
  • 用df['text'].apply(match_phrases)将函数应用到每一行的text值上,直接生成对应的terms列。

内容的提问来源于stack exchange,提问作者S_S

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 00:05:46