You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas:筛选匹配指定字符串的行并创建含匹配词的新列

Pandas筛选匹配指定词汇并提取匹配项的实现方案

嘿,我来帮你搞定这个Pandas的需求!你需要筛选出text列包含指定词汇的行,同时把匹配到的词汇提取到新的found列里对吧?下面是具体的实现步骤,我会尽量讲得清晰易懂:

1. 先准备好数据和目标词汇

首先咱们把你给出的示例数据(包括期望输出里的对应行)转换成DataFrame:

import pandas as pd

# 构建原始DataFrame
data = {
    'id': ['a', 'b', 'c', 'd', 'e'],
    'text': [
        'simultaneous there the',
        'simultaneous there',
        'mul why',
        'have the',
        'then the late'
    ]
}
df = pd.DataFrame(data)

# 指定要匹配的词汇列表
target_words = ["mul","the","have", "then"]

2. 写个小函数提取匹配的词汇

我们需要一个函数来检查每行文本里包含哪些目标词汇,然后把它们用逗号连起来:

def get_matched_words(text, target_list):
    # 把文本拆成单个单词(这里按空格拆分,要是有标点的话可以用正则优化)
    text_words = text.split()
    # 找出文本和目标列表的交集词汇
    matched = [word for word in target_list if word in text_words]
    # 有匹配就返回逗号分隔的字符串,没有就返回空
    return ', '.join(matched) if matched else ''

3. 应用函数创建列并筛选结果

接下来把这个函数应用到text列,生成found列,然后筛选出有匹配结果的行(要是不需要筛选,直接跳过筛选步骤就行):

# 创建found列
df['found'] = df['text'].apply(lambda x: get_matched_words(x, target_words))

# 筛选出有匹配的行
filtered_df = df[df['found'] != ''].reset_index(drop=True)

# 把索引改成从1开始(和你的期望输出一致)
filtered_df.index = filtered_df.index + 1

最终得到的结果

运行完上面的代码,filtered_df就是你想要的结果啦:

idtextfound
1asimultaneous there thethe
2cmul whymul
3dhave thehave, the
4ethen the latethen, the

一些小提示

  • 如果你的文本里带有标点(比如"then,"这种),按空格拆分的方法会匹配不到,这时候可以用正则来提取纯单词,比如import re之后用text_words = re.findall(r'\b\w+\b', text)来获取单词。
  • 要是需要不区分大小写的匹配,就把文本和目标词汇都转成小写,比如text_words = [word.lower() for word in text.split()],同时target_list = [word.lower() for word in target_list]。

内容的提问来源于stack exchange,提问作者TRINADH NAGUBADI

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:14:27