Pandas:筛选匹配指定字符串的行并创建含匹配词的新列
Pandas筛选匹配指定词汇并提取匹配项的实现方案
嘿,我来帮你搞定这个Pandas的需求!你需要筛选出text列包含指定词汇的行,同时把匹配到的词汇提取到新的found列里对吧?下面是具体的实现步骤,我会尽量讲得清晰易懂:
1. 先准备好数据和目标词汇
首先咱们把你给出的示例数据(包括期望输出里的对应行)转换成DataFrame:
import pandas as pd # 构建原始DataFrame data = { 'id': ['a', 'b', 'c', 'd', 'e'], 'text': [ 'simultaneous there the', 'simultaneous there', 'mul why', 'have the', 'then the late' ] } df = pd.DataFrame(data) # 指定要匹配的词汇列表 target_words = ["mul","the","have", "then"]
2. 写个小函数提取匹配的词汇
我们需要一个函数来检查每行文本里包含哪些目标词汇,然后把它们用逗号连起来:
def get_matched_words(text, target_list): # 把文本拆成单个单词(这里按空格拆分,要是有标点的话可以用正则优化) text_words = text.split() # 找出文本和目标列表的交集词汇 matched = [word for word in target_list if word in text_words] # 有匹配就返回逗号分隔的字符串,没有就返回空 return ', '.join(matched) if matched else ''
3. 应用函数创建列并筛选结果
接下来把这个函数应用到text列,生成found列,然后筛选出有匹配结果的行(要是不需要筛选,直接跳过筛选步骤就行):
# 创建found列 df['found'] = df['text'].apply(lambda x: get_matched_words(x, target_words)) # 筛选出有匹配的行 filtered_df = df[df['found'] != ''].reset_index(drop=True) # 把索引改成从1开始(和你的期望输出一致) filtered_df.index = filtered_df.index + 1
最终得到的结果
运行完上面的代码,filtered_df就是你想要的结果啦:
| id | text | found | |
|---|---|---|---|
| 1 | a | simultaneous there the | the |
| 2 | c | mul why | mul |
| 3 | d | have the | have, the |
| 4 | e | then the late | then, the |
一些小提示
- 如果你的文本里带有标点(比如"then,"这种),按空格拆分的方法会匹配不到,这时候可以用正则来提取纯单词,比如
import re之后用text_words = re.findall(r'\b\w+\b', text)来获取单词。 - 要是需要不区分大小写的匹配,就把文本和目标词汇都转成小写,比如
text_words = [word.lower() for word in text.split()],同时target_list = [word.lower() for word in target_list]。
内容的提问来源于stack exchange,提问作者TRINADH NAGUBADI
相关产品推荐
相关产品推荐

