You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas DataFrame中保留指定字符串集合

问题描述

我有一个包含特定列的DataFrame,结构如下:

colA    
['work', 'time', 'money', 'home', 'good', 'financial']    
['school', 'lazy', 'good', 'math', 'sad', 'important', 'dizzy', 'go']    
['frame', 'happy', 'feel', 'youth', 'change', 'home', 'past']    
['first', 'eat', 'good', 'hungry', 'empty', 'fool']    
['meet', 'risk', 'fire', 'angry', 'go']    

其中colA是字符串类型(而非列表)。我还有一个单词列表:

word = ['good', 'sad', 'angry', 'feel', 'empty', 'dizzy', 'go', 'happy', 'fool', 'eat', 'past', 'lazy', 'youth', 'old', 'enjoy', 'free', 'time', 'hungry']   

我希望保留colA中属于该列表的单词,处理后结果应如下:

colA    
['time', 'good']    
['lazy', 'good', 'sad', 'dizzy', 'go']    
['happy', 'feel', 'youth', 'past']     
['eat', 'good', 'hungry', 'empty', 'fool']    
['angry', 'go']    

我尝试使用str.contains方法,但出现报错:

contains() takes from 2 to 6 positional arguments but 18 were given    

解决方案

错误原因

str.contains()的第一个参数需要是单个正则表达式字符串,你直接把18个单词作为位置参数传入,不符合函数的参数要求,所以抛出错误。而且str.contains()是检查子串匹配,无法满足你精确保留单词的需求,因此需要换一种处理方式。

步骤说明

  1. 将字符串形式的列表转为真实列表:因为colA的元素是字符串格式的列表(比如"['work', 'time', ...]"),需要用ast.literal_eval()安全转换为Python列表。
  2. 用集合存储目标单词:集合的成员查询效率远高于列表,适合批量匹配。
  3. 过滤并保留符合条件的单词:对每一行的列表进行过滤,只保留在目标单词集合中的元素。

完整代码示例

import pandas as pd
import ast

# 构造示例DataFrame
df = pd.DataFrame({
    'colA': [
        "['work', 'time', 'money', 'home', 'good', 'financial']",
        "['school', 'lazy', 'good', 'math', 'sad', 'important', 'dizzy', 'go']",
        "['frame', 'happy', 'feel', 'youth', 'change', 'home', 'past']",
        "['first', 'eat', 'good', 'hungry', 'empty', 'fool']",
        "['meet', 'risk', 'fire', 'angry', 'go']"
    ]
})

# 目标单词列表,转为集合提升查询效率
word_list = ['good', 'sad', 'angry', 'feel', 'empty', 'dizzy', 'go', 'happy', 'fool', 'eat', 'past', 'lazy', 'youth', 'old', 'enjoy', 'free', 'time', 'hungry']
word_set = set(word_list)

# 处理colA:转列表→过滤→保留结果
df['colA'] = df['colA'].apply(lambda x: [word for word in ast.literal_eval(x) if word in word_set])

# 查看结果
print(df)

运行结果

colA
0              [time, good]
1  [lazy, good, sad, dizzy, go]
2      [happy, feel, youth, past]
3  [eat, good, hungry, empty, fool]
4                [angry, go]

内容的提问来源于stack exchange,提问作者andryan86

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 18:10:32