如何在Pandas DataFrame中保留指定字符串集合
问题描述
我有一个包含特定列的DataFrame,结构如下:
colA ['work', 'time', 'money', 'home', 'good', 'financial'] ['school', 'lazy', 'good', 'math', 'sad', 'important', 'dizzy', 'go'] ['frame', 'happy', 'feel', 'youth', 'change', 'home', 'past'] ['first', 'eat', 'good', 'hungry', 'empty', 'fool'] ['meet', 'risk', 'fire', 'angry', 'go']
其中colA是字符串类型(而非列表)。我还有一个单词列表:
word = ['good', 'sad', 'angry', 'feel', 'empty', 'dizzy', 'go', 'happy', 'fool', 'eat', 'past', 'lazy', 'youth', 'old', 'enjoy', 'free', 'time', 'hungry']
我希望保留colA中属于该列表的单词,处理后结果应如下:
colA ['time', 'good'] ['lazy', 'good', 'sad', 'dizzy', 'go'] ['happy', 'feel', 'youth', 'past'] ['eat', 'good', 'hungry', 'empty', 'fool'] ['angry', 'go']
我尝试使用str.contains方法,但出现报错:
contains() takes from 2 to 6 positional arguments but 18 were given
解决方案
错误原因
str.contains()的第一个参数需要是单个正则表达式字符串,你直接把18个单词作为位置参数传入,不符合函数的参数要求,所以抛出错误。而且str.contains()是检查子串匹配,无法满足你精确保留单词的需求,因此需要换一种处理方式。
步骤说明
- 将字符串形式的列表转为真实列表:因为
colA的元素是字符串格式的列表(比如"['work', 'time', ...]"),需要用ast.literal_eval()安全转换为Python列表。 - 用集合存储目标单词:集合的成员查询效率远高于列表,适合批量匹配。
- 过滤并保留符合条件的单词:对每一行的列表进行过滤,只保留在目标单词集合中的元素。
完整代码示例
import pandas as pd import ast # 构造示例DataFrame df = pd.DataFrame({ 'colA': [ "['work', 'time', 'money', 'home', 'good', 'financial']", "['school', 'lazy', 'good', 'math', 'sad', 'important', 'dizzy', 'go']", "['frame', 'happy', 'feel', 'youth', 'change', 'home', 'past']", "['first', 'eat', 'good', 'hungry', 'empty', 'fool']", "['meet', 'risk', 'fire', 'angry', 'go']" ] }) # 目标单词列表,转为集合提升查询效率 word_list = ['good', 'sad', 'angry', 'feel', 'empty', 'dizzy', 'go', 'happy', 'fool', 'eat', 'past', 'lazy', 'youth', 'old', 'enjoy', 'free', 'time', 'hungry'] word_set = set(word_list) # 处理colA:转列表→过滤→保留结果 df['colA'] = df['colA'].apply(lambda x: [word for word in ast.literal_eval(x) if word in word_set]) # 查看结果 print(df)
运行结果
colA 0 [time, good] 1 [lazy, good, sad, dizzy, go] 2 [happy, feel, youth, past] 3 [eat, good, hungry, empty, fool] 4 [angry, go]
内容的提问来源于stack exchange,提问作者andryan86
相关产品推荐
相关产品推荐

