如何在DataFrame列中移除对应列表中存在的子字符串?
解决方法
要实现将string列中属于对应lists列的单词移除并生成new_string列,可以通过逐行过滤单词的方式实现,具体步骤如下:
1. 构造示例数据(可选,用于复现)
import pandas as pd data = { 'string': ['i have a dog', 'there is a cat', 'hello everyone', 'hi my name is Joe'], 'lists': [['fox', 'dog', 'cat'], ['dog', 'house', 'car'], ['hi', 'hello', 'everyone'], ['name', 'was', 'Joe']] } df1 = pd.DataFrame(data)
2. 定义过滤函数
编写一个函数,接收每行数据,拆分string为单词列表后,过滤掉存在于对应lists中的单词,最后拼接剩余单词:
def filter_words(row): # 拆分字符串为单词 word_list = row['string'].split() # 过滤掉在当前行lists中的单词 filtered = [word for word in word_list if word not in row['lists']] # 拼接为字符串,无剩余单词则返回空 return ' '.join(filtered) if filtered else ''
3. 生成new_string列
使用apply(axis=1)逐行应用函数,生成目标列:
df1['new_string'] = df1.apply(filter_words, axis=1)
执行后得到的结果与你期望的df2完全一致:
string lists new_string 0 i have a dog ['fox', 'dog', 'cat'] i have a 1 there is a cat ['dog', 'house', 'car'] there is a cat 2 hello everyone ['hi', 'hello', 'everyone'] 3 hi my name is Joe ['name', 'was', 'Joe'] hi my is
关键说明
因为每行的lists是独立的动态列表,无法用固定列表的批量过滤方法,必须通过apply(axis=1)逐行获取当前行的lists来过滤对应string的单词,这是解决问题的核心。
内容的提问来源于stack exchange,提问作者mjp
相关产品推荐
相关产品推荐

