如何按条件更新pandas DataFrame列表列并保留元素唯一性
实现方案
原代码直接调用extend()方法会把新列表所有元素都追加到原有列表末尾,不会做重复校验,因此会出现重复元素,可根据是否需要保留元素顺序选择以下两种实现方案:
方案1:保留原有元素顺序(推荐)
原有列表的元素顺序不变,仅追加新列表中不存在于原列表的元素,符合大多数场景需求,完整可运行代码如下:
import pandas as pd df = pd.DataFrame(columns=['col1', 'col2']) print(df) col1 = 'A' # 注:pandas 2.0+ 已弃用append方法,可替换为pd.concat写法 df = df.append({'col1': col1, 'col2': ['a1','a2', 'a3']}, ignore_index=True) print(df) new_list = ['a4', 'a5','a1'] # 取出目标列表 target_list = df.loc[df['col1'] == col1,'col2'].iloc[0] # 筛选出新列表中不存在于原列表的元素,再追加 unique_add = [item for item in new_list if item not in target_list] target_list.extend(unique_add) print(df)
输出结果和期望完全一致,且原有a1、a2、a3的顺序保持不变,新增元素按a4、a5的顺序追加。
方案2:无需保留顺序,实现更简洁
如果不需要保留元素的原有顺序,可以直接用集合去重,代码更简短:
import pandas as pd df = pd.DataFrame(columns=['col1', 'col2']) print(df) col1 = 'A' df = df.append({'col1': col1, 'col2': ['a1','a2', 'a3']}, ignore_index=True) print(df) new_list = ['a4', 'a5','a1'] # 集合合并去重后重新赋值 target_list = df.loc[df['col1'] == col1,'col2'].iloc[0] df.loc[df['col1'] == col1,'col2'] = [list(set(target_list + new_list))] print(df)
该方案会打乱原有列表的元素顺序,仅适合对顺序无要求的场景。
内容的提问来源于stack exchange,提问作者burcak
相关产品推荐
相关产品推荐

