基于Pandas DataFrame,依据列表列设置对应列值的技术需求
解决Pandas根据列表内容更新对应列值的问题
嘿,这里有两种实用的方法来实现你要的需求,从直观易懂到高效优化的都有,一起来看看:
方法1:逐行遍历处理(适合新手,逻辑清晰)
这种方法逻辑直接,很容易理解,适合小数据集:
首先先导入Pandas并创建你的DataFrame:
import pandas as pd test = {'Col1':[2,5], 'Col2':[5,7], 'Col_List':[['One','Two','Three','Four','Five'], ['Two', 'Four']], 'One':[0,0], 'Two':[0,0], 'Three':[0,0], 'Four':[0,0], 'Five':[0,0],} df=pd.DataFrame.from_dict(test)
然后执行更新操作:
# 先定义我们要操作的目标列(就是One到Five这几列) target_cols = ['One', 'Two', 'Three', 'Four', 'Five'] # 逐行遍历每一行数据 for idx, row in df.iterrows(): # 获取当前行Col_List里的列名列表 cols_to_update = row['Col_List'] # 把这些列对应的位置赋值为当前行Col1的数值 df.loc[idx, cols_to_update] = row['Col1']
运行后你就能得到想要的结果:
Col1 Col2 Col_List One Two Three Four Five 0 2 5 [One, Two, Three, Four, Five] 2 2 2 2 2 1 5 7 [Two, Four] 0 5 0 5 0
方法2:向量化操作(高效处理大数据集)
如果你的DataFrame行数很多,逐行遍历的效率会比较低,推荐用这种向量化的方式,速度快很多:
# 构建一个布尔掩码,标记哪些单元格需要被更新 mask = df['Col_List'].apply(lambda x: pd.Series(1, index=x)).reindex(columns=target_cols, fill_value=0).astype(bool) # 用Col1的值替换掩码中标记为True的位置,其他保持原值 df[target_cols] = df[target_cols].mask(mask, df['Col1'], axis=0)
简单解释下这个方法的逻辑:
df['Col_List'].apply(lambda x: pd.Series(1, index=x)):把每个Col_List里的元素转成Series的索引,值设为1,这样就对应了需要更新的列。reindex(columns=target_cols, fill_value=0):补全所有目标列,没有出现在Col_List里的列就填充0。astype(bool):把数值转成布尔值,这样就得到了一个"哪些位置要更新"的掩码。mask(mask, df['Col1'], axis=0):用Col1的值替换掩码中为True的位置,其他位置保持原来的0不变。
内容的提问来源于stack exchange,提问作者kopsman
相关产品推荐
相关产品推荐

