如何将Pandas中新增列的for循环改写为单行代码?
问题
我有一个DataFrame df,可通过以下代码生成:
import pandas as pd data = [10,20,30,40,50,60] df = pd.DataFrame(data, columns=['Numbers'])
现在需要检查给定列表中的列名是否存在于df中,若不存在则创建同名新列并将值设为0,原循环代码如下:
columns_list=["3","5","8","9","12"] for i in columns_list: if i not in df.columns.to_list(): df[i]=0
我尝试将其改写为单行代码:
[df[i]=0 for i in columns_list if i not in df.columns.to_list()]
但IDE返回语法错误:
SyntaxError: cannot assign to subscript here. Maybe you meant '==' instead of '='?
请问该如何正确实现?
解决方案
报错原因
列表推导式的核心是生成列表元素,它内部不能包含赋值语句(=)——赋值属于副作用操作,并非用来生成列表项,这就是你收到语法错误的直接原因。
方法1:使用Pandas原生assign()方法(推荐)
这是最符合Pandas风格的写法,通过字典推导生成需要添加的列与对应值,再用assign()批量处理:
# 生成新的DataFrame(不修改原df) df_new = df.assign(**{col: 0 for col in columns_list if col not in df.columns}) # 如果需要原地修改原df,直接重新赋值即可 df = df.assign(**{col: 0 for col in columns_list if col not in df.columns})
方法2:单行循环写法(不推荐)
如果一定要压缩成单行循环形式,可直接简化原循环逻辑,但这种写法可读性较差:
for col in columns_list: df[col] = 0 if col not in df.columns else df[col]
也可以利用__setitem__方法在列表推导中执行赋值,但会生成一个无意义的None列表,不推荐使用:
[df.__setitem__(col, 0) for col in columns_list if col not in df.columns]
方法3:高效列存在性检查(适用于大列表)
若columns_list元素较多,建议先将df.columns转为集合——集合的成员检查速度远快于列表:
existing_cols = set(df.columns) df = df.assign(**{col: 0 for col in columns_list if col not in existing_cols})
内容的提问来源于stack exchange,提问作者William
相关产品推荐
相关产品推荐

