使用map提取Pandas DataFrame列时二次调用返回空结果的原因
问题:使用map生成列名后,重复调用返回空列表
我用map处理DataFrame列名,想提取带_a后缀的列,以下是复现代码:
import pandas as pd df = pd.DataFrame() df['Col1'] = [197, 1600, 1200] df['Col2'] = [297, 2600, 2200] df['Col1_a'] = [198, 1599, 1199] df['Col2_a'] = [296, 2599, 2199] print(df)
输出:
Col1 Col2 Col1_a Col2_a 0 197 297 198 296 1 1600 2600 1599 2599 2 1200 2200 1199 2199
我用下面的方式生成目标列名:
list_col = ["Col1","Col2"] cols_w_suffix = map(lambda x: x + '_a', list_col) # 第一次调用正常返回结果 print(df[cols_w_suffix].to_dict('records'))
得到预期输出:[{'Col1_a': 198, 'Col2_a': 296}, {'Col1_a': 1599, 'Col2_a': 2599}, {'Col1_a': 1199, 'Col2_a': 2199}]
但再次执行相同语句时,返回空列表:
print(df[cols_w_suffix].to_dict('records'))
输出:[]
而直接传入列名列表时,每次调用都能得到正确结果:
df[["Col1_a","Col2_a"]].to_dict('records')
输出:[{'Col1_a': 198, 'Col2_a': 296}, {'Col1_a': 1599, 'Col2_a': 2599}, {'Col1_a': 1199, 'Col2_a': 2199}]
原因分析
问题出在map返回的对象特性上:Python中map()函数返回的是迭代器(iterator),而非列表这种可重复遍历的容器。
迭代器的核心特点是只能被遍历一次——第一次通过df[cols_w_suffix]遍历这个迭代器时,它内部的指针已经移动到序列末尾,再次遍历就没有元素可以返回了,Pandas找不到对应列,最终返回空列表。
而直接用列表["Col1_a","Col2_a"]时,列表是可重复遍历的可迭代对象,每次调用都会重新遍历所有元素,因此每次都能正确匹配列名。
解决方法
把map返回的迭代器转换成列表,就能多次复用了:
list_col = ["Col1","Col2"] # 用list()将map对象转为列表 cols_w_suffix = list(map(lambda x: x + '_a', list_col)) # 现在无论调用多少次都能得到正确结果 print(df[cols_w_suffix].to_dict('records')) print(df[cols_w_suffix].to_dict('records'))
这样两次调用都会返回预期的字典列表。
内容的提问来源于stack exchange,提问作者honeybadger
相关产品推荐
相关产品推荐

