使用pandas apply返回多列时遇ValueError:无法解包过多值
问题解决:Pandas Apply生成两列触发ValueError的排查
报错原因
你用apply(axis=1)返回元组时,默认会生成一个每个元素都是元组的Series,直接用df_test['a'], df_test['b']去解包这个Series,相当于要把Series里的5个元组分配给2个变量,自然触发ValueError: too many values to unpack (expected 2)。
两种可行修复方案
方案1:让函数返回Series
修改calc_test函数,使其返回pd.Series,这样apply会直接生成DataFrame,可直接赋值给新列:
import pandas as pd def calc_test(row): a = row['col1'] + row['col2'] b = row['col1'] / row['col2'] # 返回带列名的Series return pd.Series([a, b], index=['a', 'b']) df_test_dict={'col1':[1,2,3,4,5],'col2':[10,20,30,40,50]} df_test=pd.DataFrame(df_test_dict) # 直接赋值给多列 df_test[['a', 'b']] = df_test.apply(calc_test, axis=1) print(df_test)
方案2:使用result_type='expand'参数
在apply中指定result_type='expand',让Pandas自动把元组展开成多列:
import pandas as pd def calc_test(row): a = row['col1'] + row['col2'] b = row['col1'] / row['col2'] return (a,b) df_test_dict={'col1':[1,2,3,4,5],'col2':[10,20,30,40,50]} df_test=pd.DataFrame(df_test_dict) # 添加result_type参数展开元组 df_test[['a', 'b']] = df_test.apply(lambda row: calc_test(row), axis=1, result_type='expand') print(df_test)
两种方案执行后都会得到正确结果:
col1 col2 a b 0 1 10 11 0.1 1 2 20 22 0.1 2 3 30 33 0.1 3 4 40 44 0.1 4 5 50 55 0.1
内容的提问来源于stack exchange,提问作者roudan
相关产品推荐
相关产品推荐

