DataFrame相乘遇NaN问题:直接用数值正常,取数相乘返回NaN
Pandas Series相乘返回全NaN问题排查及解决
问题场景
正常情况
直接使用数值变量与impact_df['Land Use']相乘,结果正常:
c = -10245.945396 print(impact_df['Land Use'].mul(c))
输出:
2015 -1.711073e+06 2016 -1.711073e+06 2017 -1.711073e+06 2018 -1.711073e+06 2019 -1.711073e+06 Name: Land Use, dtype: float64
异常情况
从country_df筛选得到的coefficients与impact_df['Land Use']相乘,结果全为NaN:
coefficients = country_df[country_df['GeoRegion'] == region]['Coefficients'] print(coefficients) print(impact_df['Land Use'].mul(coefficients))
输出:
7150 -9649.082503 Name: Coefficients, dtype: float64 2015 NaN 2016 NaN 2017 NaN 2018 NaN 2019 NaN 7150 NaN dtype: float64
核心原因
Pandas的mul()方法默认会按索引对齐数据:
impact_df['Land Use']的索引是年份2015-2019- 筛选出的
coefficients是一个带索引7150的Series,不是单纯的数值
两者索引完全不匹配,相乘时每个位置都找不到对应索引的匹配值,所以返回全NaN。而直接用数值变量相乘时,Pandas会把数值广播到整个Series,不需要索引对齐,因此结果正常。
解决方案
只需提取coefficients中的实际数值,而非保留Series结构,两种常用方式:
方式1:用.iloc[0]提取首个元素(确保筛选结果仅一行)
coefficients = country_df[country_df['GeoRegion'] == region]['Coefficients'].iloc[0] print(impact_df['Land Use'].mul(coefficients))
方式2:用.values[0]获取数值数组的首个元素
coefficients = country_df[country_df['GeoRegion'] == region]['Coefficients'].values[0] print(impact_df['Land Use'].mul(coefficients))
执行后即可得到正常的计算结果,不会再出现NaN。
内容的提问来源于stack exchange,提问作者Dima
相关产品推荐
相关产品推荐

