You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas中year列(int)自动转为float的问题及绘图异常求助

问题解决:Pandas中year列变为float类型导致绘图显示小数

核心问题

处理Pandas DataFrame时,原本为int类型的year列自动转为float类型,绘图时出现不必要的小数。添加通过gdp和population计算得到的gdp_per_capita后出现该问题,尝试转int时报错:

IntCastingNaNError: Cannot convert non-finite values (NA or inf) to integer

检查发现year列末尾存在行标签为gdp_per_capita的NaN值:

100               2000.0
101               2001.0
...
17431             2019.0
gdp_per_capita       NaN
Name: year, Length: 4854, dtype: float64

错误原因

代码中添加gdp_per_capita的方式错误:

d.loc['gdp_per_capita']=d['gdp']/d['population']

这里的.loc['gdp_per_capita']是添加了一行而非一列,导致year列新增了一个行索引为gdp_per_capita的NaN值(因为该行没有对应的year数据)。而Pandas的int类型无法存储NaN,所以year列自动转为float类型。

解决步骤

1. 纠正新列的添加方式

把添加行改为添加列:

d['gdp_per_capita'] = d['gdp'] / d['population']

2. 清理数据并转换year列类型

先处理计算产生的inf/nan,再转换year为int类型:

# 替换无穷值为NaN,删除含NaN的行(或根据需求填充)
d.replace([np.inf, -np.inf], np.nan, inplace=True)
d = d.dropna(subset=['year'])  # 确保year列无缺失值

# 转换year为int类型
d['year'] = d['year'].astype(int)

3. 调整绘图代码

分组后的数据需要重置索引,确保year作为列参与绘图:

e = d.loc[d['country'].isin(regions)]
e = e.groupby(['country', 'year']).aggregate({'energy_cons_change_pct': 'mean'}).rename(columns={'energy_cons_change_pct': 'arccp'})
e = e.reset_index()  # 重置索引,让year成为普通列

plt.xlim(2000, 2020)
sns.lineplot(x='year', y='arccp', data=e, hue='country')
plt.show()

完整修正后的代码

import pandas as pd
import numpy as np
import seaborn as sns
import matplotlib.pyplot as plt

d = pd.read_csv('/kaggle/input/world-energy-consumption/World Energy Consumption.csv')
# 注意:原代码中41-42是计算值,需确认对应列索引正确性,此处保留原写法
d = d.iloc[:, [1, 2, 9, 13, 20, 43, 51, 60, 69, 75, 82, 108, 106, 92, 98, 114, 102, 41-42]]
d = d[d['year'] >= 2000]
d = d.fillna(method='ffill').fillna(method='bfill')

# 正确添加新列
d['gdp_per_capita'] = d['gdp'] / d['population']

d.replace([np.inf, -np.inf], np.nan, inplace=True)
d = d.dropna(subset=['year'])
d['year'] = d['year'].astype(int)

# 绘图部分
regions=['Asia Pacific', 'CIS', 'World', 'Central America', 'Eastern Africa', 'Europe', 'Europe (other)','North America', 'Other Asia & Pacific', 'Other CIS', 'Other Caribbean',
 'Other Middle East', 'Other Northern Africa', 'Other South America', 'Other Southern Africa',  'South & Central America',
 'South Africa', 'World']

e = d.loc[d['country'].isin(regions)]
e = e.groupby(['country', 'year']).aggregate({'energy_cons_change_pct': 'mean'}).rename(columns={'energy_cons_change_pct': 'arccp'})
e = e.reset_index()

plt.xlim(2000, 2020)
sns.lineplot(x='year', y='arccp', data=e, hue='country')
plt.show()

额外说明

  • 原代码中d['year'].dropna只是调用方法但未执行,需用d = d.dropna(subset=['year'])来实际删除缺失值
  • 分组后调用reset_index()可以让year从索引转为普通列,确保绘图时识别为离散的年份值而非连续数值

内容的提问来源于stack exchange,提问作者Luca

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 21:27:12