You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

处理美国人口普查数据遇ValueError:带逗号字符串无法转整数

问题:人口数据转换整数时触发ValueError错误

从美国人口普查局下载州人口数据后,使用pandas创建DataFrame,需完成全美各州统计汇总并为任意州绘制折线图,但反复触发错误:

ValueError: invalid literal for int() with base 10: '989,415'

当前代码

import numpy as np
import pandas as pd  # 原代码遗漏pandas导入
import matplotlib as mpl
import matplotlib.pyplot as plt

# 打开文件
with open('/Users/Documents/population data.csv', 'r') as data_file:
      
     # 创建DataFrame并读取文件
     df= pd.read_csv(data_file)

     # 重命名/清理表头
     df= df.set_axis(['state', '4/1/2010 census population', 
                  '4/1/2010', 
                  '7/1/2010', 
                  '7/1/2011',
                  '7/1/2012',
                  '7/1/2013',
                  '7/1/2014', 
                  '7/1/2015', 
                  '7/1/2016', 
                  '7/1/2017', 
                  '7/1/2018',
                  '7/1/2019'],  axis= 'columns')
     
     # 以state列为索引
     df.set_index('state', inplace= True)
     
     # 以下为已尝试的转换方案(均失败)
     # df["4/1/2010 census population"] = df["4/1/2010 census population"].astype("int")
     # df.astype({"4/1/2010 census population":'int', "4/1/2010":'int'}) 
     # df['4/1/2010 census population']= df['4/1/2010 census population'].apply(np.int64)
     # pd.to_numeric(df)

问题原因

错误提示中的'989,415'是带千分位逗号的字符串,Python无法直接将包含逗号的字符串转换为整数,必须先清除逗号格式。

解决方案

方案1:读取CSV时直接处理千分位逗号

在pd.read_csv中添加thousands=','参数,读取时自动将带逗号的字符串转为数值类型:

import numpy as np
import pandas as pd
import matplotlib.pyplot as plt

# 读取时直接解析千分位逗号
df = pd.read_csv('/Users/Documents/population data.csv', thousands=',')

# 重命名表头
df = df.set_axis(['state', '4/1/2010 census population', 
                  '4/1/2010', '7/1/2010', '7/1/2011',
                  '7/1/2012', '7/1/2013', '7/1/2014', 
                  '7/1/2015', '7/1/2016', '7/1/2017', 
                  '7/1/2018', '7/1/2019'], axis='columns')

df.set_index('state', inplace=True)

方案2:已读取数据后批量处理列

如果已经读取了原始数据,可遍历数值列清除逗号后转换类型:

# 遍历所有数值列(除state索引外)
for col in df.columns:
    df[col] = df[col].str.replace(',', '').astype(int)

后续操作示例

1. 全美各州统计汇总

# 计算各年份全美总人口
total_population = df.sum()
print(total_population)

# 生成各州人口统计量(均值、最值等)
state_pop_stats = df.describe()
print(state_pop_stats)

2. 绘制任意州人口折线图

# 以加利福尼亚州为例
target_state = 'California'
state_pop_trend = df.loc[target_state]

# 绘制折线图
plt.figure(figsize=(10,6))
state_pop_trend.plot(kind='line', marker='o')
plt.title(f'{target_state} Population Trend (2010-2019)')
plt.xlabel('Year')
plt.ylabel('Population')
plt.grid(True)
plt.show()

内容的提问来源于stack exchange,提问作者bbb12543

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 06:31:15