处理预期寿命CSV数据集遇ValueError,求代码纠错与功能实现
问题:处理预期寿命数据集时触发ValueError错误
运行代码时出现如下错误:
ValueError: invalid literal for int() with base 10: 'Year'
需求是从指定CSV文件提取并输出整体最高预期寿命,同时实现年份查询统计功能,但当前代码无法正常运行,求错误原因及修正方案。
当前代码:
import math import csv print() with open("life-expectancy.csv") as data_set: for row in data_set: row = row.split(",") entity = row[0].strip() code = row[1].strip() year = int(row[2]) life_expectancy = (row[3]) max_life = -1 min_life = min(life_expectancy) avg_life = sum(life_expectancy) / len(life_expectancy) max_country = "" max_year = max(year) min_country = min(entity) min_year = min(year) chosen_year = "" print() year_lookup = input(float("Enter the year of interest: ")) print() if life_expectancy > max_life: max_life = life_expectancy max_country = entity print() print(f"The overall max life expectancy is:{max_life:.2f} from {max_country} in {max_year:.2f}.\n") print(f"The overall min life expectancy is:{min_life:.2f} from {min_country} in {min_year:.2f}.\n") if chosen_year == year_lookup: chosen_year = year_lookup print(f"For the year: {year_lookup}.\n") print(f"The average life expectancy across all countries was str{avg_life:.2f}\n") print(f"The max life expectancy was in {max_country} with {max_life:.2f}\n") print(f"The min life expectancy was in {min_country} with {min_life:.2f}\n")
错误原因拆解
- 未跳过CSV表头:CSV第一行是表头(含"Year"字符串),直接转int必然报错,这是触发
ValueError的直接原因。 - 变量初始化逻辑错误:
max_life、max_country等统计变量在循环内重复初始化,每次循环都会重置,根本没法累积统计结果。 - 数值类型未转换:
life_expectancy保留成字符串类型,后续的比较、求和操作全是对字符串的无效操作。 - 统计逻辑混乱:
min(life_expectancy)、max(year)这类代码完全逻辑错误——前者是对单个字符串求最小,后者是对单个年份数值求最大,毫无意义。 - 输入处理语法错误:
input(float("Enter..."))写法错误,input的参数只能是提示字符串,不能直接套float转换。 - 代码结构混乱:统计逻辑和输出逻辑位置完全错误,循环内没累积数据,循环外的判断和输出根本拿不到正确结果。
修正后的代码
import csv # 初始化全局统计变量 overall_max_life = -1.0 overall_max_country = "" overall_max_year = 0 overall_min_life = float('inf') overall_min_country = "" overall_min_year = 0 all_life_expectancies = [] year_data = {} # 按年份存储数据:键=年份,值=(寿命列表, 对应国家列表) # 读取并处理CSV数据 with open("life-expectancy.csv") as data_set: reader = csv.reader(data_set) next(reader) # 跳过表头行 for row in reader: entity = row[0].strip() code = row[1].strip() year = int(row[2]) life_expectancy = float(row[3].strip()) # 更新全局最高/最低寿命 if life_expectancy > overall_max_life: overall_max_life = life_expectancy overall_max_country = entity overall_max_year = year if life_expectancy < overall_min_life: overall_min_life = life_expectancy overall_min_country = entity overall_min_year = year # 收集所有寿命数据用于计算全局平均 all_life_expectancies.append(life_expectancy) # 按年份分组存储数据 if year not in year_data: year_data[year] = ([], []) year_data[year][0].append(life_expectancy) year_data[year][1].append(entity) # 输出全局统计结果 print("=== 全局预期寿命统计 ===") print(f"最高预期寿命: {overall_max_life:.2f},来自{overall_max_country},年份{overall_max_year}") print(f"最低预期寿命: {overall_min_life:.2f},来自{overall_min_country},年份{overall_min_year}") print(f"所有国家平均预期寿命: {sum(all_life_expectancies)/len(all_life_expectancies):.2f}\n") # 处理年份查询(带输入验证) while True: try: year_lookup = int(input("请输入要查询的年份: ")) break except ValueError: print("请输入有效的整数年份!") if year_lookup in year_data: life_list, country_list = year_data[year_lookup] year_avg = sum(life_list)/len(life_list) year_max = max(life_list) year_max_country = country_list[life_list.index(year_max)] year_min = min(life_list) year_min_country = country_list[life_list.index(year_min)] print(f"\n=== {year_lookup}年预期寿命统计 ===") print(f"该年所有国家平均预期寿命: {year_avg:.2f}") print(f"该年最高预期寿命: {year_max:.2f},来自{year_max_country}") print(f"该年最低预期寿命: {year_min:.2f},来自{year_min_country}") else: print(f"\n未找到{year_lookup}年的数据!")
修正说明
- 跳过表头:用
next(reader)跳过CSV第一行表头,避免字符串转整数错误。 - 正确初始化变量:把统计变量放在循环外,确保循环中能累积更新数据。
- 数值类型转换:将
life_expectancy转为float,保证后续数值计算正确。 - 分组统计:用字典按年份存储对应数据,方便后续年份查询。
- 输入验证:添加异常处理,确保用户输入有效的整数年份。
- 清晰结构:拆分数据读取、全局统计、年份查询三个模块,逻辑更清晰。
内容的提问来源于stack exchange,提问作者RP116
相关产品推荐
相关产品推荐

