Python调用世界银行API生成DataFrame时IndexError报错的代码修正求助
解决IndexError及代码优化建议
看起来你的代码触发IndexError: list index out of range的核心原因是API请求的URL格式错误,导致返回的JSON数据没有你期望的索引为1的元素(即数据部分)。我们一步步来修复并优化代码:
1. 修复URL构造错误
你的wbURL函数有两个关键问题:
countries{contcode}缺少了斜杠,应该是countries/{contcode},否则API无法正确识别国家代码参数- URL里的
&是HTML转义字符,API服务无法解析,需要直接用&
修正后的wbURL函数:
def wbURL(country_code, indicator, begin_year, end_year): # 调整参数名更清晰,同时修正URL格式 return f'http://api.worldbank.org/countries/{country_code}/indicators/{indicator}?format=json&date={begin_year}:{end_year}'
2. 优化wbDF函数,避免索引错误并简化逻辑
原函数的问题包括:
- 没有检查API响应是否成功,若请求失败(比如URL错误),返回的JSON只有错误信息(仅1个元素),此时
wb[1]就会触发索引错误 - 重复遍历
data列表来构造列,效率低且冗余 - 手动重新构造所有列,忽略了原DataFrame已有的字段
修正并优化后的wbDF函数:
import requests import json import pandas as pd def wbDF(country_code, indicator, begin_year, end_year): url = wbURL(country_code, indicator, begin_year, end_year) response = requests.get(url) # 先检查请求是否成功 if response.status_code != 200: raise Exception(f"API请求失败,状态码:{response.status_code}") wb = json.loads(response.content) # 检查返回的JSON是否包含数据部分(索引1的元素) if len(wb) < 2: raise Exception("API未返回有效数据,请检查参数是否正确") data = wb[1] df = pd.DataFrame(data) # 只保留需要的列 df = df[['indicator', 'country', 'date', 'value']] # 提取字典中的对应值 df['indicator'] = df['indicator'].apply(lambda x: x['id']) df['country'] = df['country'].apply(lambda x: x['value']) # value列本身就是值(可能为None),无需额外处理 return df
3. 测试代码
现在运行你的测试代码:
test = wbDF('GBR', 'SP.DYN.LE00.IN', 2000, 2019) print(test)
应该能正常返回包含英国2000-2019年预期寿命数据的DataFrame了。
额外建议
- 可以给函数参数添加类型提示,让代码更易读
- 处理
value列的空值(比如用df['value'] = df['value'].fillna(0)或者根据需求处理) - 可以添加
per_page参数到URL中,默认获取更多数据(World Bank API默认返回100条,若数据超过100条需要分页)
内容的提问来源于stack exchange,提问作者Qian Xiaotong
相关产品推荐
相关产品推荐

