Python循环调用API合并数据集异常:仅返回首个数据集问题排查
问题原因与修复方案
你的代码只返回第一个数据集的核心问题是**return语句被放在了for循环内部**——第一次循环执行完毕后,函数就直接返回了结果,后续的API请求和数据拼接逻辑根本没有机会运行。
修复后的代码
import requests import pandas as pd # Define the APIs and category fields api_endpoints = [ "https://data.cityofchicago.org/resource/3pjd-kfu4.json", "https://data.cityofchicago.org/resource/7xqa-efsz.json", "https://data.cityofchicago.org/resource/anck-gptm.json" ] indicators = ["2009-2010", "2010-2011", "2011-2012"] # Stack data from API calls, indicate dataset year def stacky(api_endpoints, indicators): dataframes = [] for endpoint, indicator in zip(api_endpoints, indicators): response = requests.get(endpoint) data = response.json() df = pd.DataFrame(data) df['Category'] = indicator dataframes.append(df) # 所有API请求完成后再合并数据 combined = pd.concat(dataframes, ignore_index=True) return combined # Create the DataFrame using the function addresses = stacky(api_endpoints, indicators)
关键修改说明
- 将
combined = pd.concat(...)和return语句移到for循环外部,确保循环能完整遍历所有API端点,收集全部年份的数据集后再进行合并。 - 移除了不必要的循环内拼接操作,只在所有数据收集完成后执行一次合并,提升代码效率。
验证方法
你可以通过以下代码确认三个年份的数据都被正确合并:
print(addresses['Category'].value_counts())
内容的提问来源于stack exchange,提问作者Sandra T
相关产品推荐
相关产品推荐

