API JSON转DataFrame遇NotImplementedError,求提取指定列方案
解决JSON展平为DataFrame的问题
首先,修正代码逻辑,直接用response.json()解析返回数据更简洁。针对API返回的深层嵌套结构,需要手动遍历层级,将每个温度值与对应的站点、时间信息一一配对,避免json_normalize因复杂嵌套抛出的错误。
完整代码如下:
import requests import pandas as pd # 请求API并解析JSON数据 response_API = requests.get('https://dwd.api.proxy.bund.dev/v30/stationOverviewExtended?stationIds=10865,G005') data = response_API.json() # 初始化存储数据的列表 rows = [] # 遍历每个站点 for station in data['stations']: station_id = station['stationId'] # 遍历该站点的时间序列条目 for ts in station['timeSeries']: start_time = ts['start'] timestep = ts['timestep'] # 提取温度值数组 temp_values = ts['parameters']['temperature']['values'] # 为每个温度值生成一行关联数据 for temp in temp_values: rows.append({ 'stationId': station_id, 'start': start_time, 'timestep': timestep, 'temperature': temp }) # 转换为目标DataFrame df = pd.DataFrame(rows) print(df.head())
关键说明
- API返回结构中,
temperature是数组形式,需要将数组中的每个值与上层的stationId、start、timestep字段绑定,手动遍历嵌套结构能精准实现这种关联 - 直接使用
json_normalize会因多层嵌套的数组结构触发NotImplementedError,手动遍历的方式更灵活可控
内容的提问来源于stack exchange,提问作者Chhandosee Bhattacharya
相关产品推荐
相关产品推荐

