API转Dataframe问题:如何将接口返回数据转为列格式
问题分析
你的问题出在API返回数据的结构处理上:API返回的是一个包含两个元素的列表,第一个元素是IP监控数据的列表,第二个是元数据(Meta)和分页链接(Links)的字典。而你直接把整个返回结果追加到response_all中,导致response_all里混合了数据列表和元数据字典,最终生成DataFrame时,这些嵌套结构被当成了单个单元格的JSON内容,而非拆分后的列。
解决方案
修改函数逻辑,只提取API返回中的数据部分(第一个元素)进行收集,元数据可单独存储或按需忽略。同时优化异常处理和循环终止条件:
修正后的代码
def get_ip_list(self): response_all = [] while True: url = "www.apitest.com/testing/" response = requests.request("GET", url) try: if response.status_code != 200: self.logger.log_info(f"API请求出错,状态码: {response.status_code}") break # 非200状态直接终止循环,避免无效处理 response_json = response.json() # 检查是否无数据,终止循环 if isinstance(response_json, list) and len(response_json) > 0: # 提取数据部分(第一个元素) data_part = response_json[0] # 检查数据部分是否为空或提示无数据 if isinstance(data_part, list) and not data_part: break if isinstance(data_part, str) and 'no blacklist monitors found in account' in data_part: break # 只追加数据条目到response_all response_all.extend(data_part) else: break # 处理分页:根据API返回的分页链接判断是否继续请求 meta_part = response_json[1] if len(response_json) > 1 else {} next_page_url = meta_part.get('Links', {}).get('Pages', {}).get('Next') if not next_page_url: break url = next_page_url self.counter += 1 except Exception as e: self.logger.log_info(f"处理API响应出错: {str(e)}") break # 生成标准列结构的DataFrame df = pd.DataFrame.from_records(response_all) # 可选:展开嵌套的Links字段为单独列 if not df.empty and 'Links' in df.columns: links_df = df['Links'].apply(pd.Series) df = pd.concat([df.drop('Links', axis=1), links_df], axis=1) print(df.head()) return df
关键改动说明
- 只收集有效数据:每次API返回后,提取第一个元素(数据列表)追加到
response_all,跳过元数据部分。 - 优化循环终止条件:不仅检查无数据提示,还判断数据部分是否为空,同时根据分页链接判断是否继续请求。
- 可选展开嵌套字段:如果需要把
Links里的Report_Link和Whitelabel_Report_Link拆成单独列,添加对应的处理逻辑。 - 增强错误处理:非200状态直接终止循环,并记录错误信息。
验证结果
修正后,response_all会是一个纯IP监控字典的列表,用pd.DataFrame.from_records生成的DataFrame会自动将每个字典的键作为列名,值作为对应行的内容,不再出现JSON格式的嵌套列。
内容的提问来源于stack exchange,提问作者Caroline Leite
相关产品推荐
相关产品推荐

