API分页循环完成后如何合并多个DataFrame?技术求助
解决API分页后合并多个DataFrame的问题
你只需在循环中收集每个页面生成的DataFrame,最后用pd.concat()合并即可,具体修改如下:
修改后的完整代码
import requests import pandas as pd from pandas import json_normalize url = 'https://vendors.paddle.com/api/2.0/subscription/users' headers = { 'Content-Type': 'application/x-www-form-urlencoded', 'User-Agent': 'PostmanRuntime/7.29.2', 'Accept': '*/*', 'Accept-Encoding': 'gzip', 'Connection': 'keep-alive'} page = 1 data_nested = [] # 初始化空列表存储每个页面的DataFrame dfs = [] while data_nested is not None: data = [('vendor_id', xxxxx),('vendor_auth_code','4803155ec5f5a17d589b650cxxxxxxxxx'),('results_per_page',200),('page',page)] response = requests.post(url, headers=headers , data = data) data_nested = response.json()['response'] data_flattened = pd.json_normalize(data_nested) df = pd.DataFrame.from_dict(data_flattened) # 将当前页面的非空DataFrame加入列表 if not df.empty: dfs.append(df) if len(df.index)==0: break page += 1 # 合并所有DataFrame为一个完整的DataFrame final_df = pd.concat(dfs, ignore_index=True) # 查看合并结果 print(final_df)
关键修改说明
- 新增
dfs = []:专门用来存储每个分页返回的DataFrame,避免循环中每次生成的df被覆盖。 - 循环内添加
if not df.empty: dfs.append(df):确保只收集非空数据,避免空DataFrame影响合并结果。 - 循环结束后执行
pd.concat(dfs, ignore_index=True):将列表中的所有DataFrame纵向合并,ignore_index=True会重置合并后的索引,避免出现重复索引问题。
如果合并后需要进一步数据清洗(比如去重、处理缺失值),直接在final_df上操作即可。
内容的提问来源于stack exchange,提问作者Muyukani Kizito
相关产品推荐
相关产品推荐

