You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多页API仅读取最后一页,如何合并所有页content到DataFrame?

解决多页API数据合并到DataFrame的问题

核心思路

不要每次处理单页数据就直接覆盖DataFrame,而是先把所有页的content数据收集到一个全局列表里,最后一次性转换成DataFrame。

具体实现步骤

1. 初始化空列表存储所有数据

import http.client
import json
import pandas as pd

all_content = []

2. 循环请求所有页码的API

根据API的分页规则(比如你提到的number参数),逐页请求并收集数据:

# 示例:假设要请求第1到第5页
for page_num in range(1, 6):
    conn = http.client.HTTPSConnection("api.example.com")
    headers = {'Content-Type': 'application/json'}
    # 替换成实际的API路径和分页参数,这里用number作为页码参数
    conn.request("GET", f"/your-api-path?number={page_num}", headers=headers)
    
    res = conn.getresponse()
    data = res.read()
    response_json = json.loads(data.decode("utf-8"))
    
    # 将当前页的content数据追加到全局列表
    all_content.extend(response_json['content'])
    
    conn.close()

注意:如果API返回的content是单个字典而非列表,把extend改成append即可。

3. 批量转换为DataFrame

等所有页的数据都收集完成后,统一用pd.json_normalize转换:

final_df = pd.json_normalize(all_content)
print(final_df)

进阶处理:自动识别总页数

如果不知道总页数,可以先请求第一页,从返回结果中提取总页数再循环:

# 先请求第一页获取分页信息
conn = http.client.HTTPSConnection("api.example.com")
conn.request("GET", "/your-api-path?number=1", headers=headers)
res = conn.getresponse()
first_page_data = json.loads(res.read().decode("utf-8"))
total_pages = first_page_data['total_pages']  # 替换成API实际返回的总页数字段名
all_content.extend(first_page_data['content'])
conn.close()

# 循环请求剩余页码
for page_num in range(2, total_pages + 1):
    # 重复上述请求逻辑,将content追加到all_content
    conn = http.client.HTTPSConnection("api.example.com")
    conn.request("GET", f"/your-api-path?number={page_num}", headers=headers)
    res = conn.getresponse()
    page_data = json.loads(res.read().decode("utf-8"))
    all_content.extend(page_data['content'])
    conn.close()

# 生成最终DataFrame
final_df = pd.json_normalize(all_content)

内容的提问来源于stack exchange,提问作者Ron Kieftenbeld

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 04:24:32