You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python分页调用API遇ValueError: Unterminated string问题求助

API分页数据获取报错与优化方案

问题背景

调用分页API(共200页,总计41500条数据),循环收集响应并写入JSON文件时,多次调用后出现ValueError: Unterminated string starting at:错误。

错误原因分析

这个错误本质是JSON解析失败,常见触发场景:

  • 网络波动导致API响应不完整、被截断
  • 请求过于频繁触发API限流,返回非JSON格式的限流提示
  • 手动拼接URL参数的方式错误,导致请求参数格式异常,API返回无效响应
  • 未正确判断分页终止条件,请求了不存在的页码,返回错误内容

解决与优化方案

1. 规范参数传递方式

不要手动拼接page参数到URL,将其加入params字典,由requests自动处理URL编码和拼接,避免格式错误:

params = {'pagesize': '200', 'versions': 'true', 'page': page}

2. 添加异常处理与重试机制

捕获网络异常、JSON解析错误,对失败请求进行重试,并记录错误页码,方便排查:

  • 使用try-except块包裹请求与解析逻辑
  • 对失败请求设置重试次数上限

3. 控制请求频率

添加短延迟,避免触发API限流:

import time
time.sleep(0.5)  # 每次请求后延迟0.5秒,可根据API限流规则调整

4. 正确判断分页终止条件

原代码new_results始终为True会无限循环,需根据API返回结果判断是否终止:

  • 若返回数据长度小于pagesize,说明已到最后一页
  • 若API返回total_pages字段,直接对比当前页码与总页数

5. 标准JSON文件写入

原代码逐行写入对象的方式生成的不是有效JSON数组,使用json.dump()直接写入整个列表:

json.dump(organisations, f, indent=2)

修改后的完整代码

import requests
import json
import time
from requests.exceptions import RequestException

url = "https://abcd../v1/org"
headers = {'Accept': 'application/json'}
auth = ('asdff', '364gbudgsh$')
pagesize = 200
params = {'pagesize': str(pagesize), 'versions': 'true'}

organisations = []
page = 1
max_retries = 3

while True:
    params['page'] = page
    success = False
    retry_count = 0
    
    while not success and retry_count < max_retries:
        try:
            response = requests.get(url, auth=auth, params=params, headers=headers)
            response.raise_for_status()  # 捕获HTTP状态码错误(如404、500)
            response_api = response.json()
            success = True
        except RequestException as e:
            retry_count += 1
            print(f"请求页码{page}失败,重试第{retry_count}次:{str(e)}")
            time.sleep(1)
        except json.JSONDecodeError as e:
            retry_count += 1
            print(f"页码{page}响应JSON解析失败,重试第{retry_count}次:{str(e)}")
            time.sleep(1)
    
    if not success:
        print(f"页码{page}多次重试失败,跳过该页")
        page += 1
        continue
    
    if not response_api:
        break
    
    organisations.extend(response_api)
    print(f"已获取页码{page},累计数据{len(organisations)}条")
    
    if len(response_api) < pagesize:
        break
    
    page += 1
    time.sleep(0.5)

# 写入标准JSON文件
with open('organisations.json', 'w', encoding='utf-8') as f:
    json.dump(organisations, f, indent=2, ensure_ascii=False)

print(f"数据写入完成,共{len(organisations)}条数据")

额外注意事项

  • 若API提供total或total_pages字段,建议直接用该字段判断循环终止,比判断返回数据长度更可靠
  • 确保auth参数的账号权限足够访问所有分页数据
  • 可将错误日志写入文件,方便后续排查问题

内容的提问来源于stack exchange,提问作者rshar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 12:22:43