遍历DataFrame调用Pipedrive API POST数据的脚本中途挂起无报错如何处理?
原因分析
- 未设置请求超时:
requests库默认无超时逻辑,当网络波动、API服务端无返回时,请求会无限等待,这是该场景下最常见的诱因。 - 触发API速率限制:Pipedrive公开的接口规则对请求频率做了限制,基础版配额为每秒2次、每小时1000次请求,循环无间隔发送请求触发限流后,服务端可能会挂起连接不返回响应,不会直接抛出错误。
- 连接资源耗尽:每次请求独立新建TCP连接,短时间大量请求会占满本地连接池或触发服务端连接数限制,后续请求无法获取连接就会进入等待状态。
修复方案
1. 新增超时配置
给post方法添加timeout参数,超过指定时长无响应自动抛出异常,避免无限挂起:
# 单位为秒,可根据实际情况调整 post_response = requests.post(post_url, params=token, json=post_data, timeout=10)
2. 适配速率限制
添加请求间隔,触发限流时自动重试:
import time for i, row in df.iterrows(): post_data = {"name" : row['name']} post_url = "https://madeupcompany.pipedrive.com/api/v1/persons" post_response = requests.post(post_url, params=token, json=post_data, timeout=10) # 429是通用限流状态码 if post_response.status_code == 429: time.sleep(15) continue post_content = json.loads(post_response.content) print('Success: ', post_content['success']) # 请求间隔0.5秒,控制速率 time.sleep(0.5)
3. 复用HTTP连接
使用Session复用连接,减少连接开销,避免连接资源耗尽:
import requests # 全局初始化Session,循环内复用 session = requests.Session() token = {'api_token' : 'xxx'} for i, row in df.iterrows(): post_data = {"name" : row['name']} post_url = "https://madeupcompany.pipedrive.com/api/v1/persons" # 用session.post替代requests.post post_response = session.post(post_url, params=token, json=post_data, timeout=10) # 后续逻辑不变
4. 增加异常捕获
包裹请求、解析逻辑,遇到错误时打印信息继续执行,避免无提示挂起:
try: post_response = session.post(post_url, params=token, json=post_data, timeout=10) post_response.raise_for_status() post_content = json.loads(post_response.content) print('Success: ', post_content['success']) except Exception as e: print(f"第{i}行数据处理失败,错误:{str(e)}") continue
内容的提问来源于stack exchange,提问作者obi_wan_jabroni
相关产品推荐
相关产品推荐

