使用requests爬取Rarible个人页粉丝名称时重复获取前20条如何解决
问题解决说明
问题根源
你代码的核心错误是continuation参数的传值逻辑不符合接口要求:
- 该接口的
continuation是服务端返回的分页游标,用于标记下一次请求的起始位置,不能自行生成本地时间戳传入 - 你每次循环都传入当前时间戳,相当于每次都请求第一页数据,自然永远返回前20条结果
修正后的代码
import time import requests link = 'https://api-mainnet.rarible.com/marketplace/api/v4/followers' params = {'user': '0xe744d23107c9c98df5311ff8c1c8637ec3ecf9f3'} # 首次请求不需要带continuation参数 payload = {"size": 20} headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 6.1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/88.0.4324.150 Safari/537.36', 'origin': 'https://rarible.com', 'referer': 'https://rarible.com/' } with requests.Session() as s: s.headers.update(headers) while True: res = s.post(link, params=params, json=payload) # 先判断请求是否成功 if res.status_code != 200: print(f"请求出错,状态码:{res.status_code}") break follower_list = res.json() # 没有更多数据时终止循环 if not follower_list: print("所有粉丝已爬取完成") break for item in follower_list: print(item['owner'].get('name', '无名称')) # 从响应头中获取下一次请求需要的continuation值 next_continuation = res.headers.get('continuation') if not next_continuation: # 若未取到对应值可打印res.headers查看实际分页游标字段名 break payload['continuation'] = next_continuation time.sleep(2)
补充说明
- 代码增加了请求状态校验和数据为空的终止逻辑,避免死循环和异常报错
- 保留了2秒的请求间隔,避免请求频率过高触发站点反爬限制
内容的提问来源于stack exchange,提问作者SMTH
相关产品推荐
相关产品推荐

