如何遵循Spotipy的sleep_for_retry()函数避免请求阻塞?
如何合规使用Spotipy批量获取艺人数据,避免触发API速率限制并解决sleep_for_retry阻塞问题
核心问题分析
你遇到的程序阻塞,本质是触发了Spotify API的速率限制——服务器返回了Retry-After响应头,urllib3的默认重试机制会严格按照这个头指定的时间休眠,导致程序看起来“卡住”。你手动添加的固定time.sleep(15)无法从根本上解决问题:一是单请求循环的总请求量太大,二是固定休眠时间没有匹配API的实际限制规则。
解决方案
1. 优先用批量接口砍请求量
Spotipy提供了artists()批量接口,一次最多可传入50个艺人ID。48000个艺人的请求量能从48000次直接降到960次,这是规避速率限制最有效的手段。
示例代码:
import spotipy import time from spotipy.oauth2 import SpotifyClientCredentials # 初始化客户端 client_credentials_manager = SpotifyClientCredentials(client_id='你的客户端ID', client_secret='你的客户端密钥') sp = spotipy.Spotify(client_credentials_manager=client_credentials_manager) uniqueIds = [你的艺人ID列表] popularIds2 = [] # 按50个ID为一批次拆分处理 for i in range(0, len(uniqueIds), 50): batch_ids = uniqueIds[i:i+50] # 批量拉取艺人数据 batch_results = sp.artists(batch_ids) # 筛选流行度>23的艺人ID for artist in batch_results['artists']: if artist['popularity'] > 23: popularIds2.append(artist['id']) # 动态根据速率限制头调整休眠 headers = sp._session.headers remaining = int(headers.get('X-RateLimit-Remaining', 150)) if remaining < 10: reset_time = int(headers.get('X-RateLimit-Reset', 0)) sleep_time = max(reset_time - int(time.time()), 5) time.sleep(sleep_time)
2. 动态响应速率限制头
Spotify API的响应头会返回3个关键限制信息,你可以根据这些信息动态调整休眠,而不是用固定值:
X-RateLimit-Limit:周期内允许的最大请求数X-RateLimit-Remaining:当前剩余可用请求数X-RateLimit-Reset:限制重置的Unix时间戳
示例检查逻辑:
def check_and_wait(sp): # 用任意请求的响应头获取限制信息 test_response = sp._internal_call('GET', 'https://api.spotify.com/v1/artists/3TVXtAsR1Inumwj472S9r4') headers = test_response.headers remaining = int(headers['X-RateLimit-Remaining']) reset_time = int(headers['X-RateLimit-Reset']) # 剩余请求数不足时,等待到限制重置 if remaining < 20: wait_time = reset_time - int(time.time()) if wait_time > 0: time.sleep(wait_time + 1) # 多等1秒确保重置完成
3. 自定义重试策略,避免无限阻塞
默认urllib3会无限重试并严格遵循Retry-After休眠,你可以自定义策略限制重试次数和等待时间:
from urllib3.util.retry import Retry from requests.adapters import HTTPAdapter # 配置重试策略:最多重试3次,不自动遵循Retry-After retry_strategy = Retry( total=3, backoff_factor=1, status_forcelist=[429, 500, 502, 503, 504], respect_retry_after_header=False ) adapter = HTTPAdapter(max_retries=retry_strategy) # 给Spotipy的会话挂载自定义适配器 sp._session.mount('https://', adapter) sp._session.mount('http://', adapter)
4. 手动捕获429错误处理
如果还是触发了速率限制,直接捕获429状态码,读取Retry-After头手动休眠:
import requests try: batch_results = sp.artists(batch_ids) except requests.exceptions.HTTPError as e: if e.response.status_code == 429: retry_after = int(e.response.headers.get('Retry-After', 60)) print(f"触发速率限制,等待{retry_after}秒") time.sleep(retry_after + 2) # 重新发起请求 batch_results = sp.artists(batch_ids) else: raise
总结
最核心的优化是改用批量接口,直接将请求量削减到原有的1/100,从根源上降低触发限制的概率。配合动态监听速率限制头和自定义重试策略,就能彻底解决程序卡在sleep_for_retry()的问题。
内容的提问来源于stack exchange,提问作者boogers
相关产品推荐
相关产品推荐

