Python循环中遇HTTP错误后如何重试访问网页?
问题描述
- 手动访问网页时会出现522、525、504这类HTTP错误
- 运行以下Python代码遍历字典获取subreddit 2022年帖子时,偶尔会因
HTTP Error 525或其他同类错误终止循环:
for subredd, url in dict_last_subreddit_posts.items(): print(subredd) page = urllib.request.urlopen(url).read() dict_last_posts[subredd] = page
- 需求:读取URL时检测这类错误,自动重试直到请求成功,再处理下一个subreddit
解决方案
通过捕获HTTP错误+循环重试的方式实现,同时添加延迟避免频繁请求被限制,示例代码如下:
import urllib.request from urllib.error import HTTPError import time dict_last_posts = {} # 可根据需求调整重试次数和间隔时间 max_retries = 5 retry_delay = 3 # 单位:秒 for subredd, url in dict_last_subreddit_posts.items(): print(subredd) success = False retries = 0 while not success and retries < max_retries: try: page = urllib.request.urlopen(url).read() dict_last_posts[subredd] = page success = True except HTTPError as e: error_code = e.getcode() # 针对522、525、504这类网关类错误进行重试 if error_code in (522, 525, 504): retries += 1 print(f"遇到HTTP错误 {error_code},第 {retries} 次重试...") time.sleep(retry_delay) else: # 其他非目标错误直接跳过或记录 print(f"遇到非重试类HTTP错误 {error_code},跳过该subreddit") break except Exception as e: # 处理网络连接等其他未知异常 retries += 1 print(f"遇到未知错误:{str(e)},第 {retries} 次重试...") time.sleep(retry_delay) if not success: print(f"重试 {max_retries} 次后仍失败,跳过 {subredd}")
关键说明
- 用
urllib.error.HTTPError捕获HTTP状态码错误,精准判断是否属于需要重试的类型 - 设置
max_retries限制最大重试次数,避免无限循环 - 添加
time.sleep(retry_delay),每次重试前等待一段时间,减轻服务器压力,降低被拦截概率 - 对404、403等非目标错误直接跳过,避免无效重试
内容的提问来源于stack exchange,提问作者MichelleM
相关产品推荐
相关产品推荐

