web.archive.org连接被拒问题及批量爬取优化咨询
Web Archive访问限制与批量爬取问题解答
问题概述
- 疑问:web.archive.org是否存在访问限制?
- 现象:Python设置1秒间隔爬取,不到一分钟出现
ConnectionRefusedError等报错 - 需求:了解批量爬取数千页的正确方法,以及当前操作的问题所在
报错详情
ConnectionRefusedError: [Errno 111] Connection refused The above exception was the direct cause of the following exception: NewConnectionError Traceback (most recent call last) NewConnectionError: <urllib3.connection.HTTPSConnection object at 0x7ae32e962fe0>: Failed to establish a new connection: [Errno 111] Connection refused The above exception was the direct cause of the following exception: MaxRetryError Traceback (most recent call last) MaxRetryError: HTTPSConnectionPool(host='web.archive.org', port=443): Max retries exceeded with url: /web/20230331062349/http://100vann.ru/ (Caused by NewConnectionError('<urllib3.connection.HTTPSConnection object at 0x7ae32e962fe0>: Failed to establish a new connection: [Errno 111] Connection refused')) During handling of the above exception, another exception occurred: ConnectionError Traceback (most recent call last) /usr/local/lib/python3.10/dist-packages/requests/adapters.py in send(self, request, stream, timeout, verify, cert, proxies) 517 raise SSLError(e, request=request) 518 --> 519 raise ConnectionError(e, request=request) 520 521 except ClosedPoolError as e: ConnectionError: HTTPSConnectionPool(host='web.archive.org', port=443): Max retries exceeded with url: /web/20230331062349/http://100vann.ru/ (Caused by NewConnectionError('<urllib3.connection.HTTPSConnection object at 0x7ae32e962fe0>: Failed to establish a new connection: [Errno 111] Connection refused'))
当前使用代码
headers = requests.utils.default_headers() time.sleep(1) headers.update( { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/108.0.0.0 Safari/537.36', } ) webarchive = requests.get(url, headers=headers, timeout=15) webarchive.raise_for_status() if webarchive.status_code == 200:
问题分析与解决方案
1. Web Archive确实存在访问限制
Web Archive(Wayback Machine)有严格的爬虫限制,目的是保护服务器资源。即使设置1秒间隔,也可能因为以下原因被拦截:
- 短时间内请求频率仍超出阈值:1秒间隔看似合理,但服务器会基于IP、请求特征等综合判断,批量请求时容易触发拦截
- 缺少必要请求头:仅设置User-Agent不够,未模拟正常浏览器的完整请求头,容易被识别为爬虫
- 未处理重试与异常:遇到临时错误直接终止,没有重试机制,导致连接失败累积
2. 当前代码的问题
time.sleep(1)位置错误:应该放在请求完成后,而非请求前,否则第一次请求无间隔就发送,容易触发限制- 请求头不完整:仅更新了User-Agent,缺少Accept、Accept-Language等关键头信息
- 异常处理缺失:未捕获
ConnectionError等异常,一旦报错就终止程序,无法继续爬取 - 无重试机制:遇到临时连接失败时,没有重试逻辑,直接放弃请求
3. 批量爬取数千页的正确做法
- 延长请求间隔:将间隔从1秒提升到3-5秒,甚至更长(根据实际情况调整),避免触发频率限制
- 完善请求头:模拟完整的浏览器请求头,示例:
headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/108.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Accept-Encoding': 'gzip, deflate, br', 'Connection': 'keep-alive', 'Upgrade-Insecure-Requests': '1' } - 添加重试机制:使用
requests的会话和重试适配器处理临时错误,示例:from requests.adapters import HTTPAdapter from urllib3.util.retry import Retry session = requests.Session() retry = Retry( total=3, backoff_factor=1, status_forcelist=[429, 500, 502, 503, 504] ) adapter = HTTPAdapter(max_retries=retry) session.mount('https://', adapter) session.mount('http://', adapter) # 使用session.get代替requests.get response = session.get(url, headers=headers, timeout=15) - 错误处理与日志记录:捕获异常并记录失败URL,后续再处理,示例:
try: response = session.get(url, headers=headers, timeout=15) response.raise_for_status() if response.status_code == 200: # 处理页面内容 pass except (ConnectionError, Timeout, requests.exceptions.HTTPError) as e: print(f"请求失败 {url}: {str(e)}") time.sleep(5) # 延迟后可考虑重试 - 分散请求目标:避免集中请求同一域名的归档页,分散请求范围,降低被识别为爬虫的概率
- 使用官方API:Web Archive提供Wayback Machine API,通过API批量获取数据更合规,不易被拦截
内容的提问来源于stack exchange,提问作者IvanS
相关产品推荐
相关产品推荐

