You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

web.archive.org连接被拒问题及批量爬取优化咨询

Web Archive访问限制与批量爬取问题解答

问题概述

  • 疑问:web.archive.org是否存在访问限制?
  • 现象:Python设置1秒间隔爬取,不到一分钟出现ConnectionRefusedError等报错
  • 需求:了解批量爬取数千页的正确方法,以及当前操作的问题所在

报错详情

ConnectionRefusedError: [Errno 111] Connection refused

The above exception was the direct cause of the following exception:

NewConnectionError                        Traceback (most recent call last) NewConnectionError: <urllib3.connection.HTTPSConnection object at 0x7ae32e962fe0>: Failed to establish a new connection: [Errno 111] Connection refused

The above exception was the direct cause of the following exception:

MaxRetryError                             Traceback (most recent call last) MaxRetryError: HTTPSConnectionPool(host='web.archive.org', port=443): Max retries exceeded with url: /web/20230331062349/http://100vann.ru/ (Caused by NewConnectionError('<urllib3.connection.HTTPSConnection object at 0x7ae32e962fe0>: Failed to establish a new connection: [Errno 111] Connection refused'))

During handling of the above exception, another exception occurred:

ConnectionError                           Traceback (most recent call last) /usr/local/lib/python3.10/dist-packages/requests/adapters.py in send(self, request, stream, timeout, verify, cert, proxies)
    517                 raise SSLError(e, request=request)
    518 
--> 519             raise ConnectionError(e, request=request)
    520 
    521         except ClosedPoolError as e:

ConnectionError: HTTPSConnectionPool(host='web.archive.org', port=443): Max retries exceeded with url: /web/20230331062349/http://100vann.ru/ (Caused by NewConnectionError('<urllib3.connection.HTTPSConnection object at 0x7ae32e962fe0>: Failed to establish a new connection: [Errno 111] Connection refused'))

当前使用代码

headers = requests.utils.default_headers()
  time.sleep(1)
  headers.update(
    {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/108.0.0.0 Safari/537.36',
    }
  )
  webarchive = requests.get(url, headers=headers, timeout=15)
  webarchive.raise_for_status()

  if webarchive.status_code == 200:

问题分析与解决方案

1. Web Archive确实存在访问限制

Web Archive(Wayback Machine)有严格的爬虫限制,目的是保护服务器资源。即使设置1秒间隔,也可能因为以下原因被拦截:

  • 短时间内请求频率仍超出阈值:1秒间隔看似合理,但服务器会基于IP、请求特征等综合判断,批量请求时容易触发拦截
  • 缺少必要请求头:仅设置User-Agent不够,未模拟正常浏览器的完整请求头,容易被识别为爬虫
  • 未处理重试与异常:遇到临时错误直接终止,没有重试机制,导致连接失败累积

2. 当前代码的问题

  • time.sleep(1)位置错误:应该放在请求完成后,而非请求前,否则第一次请求无间隔就发送,容易触发限制
  • 请求头不完整:仅更新了User-Agent,缺少Accept、Accept-Language等关键头信息
  • 异常处理缺失:未捕获ConnectionError等异常,一旦报错就终止程序,无法继续爬取
  • 无重试机制:遇到临时连接失败时,没有重试逻辑,直接放弃请求

3. 批量爬取数千页的正确做法

  • 延长请求间隔:将间隔从1秒提升到3-5秒,甚至更长(根据实际情况调整),避免触发频率限制
  • 完善请求头:模拟完整的浏览器请求头,示例:
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/108.0.0.0 Safari/537.36',
        'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8',
        'Accept-Language': 'en-US,en;q=0.5',
        'Accept-Encoding': 'gzip, deflate, br',
        'Connection': 'keep-alive',
        'Upgrade-Insecure-Requests': '1'
    }
    
  • 添加重试机制:使用requests的会话和重试适配器处理临时错误,示例:
    from requests.adapters import HTTPAdapter
    from urllib3.util.retry import Retry
    
    session = requests.Session()
    retry = Retry(
        total=3,
        backoff_factor=1,
        status_forcelist=[429, 500, 502, 503, 504]
    )
    adapter = HTTPAdapter(max_retries=retry)
    session.mount('https://', adapter)
    session.mount('http://', adapter)
    
    # 使用session.get代替requests.get
    response = session.get(url, headers=headers, timeout=15)
    
  • 错误处理与日志记录:捕获异常并记录失败URL,后续再处理,示例:
    try:
        response = session.get(url, headers=headers, timeout=15)
        response.raise_for_status()
        if response.status_code == 200:
            # 处理页面内容
            pass
    except (ConnectionError, Timeout, requests.exceptions.HTTPError) as e:
        print(f"请求失败 {url}: {str(e)}")
        time.sleep(5)  # 延迟后可考虑重试
    
  • 分散请求目标:避免集中请求同一域名的归档页,分散请求范围,降低被识别为爬虫的概率
  • 使用官方API:Web Archive提供Wayback Machine API,通过API批量获取数据更合规,不易被拦截

内容的提问来源于stack exchange,提问作者IvanS

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 02:52:47