You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python抓取WayBack Machine超时求助:获取NIH研究名单失败

解决WayBack Machine请求超时问题

问题背景

需要从WayBack Machine归档页面抓取2011年NIH所有研究小组名单:

  1. 访问指定归档链接获取所有研究小组的详情链接
  2. 进入每个详情页提取「View Roster」链接并抓取名单
    目前已成功获取小组链接,但请求详情页时触发超时错误,导致代码崩溃。

错误详情

HTTPSConnectionPool(host='web.archive.org', port=443): Max retries exceeded with url: /web/20111027104153/http://internet.csr.nih.gov/Roster_proto1/sectionI_list_detail.asp?NEWSRG=ACE&SRG=ACE&SRGDISPLAY=ACE (Caused by NewConnectionError('<urllib3.connection.HTTPSConnection object at 0x7fcd72ce0c40>: Failed to establish a new connection: [Errno 60] Operation timed out'))

现有实现代码

获取研究小组链接

## GLOBAL SET-UP
options = webdriver.ChromeOptions()
options.add_argument("headless")
options.add_experimental_option('excludeSwitches', ['enable-logging'])
driver = webdriver.Chrome(executable_path='PATH',options=options)

STARTING_URL = "https://web.archive.org/web/20111027104153/http://public.csr.nih.gov/StudySections/Standing/Pages/default.aspx"
headers={'User-Agent': 'Safari'}
starting_page=requests.get(STARTING_URL, headers=headers)
starting_soup = BeautifulSoup(starting_page.content, 'html.parser')

table = starting_soup.find('table', summary="This table contains information on CSR Meeting Roster")

overall_dict = {}
overall_dict_n_n = {}

counter = 1
counter_n_n=1

#####################################    
#STEP 1: Getting study sections link#
#####################################
for i, row in enumerate(table.find_all('tr')):
    if i == 0:
        header = [el.text.strip() for el in row.find_all('th')]
    else:
        href = row.find("a").get("href") # get the hyperlink for the person
        name = [el.text.strip() for el in row.find_all('td')] # get the name for the person
        overall_dict[counter] = [name, href]
        
        counter += 1 

尝试获取「View Roster」链接(触发超时的代码)

#####################################
#STEP 2: Getting study sections link#
#####################################
for identifier, info_list in overall_dict.items():
    print(info_list[1])

    individuals_page=requests.get(info_list[1],headers=headers)
    starting_soup = BeautifulSoup(individuals_page.content, 'lxml')
    
    column = starting_soup.find("table", {"id": "Table1"})
    people_in_column=column.find_all("tr")[1]

    href_n_n = people_in_column.find_all("td")[3].find("a").get("href") # get the hyperlink for the person
    name_n_n = people_in_column.find_all("td")[0].find("font").contents[0] 
                 
             
    overall_dict_n_n[counter_n_n] = [name_n_n,href_n_n,info_list[0]]
    counter_n_n += 1 # increment the counter for the next person

解决方法

1. 添加超时参数与重试机制

WayBack Machine对高频请求有限制,添加超时时间+自动重试可解决临时连接问题。使用requests.adapters.HTTPAdapter配置重试策略:

import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

# 配置重试策略
session = requests.Session()
retry = Retry(
    total=3,  # 总重试次数
    backoff_factor=1,  # 重试间隔(1s, 2s, 4s...)
    status_forcelist=[429, 500, 502, 503, 504]  # 需要重试的状态码
)
adapter = HTTPAdapter(max_retries=retry)
session.mount('https://', adapter)
session.mount('http://', adapter)

# 修改STEP 2的请求代码
for identifier, info_list in overall_dict.items():
    print(info_list[1])
    try:
        # 添加超时参数(10s)
        individuals_page = session.get(info_list[1], headers=headers, timeout=10)
        individuals_page.raise_for_status()  # 抛出HTTP错误
        starting_soup = BeautifulSoup(individuals_page.content, 'lxml')
        
        column = starting_soup.find("table", {"id": "Table1"})
        if not column:
            print(f"未找到Table1,跳过: {info_list[1]}")
            continue
        people_in_column = column.find_all("tr")[1]
        
        href_n_n = people_in_column.find_all("td")[3].find("a").get("href")
        name_n_n = people_in_column.find_all("td")[0].find("font").contents[0] 
        
        overall_dict_n_n[counter_n_n] = [name_n_n, href_n_n, info_list[0]]
        counter_n_n += 1
    except requests.exceptions.RequestException as e:
        print(f"请求失败: {info_list[1]}, 错误: {str(e)}")
        continue

2. 降低请求频率

在每次请求后添加延迟,避免触发WayBack的反爬限制:

import time

# 在STEP 2的循环中添加
for identifier, info_list in overall_dict.items():
    print(info_list[1])
    try:
        individuals_page = session.get(info_list[1], headers=headers, timeout=10)
        # ... 处理逻辑 ...
        time.sleep(2)  # 每次请求后暂停2秒
    except requests.exceptions.RequestException as e:
        # ... 异常处理 ...
        time.sleep(5)  # 请求失败后暂停更久再重试

3. 完善User-Agent头

使用更真实的浏览器UA字符串,避免被识别为爬虫:

headers = {
    'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/16.5 Safari/605.1.15'
}

4. 复用已有的Selenium Driver

既然已经初始化了ChromeDriver,直接用它访问详情页,避免requests的连接问题:

# 修改STEP 2代码
for identifier, info_list in overall_dict.items():
    print(info_list[1])
    try:
        driver.get(info_list[1])
        # 等待页面加载(可选,根据需要添加显式等待)
        time.sleep(1)
        starting_soup = BeautifulSoup(driver.page_source, 'lxml')
        
        column = starting_soup.find("table", {"id": "Table1"})
        if not column:
            print(f"未找到Table1,跳过: {info_list[1]}")
            continue
        people_in_column = column.find_all("tr")[1]
        
        href_n_n = people_in_column.find_all("td")[3].find("a").get("href")
        name_n_n = people_in_column.find_all("td")[0].find("font").contents[0] 
        
        overall_dict_n_n[counter_n_n] = [name_n_n, href_n_n, info_list[0]]
        counter_n_n += 1
        time.sleep(2)
    except Exception as e:
        print(f"处理失败: {info_list[1]}, 错误: {str(e)}")
        continue

内容的提问来源于stack exchange,提问作者Clara HL

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 12:51:56