Python抓取WayBack Machine超时求助:获取NIH研究名单失败
解决WayBack Machine请求超时问题
问题背景
需要从WayBack Machine归档页面抓取2011年NIH所有研究小组名单:
- 访问指定归档链接获取所有研究小组的详情链接
- 进入每个详情页提取「View Roster」链接并抓取名单
目前已成功获取小组链接,但请求详情页时触发超时错误,导致代码崩溃。
错误详情
HTTPSConnectionPool(host='web.archive.org', port=443): Max retries exceeded with url: /web/20111027104153/http://internet.csr.nih.gov/Roster_proto1/sectionI_list_detail.asp?NEWSRG=ACE&SRG=ACE&SRGDISPLAY=ACE (Caused by NewConnectionError('<urllib3.connection.HTTPSConnection object at 0x7fcd72ce0c40>: Failed to establish a new connection: [Errno 60] Operation timed out'))
现有实现代码
获取研究小组链接
## GLOBAL SET-UP options = webdriver.ChromeOptions() options.add_argument("headless") options.add_experimental_option('excludeSwitches', ['enable-logging']) driver = webdriver.Chrome(executable_path='PATH',options=options) STARTING_URL = "https://web.archive.org/web/20111027104153/http://public.csr.nih.gov/StudySections/Standing/Pages/default.aspx" headers={'User-Agent': 'Safari'} starting_page=requests.get(STARTING_URL, headers=headers) starting_soup = BeautifulSoup(starting_page.content, 'html.parser') table = starting_soup.find('table', summary="This table contains information on CSR Meeting Roster") overall_dict = {} overall_dict_n_n = {} counter = 1 counter_n_n=1 ##################################### #STEP 1: Getting study sections link# ##################################### for i, row in enumerate(table.find_all('tr')): if i == 0: header = [el.text.strip() for el in row.find_all('th')] else: href = row.find("a").get("href") # get the hyperlink for the person name = [el.text.strip() for el in row.find_all('td')] # get the name for the person overall_dict[counter] = [name, href] counter += 1
尝试获取「View Roster」链接(触发超时的代码)
##################################### #STEP 2: Getting study sections link# ##################################### for identifier, info_list in overall_dict.items(): print(info_list[1]) individuals_page=requests.get(info_list[1],headers=headers) starting_soup = BeautifulSoup(individuals_page.content, 'lxml') column = starting_soup.find("table", {"id": "Table1"}) people_in_column=column.find_all("tr")[1] href_n_n = people_in_column.find_all("td")[3].find("a").get("href") # get the hyperlink for the person name_n_n = people_in_column.find_all("td")[0].find("font").contents[0] overall_dict_n_n[counter_n_n] = [name_n_n,href_n_n,info_list[0]] counter_n_n += 1 # increment the counter for the next person
解决方法
1. 添加超时参数与重试机制
WayBack Machine对高频请求有限制,添加超时时间+自动重试可解决临时连接问题。使用requests.adapters.HTTPAdapter配置重试策略:
import requests from requests.adapters import HTTPAdapter from urllib3.util.retry import Retry # 配置重试策略 session = requests.Session() retry = Retry( total=3, # 总重试次数 backoff_factor=1, # 重试间隔(1s, 2s, 4s...) status_forcelist=[429, 500, 502, 503, 504] # 需要重试的状态码 ) adapter = HTTPAdapter(max_retries=retry) session.mount('https://', adapter) session.mount('http://', adapter) # 修改STEP 2的请求代码 for identifier, info_list in overall_dict.items(): print(info_list[1]) try: # 添加超时参数(10s) individuals_page = session.get(info_list[1], headers=headers, timeout=10) individuals_page.raise_for_status() # 抛出HTTP错误 starting_soup = BeautifulSoup(individuals_page.content, 'lxml') column = starting_soup.find("table", {"id": "Table1"}) if not column: print(f"未找到Table1,跳过: {info_list[1]}") continue people_in_column = column.find_all("tr")[1] href_n_n = people_in_column.find_all("td")[3].find("a").get("href") name_n_n = people_in_column.find_all("td")[0].find("font").contents[0] overall_dict_n_n[counter_n_n] = [name_n_n, href_n_n, info_list[0]] counter_n_n += 1 except requests.exceptions.RequestException as e: print(f"请求失败: {info_list[1]}, 错误: {str(e)}") continue
2. 降低请求频率
在每次请求后添加延迟,避免触发WayBack的反爬限制:
import time # 在STEP 2的循环中添加 for identifier, info_list in overall_dict.items(): print(info_list[1]) try: individuals_page = session.get(info_list[1], headers=headers, timeout=10) # ... 处理逻辑 ... time.sleep(2) # 每次请求后暂停2秒 except requests.exceptions.RequestException as e: # ... 异常处理 ... time.sleep(5) # 请求失败后暂停更久再重试
3. 完善User-Agent头
使用更真实的浏览器UA字符串,避免被识别为爬虫:
headers = { 'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/16.5 Safari/605.1.15' }
4. 复用已有的Selenium Driver
既然已经初始化了ChromeDriver,直接用它访问详情页,避免requests的连接问题:
# 修改STEP 2代码 for identifier, info_list in overall_dict.items(): print(info_list[1]) try: driver.get(info_list[1]) # 等待页面加载(可选,根据需要添加显式等待) time.sleep(1) starting_soup = BeautifulSoup(driver.page_source, 'lxml') column = starting_soup.find("table", {"id": "Table1"}) if not column: print(f"未找到Table1,跳过: {info_list[1]}") continue people_in_column = column.find_all("tr")[1] href_n_n = people_in_column.find_all("td")[3].find("a").get("href") name_n_n = people_in_column.find_all("td")[0].find("font").contents[0] overall_dict_n_n[counter_n_n] = [name_n_n, href_n_n, info_list[0]] counter_n_n += 1 time.sleep(2) except Exception as e: print(f"处理失败: {info_list[1]}, 错误: {str(e)}") continue
内容的提问来源于stack exchange,提问作者Clara HL
相关产品推荐
相关产品推荐

