You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为Selenium爬虫集成分页功能?求替代方案建议

一、Selenium分页功能集成实现

针对clinicaltrials.gov的分页抓取,核心逻辑是循环抓取当前页链接 → 检查下一页可点击性 → 点击下一页重复操作,以下是具体实现步骤和代码示例:

核心逻辑说明

  1. 页面加载完成后,定位并提取当前页所有试验链接(通常为href="/ct2/show/NCTxxxxxx"格式的元素)
  2. 判断底部「Next」按钮是否处于可点击状态(未被禁用)
  3. 若可点击则跳转下一页重复抓取;若不可点击则终止循环

代码示例

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import NoSuchElementException, ElementNotInteractableException
import time
import random

driver = webdriver.Chrome()
driver.get("https://clinicaltrials.gov/search?term=your-search-keyword")  # 替换为你的搜索URL

all_links = []

while True:
    # 等待当前页链接加载完成
    WebDriverWait(driver, 10).until(
        EC.presence_of_all_elements_located((By.CSS_SELECTOR, "a[href^='/ct2/show/NCT']"))
    )
    
    # 提取当前页所有链接
    current_page_links = [elem.get_attribute("href") for elem in driver.find_elements(By.CSS_SELECTOR, "a[href^='/ct2/show/NCT']")]
    all_links.extend(current_page_links)
    print(f"累计抓取 {len(all_links)} 条链接")
    
    try:
        # 定位下一页按钮
        next_btn = driver.find_element(By.CSS_SELECTOR, "button.pagination-next")
        if "disabled" in next_btn.get_attribute("class"):
            break  # 下一页按钮禁用,终止循环
        
        next_btn.click()
        # 等待页面跳转完成(通过元素失效判断)
        WebDriverWait(driver, 10).until(
            EC.staleness_of(driver.find_element(By.CSS_SELECTOR, "a[href^='/ct2/show/NCT']"))
        )
        # 随机延迟,规避反爬
        time.sleep(random.uniform(1, 3))
    except (NoSuchElementException, ElementNotInteractableException):
        break  # 找不到下一页按钮或无法点击,终止循环

driver.quit()
# 保存结果到文件
with open("clinical_trials_links.txt", "w") as f:
    f.write("\n".join(all_links))

注意事项

  • 超万条数据建议每抓取N页就保存一次结果,避免程序崩溃丢失数据
  • 可根据网站反爬强度调整延迟时间或添加代理IP

二、非Selenium替代方案

1. 官方API调用(最推荐)

clinicaltrials.gov提供结构化数据API,无需解析HTML,直接获取标准化数据,支持分页参数,稳定性最高:

核心参数

  • limit:单次返回结果数量(最大支持1000)
  • offset:偏移量,用于控制分页位置
  • term:搜索关键词(与网页搜索逻辑一致)

代码示例

import requests

base_url = "https://clinicaltrials.gov/api/v2/studies"
params = {
    "term": "your-search-keyword",  # 替换为你的搜索词
    "limit": 100,  # 单次取100条,可按需调整
    "offset": 0
}

all_links = []

while True:
    response = requests.get(base_url, params=params)
    data = response.json()
    
    if not data.get("studies"):
        break  # 无更多结果,终止循环
    
    # 提取每条试验的详情页链接
    for study in data["studies"]:
        nct_id = study["protocolSection"]["identificationModule"]["nctId"]
        all_links.append(f"https://clinicaltrials.gov/ct2/show/{nct_id}")
    
    params["offset"] += params["limit"]
    print(f"累计抓取 {len(all_links)} 条链接")

# 保存结果
with open("clinical_trials_api_links.txt", "w") as f:
    f.write("\n".join(all_links))

2. Requests + BeautifulSoup(轻量爬取)

直接请求网页HTML并解析,适合中规模数据抓取,资源消耗远低于Selenium:

代码示例

import requests
from bs4 import BeautifulSoup
import time
import random

base_url = "https://clinicaltrials.gov/search?term=your-search-keyword&page="
page_num = 1
all_links = []
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

while True:
    url = f"{base_url}{page_num}"
    response = requests.get(url, headers=headers)
    soup = BeautifulSoup(response.text, "html.parser")
    
    # 提取当前页链接
    links = soup.select("a[href^='/ct2/show/NCT']")
    if not links:
        break
    
    current_links = [f"https://clinicaltrials.gov{link['href']}" for link in links]
    all_links.extend(current_links)
    print(f"累计抓取 {len(all_links)} 条链接")
    
    # 检查是否有下一页
    next_page_btn = soup.select_one("button.pagination-next:not([disabled])")
    if not next_page_btn:
        break
    
    page_num += 1
    time.sleep(random.uniform(2, 4))  # 添加延迟,规避反爬

# 保存结果
with open("clinical_trials_bs_links.txt", "w") as f:
    f.write("\n".join(all_links))

3. Scrapy框架(大规模分布式爬取)

适合超大规模数据抓取,自带并发控制、反爬中间件、分页处理机制,效率最高:

爬虫示例代码

import scrapy

class ClinicalTrialsSpider(scrapy.Spider):
    name = "clinicaltrials"
    start_urls = ["https://clinicaltrials.gov/search?term=your-search-keyword"]

    def parse(self, response):
        # 提取当前页所有链接
        links = response.css("a[href^='/ct2/show/NCT']::attr(href)").getall()
        for link in links:
            yield {"url": f"https://clinicaltrials.gov{link}"}
        
        # 定位下一页按钮并构造跳转URL
        next_btn = response.css("button.pagination-next:not([disabled])")
        if next_btn:
            next_page_path = next_btn.attrib.get("onclick").split("'")[1]
            next_page_url = response.urljoin(next_page_path)
            yield scrapy.Request(next_page_url, callback=self.parse)

内容的提问来源于stack exchange,提问作者Kaan Turgay

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 05:43:01