You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium+Edge WebDriver获取Web of Science全部结果的href值

问题原因

Web of Science的搜索结果采用懒加载机制——只有当条目滚动到浏览器可视区域内时,才会生成对应的DOM元素。你直接调用find_elements时,页面只渲染了前4条可见结果,所以只能抓到这4条的链接。

解决方法

需要通过滚动页面触发所有结果的渲染,再提取链接。以下是修改后的代码:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

driver = webdriver.Edge()
driver.get("https://www.webofscience.com/wos/woscc/summary/6224cee5-30d2-4e62-8729-212ea8f51522-c6c4cc63/times-cited-descending/1")

# 等待初始结果加载完成
wait = WebDriverWait(driver, 20)
wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, 'a[data-ta="summary-record-title-link"]')))

# 循环滚动页面,加载所有50条结果
previous_count = 0
target_count = 50
while True:
    # 滚动到页面底部
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    # 等待新元素加载
    time.sleep(2)
    # 获取当前已加载的元素数量
    current_elements = driver.find_elements(By.CSS_SELECTOR, 'a[data-ta="summary-record-title-link"]')
    current_count = len(current_elements)
    # 当元素数量不再增加或达到目标数量时停止
    if current_count == previous_count or current_count >= target_count:
        break
    previous_count = current_count

# 提取所有href值
for element in current_elements:
    href = element.get_attribute('href')
    if href:  # 过滤空值
        print(href)

driver.quit()
关键说明
  • 改用By.CSS_SELECTOR定位元素,比By.TAG_NAME更精准
  • 用循环滚动+等待的方式触发懒加载,确保所有50条结果都被渲染
  • 添加了结果数量判断,避免无限滚动
  • 增加空值过滤,防止输出无效链接

内容的提问来源于stack exchange,提问作者Pin 8

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 21:04:56