You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium WebDriverWait爬取ScienceDirect期刊链接及文本遇异常求助

问题排查与修复方案

一、总页数计算不稳定的修复

原代码中总页数计算依赖手动截取字符串,且未确保文本完全加载,导致结果不稳定。修复思路:

  • 用text_to_be_present_in_element显式等待分页文本加载完成(确保包含"of"关键字)
  • 用正则表达式提取总页数,避免手动截取的脆弱性

二、其他代码问题修复

  1. 移除重复的元素等待逻辑,复用已获取的分页元素
  2. 调整期刊链接的等待条件为visibility_of_all_elements_located,确保元素可见后再提取链接
  3. 修正Aims & Scope的定位XPATH,原定位不准确,改为定位展开后的内容容器
  4. 完全移除time.sleep,全部使用显式等待控制页面交互节奏

修改后的完整代码

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from webdriver_manager.chrome import ChromeDriverManager
import re

options = Options()
options.add_argument("start-maximized")
driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options)
wait = WebDriverWait(driver, 20)

# 初始页面加载
url = "https://www.sciencedirect.com/browse/journals-and-books?accessType=openAccess&accessType=containsOpenAccess"
driver.get(url)

# 等待分页文本加载完成并提取总页数
page_desc_locator = (By.XPATH, "//span[@class='pagination-pages-label u-margin-s-left-from-sm u-margin-s-right-from-sm']")
wait.until(EC.text_to_be_present_in_element(page_desc_locator, "of"))
page_description = driver.find_element(*page_desc_locator)
# 用正则提取"of"后面的数字
match = re.search(r'of (\d+)', page_description.text)
pages = int(match.group(1)) if match else 0

all_journal_links = []

# 遍历所有分页获取期刊链接
for page_num in range(1, pages + 1):
    current_page_url = f"https://www.sciencedirect.com/browse/journals-and-books?page={page_num}&accessType=containsOpenAccess&accessType=openAccess"
    driver.get(current_page_url)
    # 等待所有期刊标题链接可见
    journal_links = wait.until(EC.visibility_of_all_elements_located((By.XPATH, "//a[@class='anchor js-publication-title anchor-default']")))
    for link in journal_links:
        href = link.get_attribute('href')
        if href:
            all_journal_links.append(href)

# 遍历期刊链接提取Aims & Scope
for link in all_journal_links:
    driver.get(link)
    # 等待"View Full Aims & Scope"按钮可点击并点击
    view_button = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[contains(text(), 'View Full Aims & Scope')]")))
    view_button.click()
    # 等待Aims & Scope内容加载并提取
    aims_scope_content = wait.until(EC.visibility_of_element_located((By.XPATH, "//div[contains(@class, 'aims-and-scope')]/div[@class='content']")))
    print("Aims & Scope:", aims_scope_content.text)

driver.quit()

关键改进说明

  • 总页数提取:用正则匹配of (\d+)直接获取数字,避免手动截取字符串的错误风险,同时等待文本包含"of"确保内容加载完成
  • 元素等待优化:将期刊链接的等待条件改为visibility_of_all_elements_located,确保元素可见后再提取链接,避免获取到未渲染完成的元素
  • 按钮定位优化:用contains(text(), 'View Full Aims & Scope')定位按钮,比仅靠class更稳定,避免class变更导致定位失败
  • 内容提取优化:直接定位Aims & Scope的内容容器,确保提取到完整的目标文本

内容的提问来源于stack exchange,提问作者trgjk yfojn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 16:45:31