Selenium WebDriverWait爬取ScienceDirect期刊链接及文本遇异常求助
问题排查与修复方案
一、总页数计算不稳定的修复
原代码中总页数计算依赖手动截取字符串,且未确保文本完全加载,导致结果不稳定。修复思路:
- 用
text_to_be_present_in_element显式等待分页文本加载完成(确保包含"of"关键字) - 用正则表达式提取总页数,避免手动截取的脆弱性
二、其他代码问题修复
- 移除重复的元素等待逻辑,复用已获取的分页元素
- 调整期刊链接的等待条件为
visibility_of_all_elements_located,确保元素可见后再提取链接 - 修正Aims & Scope的定位XPATH,原定位不准确,改为定位展开后的内容容器
- 完全移除
time.sleep,全部使用显式等待控制页面交互节奏
修改后的完整代码
from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.chrome.service import Service from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By from webdriver_manager.chrome import ChromeDriverManager import re options = Options() options.add_argument("start-maximized") driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options) wait = WebDriverWait(driver, 20) # 初始页面加载 url = "https://www.sciencedirect.com/browse/journals-and-books?accessType=openAccess&accessType=containsOpenAccess" driver.get(url) # 等待分页文本加载完成并提取总页数 page_desc_locator = (By.XPATH, "//span[@class='pagination-pages-label u-margin-s-left-from-sm u-margin-s-right-from-sm']") wait.until(EC.text_to_be_present_in_element(page_desc_locator, "of")) page_description = driver.find_element(*page_desc_locator) # 用正则提取"of"后面的数字 match = re.search(r'of (\d+)', page_description.text) pages = int(match.group(1)) if match else 0 all_journal_links = [] # 遍历所有分页获取期刊链接 for page_num in range(1, pages + 1): current_page_url = f"https://www.sciencedirect.com/browse/journals-and-books?page={page_num}&accessType=containsOpenAccess&accessType=openAccess" driver.get(current_page_url) # 等待所有期刊标题链接可见 journal_links = wait.until(EC.visibility_of_all_elements_located((By.XPATH, "//a[@class='anchor js-publication-title anchor-default']"))) for link in journal_links: href = link.get_attribute('href') if href: all_journal_links.append(href) # 遍历期刊链接提取Aims & Scope for link in all_journal_links: driver.get(link) # 等待"View Full Aims & Scope"按钮可点击并点击 view_button = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[contains(text(), 'View Full Aims & Scope')]"))) view_button.click() # 等待Aims & Scope内容加载并提取 aims_scope_content = wait.until(EC.visibility_of_element_located((By.XPATH, "//div[contains(@class, 'aims-and-scope')]/div[@class='content']"))) print("Aims & Scope:", aims_scope_content.text) driver.quit()
关键改进说明
- 总页数提取:用正则匹配
of (\d+)直接获取数字,避免手动截取字符串的错误风险,同时等待文本包含"of"确保内容加载完成 - 元素等待优化:将期刊链接的等待条件改为
visibility_of_all_elements_located,确保元素可见后再提取链接,避免获取到未渲染完成的元素 - 按钮定位优化:用
contains(text(), 'View Full Aims & Scope')定位按钮,比仅靠class更稳定,避免class变更导致定位失败 - 内容提取优化:直接定位Aims & Scope的内容容器,确保提取到完整的目标文本
内容的提问来源于stack exchange,提问作者trgjk yfojn
相关产品推荐
相关产品推荐

