使用Python+Selenium提取Sciencedirect机构信息返回空输出问题排查
排查Sciencedirect机构信息提取代码返回空输出的问题
代码返回空输出的核心原因是Sciencedirect页面结构更新,导致原代码依赖的元素定位器(XPATH、CSS选择器)失效,另外固定的time.sleep也可能因网络延迟导致元素未加载完成就执行后续操作。以下是具体排查点和修正方案:
1. Show more按钮定位失效
原代码用//span[@class="button-link-text" and contains(text(), "Show more")]定位按钮,现在该元素的标签、class或文本可能已变更:
- 按钮可能从
<span>变为<button>元素 - class名可能更新(比如变为
show-more-link) - 页面加入动态加载逻辑后,固定
time.sleep(3)不足以等待按钮加载完成
2. 机构信息选择器失效
原代码用[class="AuthorGroups text-s"] dl dd提取机构信息,现在页面的类名可能已修改(比如AuthorGroups改为affiliation-group),导致BeautifulSoup无法匹配到目标元素。
修正后的代码示例
改用WebDriverWait等待元素加载(替代不稳定的time.sleep),并根据当前页面结构更新定位器:
import time from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup # 初始化Chrome浏览器 service = Service(r"Z:\Private\hbasamh\ACWA Power\Files\Jupyter\Web Scraping\chromedriver.exe") options = webdriver.ChromeOptions() options.add_argument("start-maximized") driver = webdriver.Chrome(service=service, options=options) url = 'https://www.sciencedirect.com/science/article/abs/pii/S0011916424004600' driver.get(url) try: # 等待Show more按钮可点击并点击(根据当前页面元素调整定位) show_more_btn = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, '//button[contains(@class, "show-more-link") and contains(text(), "Show more")]')) ) show_more_btn.click() # 等待机构信息区域加载完成 WebDriverWait(driver, 5).until( EC.presence_of_element_located((By.CSS_SELECTOR, 'div.affiliation-group dl')) ) except Exception as e: print(f"加载或点击按钮出错: {e}") # 解析页面提取机构信息 soup = BeautifulSoup(driver.page_source, "html.parser") txt = [x.get_text().strip() for x in soup.select('div.affiliation-group dl dd')] print(txt) driver.quit()
关键调整说明
- 替换定位器:打开目标页面,用浏览器开发者工具(F12)查看Show more按钮和机构信息区域的最新HTML结构,更新XPATH/CSS选择器
- 用WebDriverWait替代time.sleep:根据元素加载状态动态等待,避免因网络波动导致的元素未加载问题
内容的提问来源于stack exchange,提问作者user17356493
相关产品推荐
相关产品推荐

