使用Selenium与BeautifulSoup爬取网页数据失败求助
问题解决思路与修正代码
核心问题排查
- 缺失页面访问步骤:代码未调用
driver.get()加载目标网页,直接查找元素必然返回空 - 无头模式特征暴露:默认无头模式易被网站反爬机制识别,导致页面加载异常
- 未等待动态元素:页面内容是动态渲染的,直接获取元素时可能还未加载完成
- 定位逻辑冗余:无需先找父级
ul再遍历li,可直接定位目标class的元素
修正后的Selenium代码
from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 配置Chrome选项,规避反爬检测 options = Options() options.add_argument("--headless=new") # 新版无头模式,更贴近正常浏览器行为 options.add_argument("--disable-blink-features=AutomationControlled") options.add_argument("--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") # 初始化浏览器并加载目标页面 driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options) target_url = "https://www.sciencedirect.com/browse/journals-and-books?accessType=openAccess&accessType=containsOpenAccess" driver.get(target_url) # 等待目标元素加载完成(最长等待10秒) wait = WebDriverWait(driver, 10) publications = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, ".publication.u-padding-xs-ver.js-publication"))) # 遍历提取内容 count = 1 for pub in publications: content = pub.text.strip() print(f"Links {count} {content}") count += 1 driver.quit()
补充方案:结合BeautifulSoup解析
如果偏好使用BeautifulSoup,可先通过Selenium获取渲染后的页面源码,再解析:
# 接上述代码,在driver.get(target_url)后添加 page_source = driver.page_source from bs4 import BeautifulSoup soup = BeautifulSoup(page_source, "html.parser") publications = soup.find_all(class_="publication u-padding-xs-ver js-publication") for idx, pub in enumerate(publications, 1): print(f"Links {idx} {pub.get_text(strip=True)}")
内容的提问来源于stack exchange,提问作者trgjk yfojn
相关产品推荐
相关产品推荐

