You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium Chrome Driver page_source与控制台元素不一致的爬取问题

解决Medium文章爬取时Selenium无法获取渲染后内容的问题

针对你遇到的Selenium爬取Medium文章时,driver.page_source找不到目标内容容器的问题,我整理了几个实际爬取Medium时验证有效的解决方案,都是踩过坑后总结的经验:

1. 先换掉动态生成的类名选择器

你用的.x.y.z.ab.ac.ez.af.ag这类类名是Medium动态随机生成的,每次刷新页面都会变化,完全没法作为固定选择器使用——这才是最核心的问题!

换成基于页面结构的稳定定位方式:

  • 直接定位article标签(Medium文章内容基本都包裹在这个标签里):
    from selenium.webdriver.support.ui import WebDriverWait
    from selenium.webdriver.support import expected_conditions as EC
    from selenium.webdriver.common.by import By
    from bs4 import BeautifulSoup
    
    ARTICLE = "https://medium.com/@adrianmarkperea/demystifying-python-decorators-in-10-minutes-ffe092723c6c"
    driver.get(ARTICLE)
    
    # 等待article标签加载完成
    WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.TAG_NAME, "article")))
    
    text_soup = BeautifulSoup(driver.page_source,"html5lib")
    article = text_soup.find("article")
    if article:
        # 提取所有段落文本并拼接
        full_content = "\n".join([p.get_text(strip=True) for p in article.find_all("p")])
        print(full_content)
    
  • 或者用XPath定位语义化的内容区域(比如找带测试ID的容器):
    content_element = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.XPATH, "//div[contains(@data-testid, 'post-content')]"))
    )
    # 直接获取元素的HTML,比page_source更精准
    content_html = content_element.get_attribute("outerHTML")
    text_soup = BeautifulSoup(content_html,"html5lib")
    

2. 处理Medium的反爬与动态加载

Medium会检测机器人,而且长文章需要滚动才能完全渲染,这会导致page_source没有包含完整内容:

  • 模拟真实用户滚动行为:
    driver.get(ARTICLE)
    # 滚动到页面底部触发剩余内容加载
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    # 等待几秒让内容渲染完成
    time.sleep(3)
    # 长文章可以多滚动一次
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(2)
    
  • 配置ChromeDriver规避反爬检测:
    from selenium.webdriver.chrome.options import Options
    
    options = Options()
    # 无头模式下必须加这些参数模拟真实浏览器
    options.add_argument("--headless=new")
    options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
    # 禁用自动化特征检测
    options.add_argument("--disable-blink-features=AutomationControlled")
    options.add_experimental_option("excludeSwitches", ["enable-automation"])
    options.add_experimental_option('useAutomationExtension', False)
    
    driver = webdriver.Chrome(options=options)
    

3. 确保等待逻辑正确

虽然你排除了超时问题,但如果用的是time.sleep()硬等待,可能还是会出现页面未完全渲染就获取源码的情况。改用WebDriverWait等待目标元素出现更可靠:

WebDriverWait(driver, 15).until(
    EC.presence_of_element_located((By.TAG_NAME, "article"))
)
# 确认元素加载后再获取page_source
text_soup = BeautifulSoup(driver.page_source,"html5lib")

最后验证思路

你本地测试静态HTML正常,说明Selenium本身能处理JS渲染的内容,问题肯定出在Medium的动态类名或反爬机制上。先换稳定的定位器,再配合反反爬配置,就能拿到和Chrome控制台一致的内容了。

内容的提问来源于stack exchange,提问作者lazmond3

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:20:50