Selenium Chrome Driver page_source与控制台元素不一致的爬取问题
解决Medium文章爬取时Selenium无法获取渲染后内容的问题
针对你遇到的Selenium爬取Medium文章时,driver.page_source找不到目标内容容器的问题,我整理了几个实际爬取Medium时验证有效的解决方案,都是踩过坑后总结的经验:
1. 先换掉动态生成的类名选择器
你用的.x.y.z.ab.ac.ez.af.ag这类类名是Medium动态随机生成的,每次刷新页面都会变化,完全没法作为固定选择器使用——这才是最核心的问题!
换成基于页面结构的稳定定位方式:
- 直接定位
article标签(Medium文章内容基本都包裹在这个标签里):from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By from bs4 import BeautifulSoup ARTICLE = "https://medium.com/@adrianmarkperea/demystifying-python-decorators-in-10-minutes-ffe092723c6c" driver.get(ARTICLE) # 等待article标签加载完成 WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.TAG_NAME, "article"))) text_soup = BeautifulSoup(driver.page_source,"html5lib") article = text_soup.find("article") if article: # 提取所有段落文本并拼接 full_content = "\n".join([p.get_text(strip=True) for p in article.find_all("p")]) print(full_content) - 或者用XPath定位语义化的内容区域(比如找带测试ID的容器):
content_element = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.XPATH, "//div[contains(@data-testid, 'post-content')]")) ) # 直接获取元素的HTML,比page_source更精准 content_html = content_element.get_attribute("outerHTML") text_soup = BeautifulSoup(content_html,"html5lib")
2. 处理Medium的反爬与动态加载
Medium会检测机器人,而且长文章需要滚动才能完全渲染,这会导致page_source没有包含完整内容:
- 模拟真实用户滚动行为:
driver.get(ARTICLE) # 滚动到页面底部触发剩余内容加载 driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # 等待几秒让内容渲染完成 time.sleep(3) # 长文章可以多滚动一次 driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(2) - 配置ChromeDriver规避反爬检测:
from selenium.webdriver.chrome.options import Options options = Options() # 无头模式下必须加这些参数模拟真实浏览器 options.add_argument("--headless=new") options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") # 禁用自动化特征检测 options.add_argument("--disable-blink-features=AutomationControlled") options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option('useAutomationExtension', False) driver = webdriver.Chrome(options=options)
3. 确保等待逻辑正确
虽然你排除了超时问题,但如果用的是time.sleep()硬等待,可能还是会出现页面未完全渲染就获取源码的情况。改用WebDriverWait等待目标元素出现更可靠:
WebDriverWait(driver, 15).until( EC.presence_of_element_located((By.TAG_NAME, "article")) ) # 确认元素加载后再获取page_source text_soup = BeautifulSoup(driver.page_source,"html5lib")
最后验证思路
你本地测试静态HTML正常,说明Selenium本身能处理JS渲染的内容,问题肯定出在Medium的动态类名或反爬机制上。先换稳定的定位器,再配合反反爬配置,就能拿到和Chrome控制台一致的内容了。
内容的提问来源于stack exchange,提问作者lazmond3
相关产品推荐
相关产品推荐

