如何用Selenium高效抓取动态网站?优化代码获取Ksana Health博客详情
优化后的Selenium博客内容抓取方案
问题根源
你的原代码仅完成了打开页面就立即退出,既没等待JavaScript动态加载博客内容,也没有任何元素定位和数据提取逻辑,自然不会有输出结果。以下是分步优化方案:
1. 修正驱动初始化逻辑
改用更稳定的本地Chrome驱动启动方式,搭配无头模式(适合无界面环境运行),同时禁用冗余功能提升效率:
from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By import time # 配置Chrome运行选项 chrome_options = Options() chrome_options.add_argument("--headless=new") # 无头模式,不弹出浏览器窗口 chrome_options.add_argument("--disable-gpu") chrome_options.add_argument("--no-sandbox") chrome_options.add_argument("--disable-dev-shm-usage") # 初始化驱动(注意路径修正:原代码多写了一个s,正确路径为/usr/bin/chromedriver) service = Service('/usr/bin/chromedriver') driver = webdriver.Chrome(service=service, options=chrome_options)
2. 等待动态内容加载并抓取博客链接
使用显式等待确保博客列表加载完成,再提取所有详情页链接并去重:
try: driver.get('https://ksanahealth.com/mental-health-blog/') # 等待博客列表元素加载,最长等待10秒 WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.CSS_SELECTOR, "article a")) ) # 提取并去重博客链接 blog_links = [] links = driver.find_elements(By.CSS_SELECTOR, "article a") for link in links: href = link.get_attribute('href') if href and '/mental-health-blog/' in href and href not in blog_links: blog_links.append(href) print(f"共抓取到{len(blog_links)}篇博客链接")
3. 遍历链接抓取详情页核心数据
逐个打开博客链接,提取标题、发布日期、正文等信息(选择器需根据页面实际结构调整):
# 循环处理每个博客详情页 for url in blog_links: driver.get(url) # 等待详情页标题加载完成 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.TAG_NAME, "h1")) ) # 提取数据(示例选择器,需根据实际页面结构修改) title = driver.find_element(By.TAG_NAME, "h1").text date = driver.find_element(By.CSS_SELECTOR, ".post-date").text if driver.find_elements(By.CSS_SELECTOR, ".post-date") else "无发布日期" content = driver.find_element(By.CSS_SELECTOR, ".entry-content").text[:200] + "..." if driver.find_elements(By.CSS_SELECTOR, ".entry-content") else "无内容" # 输出抓取结果(也可写入文件或数据库) print(f"标题: {title}") print(f"发布日期: {date}") print(f"内容摘要: {content}") print("-"*50) time.sleep(1) # 增加短暂延迟,避免请求过于频繁触发反爬 finally: driver.quit() # 确保驱动正常退出,释放系统资源
重要提醒
- 页面元素选择器(如
.post-date、.entry-content)需通过浏览器开发者工具查看目标网站的实际HTML结构调整,避免因页面更新导致抓取失败。 - 若遇到反爬机制,可适当延长等待时间、添加请求头模拟真实用户行为。
内容的提问来源于stack exchange,提问作者Mugoya Dihfahsih
相关产品推荐
相关产品推荐

