You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium高效抓取动态网站?优化代码获取Ksana Health博客详情

优化后的Selenium博客内容抓取方案

问题根源

你的原代码仅完成了打开页面就立即退出,既没等待JavaScript动态加载博客内容,也没有任何元素定位和数据提取逻辑,自然不会有输出结果。以下是分步优化方案:

1. 修正驱动初始化逻辑

改用更稳定的本地Chrome驱动启动方式,搭配无头模式(适合无界面环境运行),同时禁用冗余功能提升效率:

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
import time

# 配置Chrome运行选项
chrome_options = Options()
chrome_options.add_argument("--headless=new")  # 无头模式,不弹出浏览器窗口
chrome_options.add_argument("--disable-gpu")
chrome_options.add_argument("--no-sandbox")
chrome_options.add_argument("--disable-dev-shm-usage")

# 初始化驱动(注意路径修正:原代码多写了一个s,正确路径为/usr/bin/chromedriver)
service = Service('/usr/bin/chromedriver')
driver = webdriver.Chrome(service=service, options=chrome_options)

2. 等待动态内容加载并抓取博客链接

使用显式等待确保博客列表加载完成,再提取所有详情页链接并去重:

try:
    driver.get('https://ksanahealth.com/mental-health-blog/')
    # 等待博客列表元素加载,最长等待10秒
    WebDriverWait(driver, 10).until(
        EC.presence_of_all_elements_located((By.CSS_SELECTOR, "article a"))
    )
    # 提取并去重博客链接
    blog_links = []
    links = driver.find_elements(By.CSS_SELECTOR, "article a")
    for link in links:
        href = link.get_attribute('href')
        if href and '/mental-health-blog/' in href and href not in blog_links:
            blog_links.append(href)
    print(f"共抓取到{len(blog_links)}篇博客链接")

3. 遍历链接抓取详情页核心数据

逐个打开博客链接,提取标题、发布日期、正文等信息(选择器需根据页面实际结构调整):

# 循环处理每个博客详情页
    for url in blog_links:
        driver.get(url)
        # 等待详情页标题加载完成
        WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.TAG_NAME, "h1"))
        )
        # 提取数据(示例选择器,需根据实际页面结构修改)
        title = driver.find_element(By.TAG_NAME, "h1").text
        date = driver.find_element(By.CSS_SELECTOR, ".post-date").text if driver.find_elements(By.CSS_SELECTOR, ".post-date") else "无发布日期"
        content = driver.find_element(By.CSS_SELECTOR, ".entry-content").text[:200] + "..." if driver.find_elements(By.CSS_SELECTOR, ".entry-content") else "无内容"
        
        # 输出抓取结果(也可写入文件或数据库)
        print(f"标题: {title}")
        print(f"发布日期: {date}")
        print(f"内容摘要: {content}")
        print("-"*50)
        time.sleep(1)  # 增加短暂延迟,避免请求过于频繁触发反爬
finally:
    driver.quit()  # 确保驱动正常退出,释放系统资源

重要提醒

  • 页面元素选择器(如.post-date、.entry-content)需通过浏览器开发者工具查看目标网站的实际HTML结构调整,避免因页面更新导致抓取失败。
  • 若遇到反爬机制,可适当延长等待时间、添加请求头模拟真实用户行为。

内容的提问来源于stack exchange,提问作者Mugoya Dihfahsih

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 10:33:59