You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python抓取JavaScript生成的动态DIV内容?BS4无法获取动态数据

解决BeautifulSoup无法抓取JavaScript动态生成内容的问题

我完全懂你遇到的困扰——用requests加BeautifulSoup爬页面时,那些带class="quote"的DIV明明在浏览器里能看到,但代码就是抓不到。这是因为requests只会获取页面的初始静态HTML源码,它不会像浏览器那样执行页面里的JavaScript代码,而这些quote元素正是页面加载后通过JS动态生成的,所以BeautifulSoup自然找不到它们。

下面给你两个实用的解决方案:

方案1:用Selenium模拟真实浏览器

Selenium可以操控Chrome、Firefox这类真实浏览器加载页面,等JS执行完、动态内容渲染出来后,再获取完整的页面源码,这样就能抓到你要的内容了。

操作步骤:

  1. 先安装依赖:

    pip install selenium
    

    另外要下载对应你浏览器版本的驱动(比如ChromeDriver),可以把它放到系统PATH里,或者在代码里指定路径。

  2. 修改后的代码示例:

    from selenium import webdriver
    from selenium.webdriver.chrome.options import Options
    from bs4 import BeautifulSoup
    
    URL = "https://rawgit.com/skysoft999/tableauJS/master/example.html"
    
    # 配置无头模式(不用弹出浏览器窗口,后台运行)
    chrome_options = Options()
    chrome_options.add_argument("--headless=new")
    chrome_options.add_argument("--disable-gpu")
    
    # 启动浏览器驱动
    driver = webdriver.Chrome(options=chrome_options)
    driver.get(URL)
    
    # 获取渲染完成后的页面源码
    page_source = driver.page_source
    soup = BeautifulSoup(page_source, 'html5lib')
    
    # 现在就能抓到动态生成的quote元素了
    for row in soup.find_all('div', attrs={'class':'quote'}):
        print(row.get_text(strip=True))  # 或者直接打印row看完整内容
    
    # 记得关闭浏览器
    driver.quit()
    

方案2:用Playwright(更现代的选择)

Playwright是微软出的自动化工具,API比Selenium更简洁,还能自动管理浏览器驱动,不用你手动下载。

操作步骤:

  1. 安装依赖和浏览器:

    pip install playwright
    playwright install  # 自动安装所需的浏览器包
    
  2. 代码示例:

    from playwright.sync_api import sync_playwright
    from bs4 import BeautifulSoup
    
    URL = "https://rawgit.com/skysoft999/tableauJS/master/example.html"
    
    with sync_playwright() as p:
        # 启动无头Chrome浏览器
        browser = p.chromium.launch(headless=True)
        page = browser.new_page()
        page.goto(URL)
        # 可选:等待目标元素加载完成,确保JS执行完毕
        page.wait_for_selector('div.quote')
        # 获取渲染后的页面源码
        page_source = page.content()
        soup = BeautifulSoup(page_source, 'html5lib')
    
        for row in soup.find_all('div', attrs={'class':'quote'}):
            print(row.get_text(strip=True))
    
        browser.close()
    

额外小技巧:

如果这个网站的动态内容是通过API接口获取的,你可以用浏览器开发者工具(按F12)的Network标签页,找一下页面加载时发出的XHR/fetch请求,直接用requests请求这些API接口拿数据——这种方式比模拟浏览器高效多了,还不容易触发反爬机制。

最后别忘了遵守网站的使用规则,不要过度爬取给服务器添负担哦~

内容的提问来源于stack exchange,提问作者frank hk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:24:06