如何用Python抓取JavaScript生成的动态DIV内容?BS4无法获取动态数据
解决BeautifulSoup无法抓取JavaScript动态生成内容的问题
我完全懂你遇到的困扰——用requests加BeautifulSoup爬页面时,那些带class="quote"的DIV明明在浏览器里能看到,但代码就是抓不到。这是因为requests只会获取页面的初始静态HTML源码,它不会像浏览器那样执行页面里的JavaScript代码,而这些quote元素正是页面加载后通过JS动态生成的,所以BeautifulSoup自然找不到它们。
下面给你两个实用的解决方案:
方案1:用Selenium模拟真实浏览器
Selenium可以操控Chrome、Firefox这类真实浏览器加载页面,等JS执行完、动态内容渲染出来后,再获取完整的页面源码,这样就能抓到你要的内容了。
操作步骤:
先安装依赖:
pip install selenium另外要下载对应你浏览器版本的驱动(比如ChromeDriver),可以把它放到系统PATH里,或者在代码里指定路径。
修改后的代码示例:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup URL = "https://rawgit.com/skysoft999/tableauJS/master/example.html" # 配置无头模式(不用弹出浏览器窗口,后台运行) chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("--disable-gpu") # 启动浏览器驱动 driver = webdriver.Chrome(options=chrome_options) driver.get(URL) # 获取渲染完成后的页面源码 page_source = driver.page_source soup = BeautifulSoup(page_source, 'html5lib') # 现在就能抓到动态生成的quote元素了 for row in soup.find_all('div', attrs={'class':'quote'}): print(row.get_text(strip=True)) # 或者直接打印row看完整内容 # 记得关闭浏览器 driver.quit()
方案2:用Playwright(更现代的选择)
Playwright是微软出的自动化工具,API比Selenium更简洁,还能自动管理浏览器驱动,不用你手动下载。
操作步骤:
安装依赖和浏览器:
pip install playwright playwright install # 自动安装所需的浏览器包代码示例:
from playwright.sync_api import sync_playwright from bs4 import BeautifulSoup URL = "https://rawgit.com/skysoft999/tableauJS/master/example.html" with sync_playwright() as p: # 启动无头Chrome浏览器 browser = p.chromium.launch(headless=True) page = browser.new_page() page.goto(URL) # 可选:等待目标元素加载完成,确保JS执行完毕 page.wait_for_selector('div.quote') # 获取渲染后的页面源码 page_source = page.content() soup = BeautifulSoup(page_source, 'html5lib') for row in soup.find_all('div', attrs={'class':'quote'}): print(row.get_text(strip=True)) browser.close()
额外小技巧:
如果这个网站的动态内容是通过API接口获取的,你可以用浏览器开发者工具(按F12)的Network标签页,找一下页面加载时发出的XHR/fetch请求,直接用requests请求这些API接口拿数据——这种方式比模拟浏览器高效多了,还不容易触发反爬机制。
最后别忘了遵守网站的使用规则,不要过度爬取给服务器添负担哦~
内容的提问来源于stack exchange,提问作者frank hk
相关产品推荐
相关产品推荐

