WebScraping问题:requests-html的Chromium无法加载JS,Chrome正常,如何处理?
问题与解决方案
问题描述
- 使用
requests-html模块爬取https://mphonline.com/collections/architecture-landscaping时,依赖的Chromium无法加载JavaScript内容,触发超时错误 - 已尝试调用
html.render()并将超时时间设为30秒,但问题依旧;但使用Selenium搭配Chrome WebDriver的脚本可正常运行 - 手动启动
%localappdata%\pyppeteer\pyppeteer\local-chromium路径下的Chromium(版本71.0.3542.0,脚本首次运行时自动安装)访问目标URL,发现JavaScript加载极慢
解决方案
1. 升级Chromium版本
requests-html依赖的pyppeteer默认安装的Chromium版本过旧(71.x),对现代网站的JavaScript兼容性差,这是加载缓慢的核心原因。可以手动指定新版本Chromium的路径:
- 先下载对应系统的最新Chromium稳定版
- 在脚本中通过
pyppeteer_kwargs的executablePath参数指定新版本路径:
from requests_html import HTMLSession webURL = "https://mphonline.com/collections/architecture-landscaping" session = HTMLSession(pyppeteer_kwargs = { 'handleSIGINT' : False, 'handleSIGTERM': False, 'handleSIGHUP': False, 'headless': False, 'executablePath': r'C:\你的路径\chromium.exe' # 替换为实际的新版本Chromium路径 }) root = session.get(webURL) root.html.render(timeout=60, keep_page=True, wait=5) # wait参数让页面稳定后再操作 titlesxpath = "//div[contains(@class, 'boost-sd__product-title')]" titles = root.html.xpath(titlesxpath) for title in titles: print(title.text) session.close()
2. 优化渲染参数,给JS足够加载时间
- 延长
timeout参数至60秒以上,适配旧版Chromium的慢加载速度 - 增加
wait参数,指定页面初始加载后等待的秒数,确保JavaScript有足够时间渲染内容 - 如果目标网站存在懒加载内容,可添加
scrolldown参数模拟滚动触发加载:
root.html.render(timeout=60, keep_page=True, wait=5, scrolldown=3)
3. 禁用冗余特性,降低资源占用
在pyppeteer_kwargs中添加args参数,关闭图片加载、GPU加速等非必要特性,减少Chromium的资源消耗,提升加载速度:
session = HTMLSession(pyppeteer_kwargs = { 'handleSIGINT' : False, 'handleSIGTERM': False, 'handleSIGHUP': False, 'headless': False, 'args': [ '--disable-images', '--disable-gpu', '--no-sandbox', '--disable-dev-shm-usage', '--disable-extensions' ] })
4. 等待目标元素加载完成后再提取
通过pyppeteer的页面等待方法,强制等待目标元素出现后再进行内容提取,避免因JS未渲染完成导致的空结果:
import asyncio from requests_html import HTMLSession async def main(): webURL = "https://mphonline.com/collections/architecture-landscaping" session = HTMLSession(pyppeteer_kwargs = { 'handleSIGINT' : False, 'handleSIGTERM': False, 'handleSIGHUP': False, 'headless': False, }) root = session.get(webURL) page = root.html.page # 等待目标元素出现,超时时间设为60秒 await page.waitForXPath("//div[contains(@class, 'boost-sd__product-title')]", timeout=60000) root.html.render(timeout=60, keep_page=True) titlesxpath = "//div[contains(@class, 'boost-sd__product-title')]" titles = root.html.xpath(titlesxpath) for title in titles: print(title.text) session.close() asyncio.run(main())
内容的提问来源于stack exchange,提问作者Jason
相关产品推荐
相关产品推荐

