You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

WebScraping问题:requests-html的Chromium无法加载JS,Chrome正常,如何处理?

问题与解决方案

问题描述

  • 使用requests-html模块爬取https://mphonline.com/collections/architecture-landscaping时,依赖的Chromium无法加载JavaScript内容,触发超时错误
  • 已尝试调用html.render()并将超时时间设为30秒,但问题依旧;但使用Selenium搭配Chrome WebDriver的脚本可正常运行
  • 手动启动%localappdata%\pyppeteer\pyppeteer\local-chromium路径下的Chromium(版本71.0.3542.0,脚本首次运行时自动安装)访问目标URL,发现JavaScript加载极慢

解决方案

1. 升级Chromium版本

requests-html依赖的pyppeteer默认安装的Chromium版本过旧(71.x),对现代网站的JavaScript兼容性差,这是加载缓慢的核心原因。可以手动指定新版本Chromium的路径:

  • 先下载对应系统的最新Chromium稳定版
  • 在脚本中通过pyppeteer_kwargs的executablePath参数指定新版本路径:
from requests_html import HTMLSession

webURL = "https://mphonline.com/collections/architecture-landscaping"
session = HTMLSession(pyppeteer_kwargs = {
    'handleSIGINT' : False,
    'handleSIGTERM': False,
    'handleSIGHUP': False,
    'headless': False,
    'executablePath': r'C:\你的路径\chromium.exe'  # 替换为实际的新版本Chromium路径
})

root = session.get(webURL)
root.html.render(timeout=60, keep_page=True, wait=5)  # wait参数让页面稳定后再操作

titlesxpath = "//div[contains(@class, 'boost-sd__product-title')]"
titles = root.html.xpath(titlesxpath)
for title in titles:
    print(title.text)

session.close()

2. 优化渲染参数,给JS足够加载时间

  • 延长timeout参数至60秒以上,适配旧版Chromium的慢加载速度
  • 增加wait参数,指定页面初始加载后等待的秒数,确保JavaScript有足够时间渲染内容
  • 如果目标网站存在懒加载内容,可添加scrolldown参数模拟滚动触发加载:
root.html.render(timeout=60, keep_page=True, wait=5, scrolldown=3)

3. 禁用冗余特性,降低资源占用

在pyppeteer_kwargs中添加args参数,关闭图片加载、GPU加速等非必要特性,减少Chromium的资源消耗,提升加载速度:

session = HTMLSession(pyppeteer_kwargs = {
    'handleSIGINT' : False,
    'handleSIGTERM': False,
    'handleSIGHUP': False,
    'headless': False,
    'args': [
        '--disable-images',
        '--disable-gpu',
        '--no-sandbox',
        '--disable-dev-shm-usage',
        '--disable-extensions'
    ]
})

4. 等待目标元素加载完成后再提取

通过pyppeteer的页面等待方法,强制等待目标元素出现后再进行内容提取,避免因JS未渲染完成导致的空结果:

import asyncio
from requests_html import HTMLSession

async def main():
    webURL = "https://mphonline.com/collections/architecture-landscaping"
    session = HTMLSession(pyppeteer_kwargs = {
        'handleSIGINT' : False,
        'handleSIGTERM': False,
        'handleSIGHUP': False,
        'headless': False,
    })

    root = session.get(webURL)
    page = root.html.page
    # 等待目标元素出现,超时时间设为60秒
    await page.waitForXPath("//div[contains(@class, 'boost-sd__product-title')]", timeout=60000)
    root.html.render(timeout=60, keep_page=True)

    titlesxpath = "//div[contains(@class, 'boost-sd__product-title')]"
    titles = root.html.xpath(titlesxpath)
    for title in titles:
        print(title.text)

    session.close()

asyncio.run(main())

内容的提问来源于stack exchange,提问作者Jason

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 08:51:00