You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何Scrapy-Splash无法返回动态JavaScript页面的预期HTML?

问题分析与解决方案

原因分析

  • 目标页面是单页应用(SPA),核心表格内容完全由app.cf91ad4bfa347ec83220.bundle.js这类前端打包脚本动态渲染到空的root div中,而非服务器端直接返回。
  • Splash默认的render.html端点可能未满足页面渲染的两个关键条件:足够的JS执行等待时间,或未通过网站的浏览器指纹检测(比如Splash默认的User-Agent、渲染环境特征被识别)。
  • 部分SPA会依赖滚动、交互等动作触发内容加载,Splash默认的静态等待可能无法触发这些逻辑。

Splash修复尝试

如果想继续使用Splash,可尝试以下优化:

  1. 延长等待时间并指定元素等待
    替换原SplashRequest的args,直接等待目标表格元素出现,而非固定时间:

    yield SplashRequest(
        url,
        self.response_parser,
        endpoint='render.html',
        args={
            'wait_for': 'table',  # 等待表格元素加载完成
            'wait': 3.0,
            'user_agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
        },
    )
    
  2. 使用自定义Lua脚本模拟真实交互
    通过execute端点执行Lua脚本,模拟浏览器滚动、等待元素等动作,绕过可能的反爬检测:

    lua_script = """
    function main(splash, args)
      -- 设置真实浏览器UA
      splash:set_user_agent('Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36')
      splash:go(args.url)
      -- 等待表格元素出现
      splash:wait_for_selector('table')
      -- 滚动页面触发可能的懒加载
      splash:runjs('window.scrollTo(0, document.body.scrollHeight)')
      splash:wait(1)
      -- 返回渲染后的HTML
      return splash:html()
    end
    """
    
    yield SplashRequest(
        url,
        self.response_parser,
        endpoint='execute',
        args={'lua_source': lua_script},
    )
    
  3. 升级Splash版本
    旧版Splash对ES6+等现代JS特性支持不足,建议升级到最新稳定版,确保能正确解析目标页面的打包脚本。

替代方案:Scrapy + Playwright

如果Splash始终无法正常渲染,推荐使用Playwright(无头浏览器),其渲染环境更接近真实浏览器,反爬绕过能力更强:

步骤1:安装依赖

pip install scrapy-playwright

步骤2:配置Scrapy Settings

在settings.py中添加以下配置:

DOWNLOAD_HANDLERS = {
    "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
    "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
}
PLAYWRIGHT_LAUNCH_OPTIONS = {
    "headless": True,
    "args": ["--no-sandbox", "--disable-dev-shm-usage"],
}

步骤3:修改爬虫代码

import os
import scrapy
import datetime
from scrapy_playwright.page import PageCoroutine

class MarketsSpider(scrapy.Spider):
    name = "markets"
    allowed_domains = ["manta.layerbank.finance"]
    start_urls = ["https://manta.layerbank.finance/bank"]
    output_directory = "webpageRepository"

    def start_requests(self):
        for url in self.start_urls:
            yield scrapy.Request(
                url,
                meta={
                    "playwright": True,
                    "playwright_page_coroutines": [
                        PageCoroutine("wait_for_selector", "table"),  # 等待表格加载
                        PageCoroutine("evaluate", "window.scrollTo(0, document.body.scrollHeight)"),  # 滚动触发懒加载
                        PageCoroutine("wait_for_timeout", 1000),  # 额外等待1秒确保渲染完成
                    ],
                },
                callback=self.response_parser,
            )

    def response_parser(self, response):
        date_time = datetime.datetime.now().strftime('%m-%d-%YT%H:%M:%S')
        filename = f"bank-page_{date_time}.html"
        output_path = os.path.join(self.output_directory, filename)
        os.makedirs(self.output_directory, exist_ok=True)
        with open(output_path, 'w', encoding='utf-8') as file:
            file.write(response.text)
        self.log(f"Saved file {output_path}")

内容的提问来源于stack exchange,提问作者Kody F

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 13:27:52