You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy+Splash无法渲染目标JavaScript网站的问题求助

问题排查与解决方法

可能的问题原因

  • 目标网站反爬检测:sreality.cz可能识别出Splash的无头浏览器特征(如默认User-Agent、特定浏览器指纹),拒绝返回正常渲染内容。
  • ARM架构兼容性问题:你使用的是Apple silicon(ARM架构),官方默认的scrapinghub/splash镜像是x86架构,通过Rosetta转译运行时可能导致JS执行异常,渲染失败。

具体解决方法

1. 绕过反爬检测(针对Splash)

在Splash请求中添加真实浏览器的User-Agent,并延长等待时间确保JS加载完成:

from scrapy_splash import SplashRequest

def start_requests(self):
    url = "https://www.sreality.cz/en/search/for-sale/apartments"
    yield SplashRequest(
        url=url,
        callback=self.parse,
        args={
            # 延长等待时间,确保页面JS渲染完成
            'wait': 5,
            # 模拟真实Mac浏览器的User-Agent
            'headers': {
                'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 13_0_1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
            },
            # 可选:等待特定元素加载完成再返回
            'js_source': """
                function main(splash) {
                    return splash.wait_for_element('.property-list');
                }
            """
        }
    )

2. 更换ARM兼容的Splash镜像

官方提供了ARM64架构的Splash镜像,替换原有镜像重新启动容器:

# 拉取ARM64版本镜像
docker pull scrapinghub/splash:latest-arm64

# 启动容器
docker run -p 8050:8050 scrapinghub/splash:latest-arm64

启动后先访问http://localhost:8050的Splash网页端,测试能否正常渲染目标网站,再重新运行Scrapy爬虫。

3. 替换为Playwright(更稳定的ARM兼容方案)

如果Splash仍有问题,推荐使用Playwright替代,它对ARM架构支持更好,反爬规避能力更强:

步骤1:安装依赖

pip install scrapy-playwright playwright
# 安装Chromium浏览器
playwright install chromium

步骤2:配置Scrapy settings.py

DOWNLOAD_HANDLERS = {
    "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
    "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
}

PLAYWRIGHT_LAUNCH_OPTIONS = {
    "headless": True,
    "args": ["--no-sandbox", "--disable-setuid-sandbox"],
}

步骤3:编写爬虫请求

from scrapy_playwright.page import PageCoroutine
from scrapy_playwright.request import PlaywrightRequest

def start_requests(self):
    url = "https://www.sreality.cz/en/search/for-sale/apartments"
    yield PlaywrightRequest(
        url=url,
        callback=self.parse,
        # 等待目标元素加载完成
        page_coroutines=[
            PageCoroutine("wait_for_selector", ".property-list"),
        ],
        meta={"playwright": True}
    )

内容的提问来源于stack exchange,提问作者jambormike

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 21:35:16