You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Windows 10下如何安装Splash?Scrapy爬虫需用到该工具

Fixing Scrapy-Splash Installation Issues on Windows

Hey there! I’ve struggled with Scrapy-Splash on Windows too, so let’s walk through three solid solutions to get you up and running with JS-rendered content scraping.

Solution 1: Fix PyQt4 Errors in Windows Subsystem for Linux (WSL)

The PyQt4 error happens because modern WSL distros (like Ubuntu 20.04+) don’t ship with PyQt4 by default, and the old Splash install scripts rely on it. Here’s how to patch this:

  1. First, update your WSL packages:
    sudo apt update && sudo apt upgrade -y
    
  2. Install the newer PyQt5 dependencies (Splash works with PyQt5 now, you just need to tweak the install process):
    sudo apt install python3-pyqt5 python3-pyqt5.qtwebkit python3-pip git -y
    
  3. Clone the Splash repo and modify the setup script to use PyQt5 instead of PyQt4:
    git clone https://github.com/scrapinghub/splash.git
    cd splash
    
    Open setup.py in a text editor (like nano setup.py) and replace all instances of PyQt4 with PyQt5. Save and exit.
  4. Install Splash using pip:
    pip3 install .
    
  5. Test if Splash runs:
    splash
    
    If it starts up without errors, you can connect to it via http://localhost:8050 from your Windows browser.

Solution 2: Get Docker + Linux Containers Working on Windows

Docker’s Linux container error usually comes from not having WSL2 enabled as the backend. Here’s how to fix that:

  1. Enable WSL2 (if not already done):
    Open PowerShell as Administrator and run:

    wsl --install
    

    Follow the prompts to restart your PC and set up a Linux distro (like Ubuntu).

  2. Configure Docker Desktop to use WSL2:

    • Open Docker Desktop, go to Settings > Resources > WSL Integration.
    • Toggle on the switch for your installed WSL distro (e.g., Ubuntu).
    • Click "Apply & Restart" to save changes.
  3. Pull and run the Splash image:
    Open a WSL terminal (or Windows Command Prompt/PowerShell) and run:

    docker pull scrapinghub/splash
    docker run -p 8050:8050 scrapinghub/splash
    

    Now you can access Splash at http://localhost:8050 and use it with your Scrapy project.

Solution 3: Ditch Splash for Scrapy-Playwright (Easier Windows Alternative)

If you want to avoid the hassle of Splash entirely, Scrapy-Playwright is a fantastic alternative with better Windows support. It uses Playwright to handle JS rendering, and setup is straightforward:

  1. Install the package:
    pip install scrapy-playwright
    
  2. Enable Playwright in your Scrapy project’s settings.py:
    PLAYWRIGHT_ENABLED = True
    DOWNLOAD_HANDLERS = {
        "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
        "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
    }
    
  3. Use it in your spider by adding playwright: True to request metadata:
    def start_requests(self):
        yield scrapy.Request(
            url="your-js-rendered-url",
            meta={"playwright": True},
            callback=self.parse
        )
    
    Playwright will automatically render the JS content before passing it to your parse function.

All three methods work reliably on Windows—my personal pick is Scrapy-Playwright because it’s less maintenance-heavy long-term.

内容的提问来源于stack exchange,提问作者Bathe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:45:01