Windows 10下如何安装Splash?Scrapy爬虫需用到该工具
Hey there! I’ve struggled with Scrapy-Splash on Windows too, so let’s walk through three solid solutions to get you up and running with JS-rendered content scraping.
Solution 1: Fix PyQt4 Errors in Windows Subsystem for Linux (WSL)
The PyQt4 error happens because modern WSL distros (like Ubuntu 20.04+) don’t ship with PyQt4 by default, and the old Splash install scripts rely on it. Here’s how to patch this:
- First, update your WSL packages:
sudo apt update && sudo apt upgrade -y - Install the newer PyQt5 dependencies (Splash works with PyQt5 now, you just need to tweak the install process):
sudo apt install python3-pyqt5 python3-pyqt5.qtwebkit python3-pip git -y - Clone the Splash repo and modify the setup script to use PyQt5 instead of PyQt4:
Opengit clone https://github.com/scrapinghub/splash.git cd splashsetup.pyin a text editor (likenano setup.py) and replace all instances ofPyQt4withPyQt5. Save and exit. - Install Splash using pip:
pip3 install . - Test if Splash runs:
If it starts up without errors, you can connect to it viasplashhttp://localhost:8050from your Windows browser.
Solution 2: Get Docker + Linux Containers Working on Windows
Docker’s Linux container error usually comes from not having WSL2 enabled as the backend. Here’s how to fix that:
Enable WSL2 (if not already done):
Open PowerShell as Administrator and run:wsl --installFollow the prompts to restart your PC and set up a Linux distro (like Ubuntu).
Configure Docker Desktop to use WSL2:
- Open Docker Desktop, go to
Settings > Resources > WSL Integration. - Toggle on the switch for your installed WSL distro (e.g., Ubuntu).
- Click "Apply & Restart" to save changes.
- Open Docker Desktop, go to
Pull and run the Splash image:
Open a WSL terminal (or Windows Command Prompt/PowerShell) and run:docker pull scrapinghub/splash docker run -p 8050:8050 scrapinghub/splashNow you can access Splash at
http://localhost:8050and use it with your Scrapy project.
Solution 3: Ditch Splash for Scrapy-Playwright (Easier Windows Alternative)
If you want to avoid the hassle of Splash entirely, Scrapy-Playwright is a fantastic alternative with better Windows support. It uses Playwright to handle JS rendering, and setup is straightforward:
- Install the package:
pip install scrapy-playwright - Enable Playwright in your Scrapy project’s
settings.py:PLAYWRIGHT_ENABLED = True DOWNLOAD_HANDLERS = { "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", } - Use it in your spider by adding
playwright: Trueto request metadata:
Playwright will automatically render the JS content before passing it to your parse function.def start_requests(self): yield scrapy.Request( url="your-js-rendered-url", meta={"playwright": True}, callback=self.parse )
All three methods work reliably on Windows—my personal pick is Scrapy-Playwright because it’s less maintenance-heavy long-term.
内容的提问来源于stack exchange,提问作者Bathe

