为何Scrapy-Splash无法返回动态JavaScript页面的预期HTML?
问题分析与解决方案
原因分析
- 目标页面是单页应用(SPA),核心表格内容完全由
app.cf91ad4bfa347ec83220.bundle.js这类前端打包脚本动态渲染到空的rootdiv中,而非服务器端直接返回。 - Splash默认的
render.html端点可能未满足页面渲染的两个关键条件:足够的JS执行等待时间,或未通过网站的浏览器指纹检测(比如Splash默认的User-Agent、渲染环境特征被识别)。 - 部分SPA会依赖滚动、交互等动作触发内容加载,Splash默认的静态等待可能无法触发这些逻辑。
Splash修复尝试
如果想继续使用Splash,可尝试以下优化:
延长等待时间并指定元素等待
替换原SplashRequest的args,直接等待目标表格元素出现,而非固定时间:yield SplashRequest( url, self.response_parser, endpoint='render.html', args={ 'wait_for': 'table', # 等待表格元素加载完成 'wait': 3.0, 'user_agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' }, )使用自定义Lua脚本模拟真实交互
通过execute端点执行Lua脚本,模拟浏览器滚动、等待元素等动作,绕过可能的反爬检测:lua_script = """ function main(splash, args) -- 设置真实浏览器UA splash:set_user_agent('Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36') splash:go(args.url) -- 等待表格元素出现 splash:wait_for_selector('table') -- 滚动页面触发可能的懒加载 splash:runjs('window.scrollTo(0, document.body.scrollHeight)') splash:wait(1) -- 返回渲染后的HTML return splash:html() end """ yield SplashRequest( url, self.response_parser, endpoint='execute', args={'lua_source': lua_script}, )升级Splash版本
旧版Splash对ES6+等现代JS特性支持不足,建议升级到最新稳定版,确保能正确解析目标页面的打包脚本。
替代方案:Scrapy + Playwright
如果Splash始终无法正常渲染,推荐使用Playwright(无头浏览器),其渲染环境更接近真实浏览器,反爬绕过能力更强:
步骤1:安装依赖
pip install scrapy-playwright
步骤2:配置Scrapy Settings
在settings.py中添加以下配置:
DOWNLOAD_HANDLERS = { "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", } PLAYWRIGHT_LAUNCH_OPTIONS = { "headless": True, "args": ["--no-sandbox", "--disable-dev-shm-usage"], }
步骤3:修改爬虫代码
import os import scrapy import datetime from scrapy_playwright.page import PageCoroutine class MarketsSpider(scrapy.Spider): name = "markets" allowed_domains = ["manta.layerbank.finance"] start_urls = ["https://manta.layerbank.finance/bank"] output_directory = "webpageRepository" def start_requests(self): for url in self.start_urls: yield scrapy.Request( url, meta={ "playwright": True, "playwright_page_coroutines": [ PageCoroutine("wait_for_selector", "table"), # 等待表格加载 PageCoroutine("evaluate", "window.scrollTo(0, document.body.scrollHeight)"), # 滚动触发懒加载 PageCoroutine("wait_for_timeout", 1000), # 额外等待1秒确保渲染完成 ], }, callback=self.response_parser, ) def response_parser(self, response): date_time = datetime.datetime.now().strftime('%m-%d-%YT%H:%M:%S') filename = f"bank-page_{date_time}.html" output_path = os.path.join(self.output_directory, filename) os.makedirs(self.output_directory, exist_ok=True) with open(output_path, 'w', encoding='utf-8') as file: file.write(response.text) self.log(f"Saved file {output_path}")
内容的提问来源于stack exchange,提问作者Kody F
相关产品推荐
相关产品推荐

