Scrapy+Splash爬取Zappos时部分目标内容时隐时现的问题求助
Hey there, let's break down why this module shows up inconsistently and fix it step by step. The issue is likely tied to dynamic loading triggers (like site anti-bot checks or lazy loading rules) or Splash not waiting long enough for content to render. Here are actionable solutions:
1. 等待目标元素加载,而非固定时长
Instead of using a static wait parameter, tell Splash to wait until the target module appears—this is way more reliable than guessing timings.
For example, if your target module uses a selector like div[data-testid="also-viewed"] (adjust this to match the actual element on Zappos), use the wait_for argument in your SplashRequest:
from scrapy_splash import SplashRequest def start_requests(self): url = "https://www.zappos.com/p/chaco-marshall-tartan-rust/product/8982802/color/725500" yield SplashRequest( url, callback=self.parse_product, args={ "wait_for": "div[data-testid='also-viewed']", # 等待目标元素出现 "timeout": 30, # 给足加载时间 } )
2. 模拟页面刷新(匹配手动刷新的行为)
Since you noticed refreshing 2-3 times makes the module appear, we can replicate this in a Splash Lua script. This helps bypass one-time anti-bot checks that block the module on first load.
Create a Lua script to reload the page multiple times before extracting content:
function main(splash, args) -- 模拟真实浏览器的User-Agent splash:set_user_agent(args.user_agent) -- 第一次加载页面 assert(splash:go(args.url)) splash:wait(2) -- 第一次刷新 assert(splash:reload()) splash:wait(2) -- 第二次刷新(可根据实际情况调整刷新次数) assert(splash:reload()) splash:wait(5) -- 等待目标模块渲染完成 assert(splash:wait_for_selector('div[data-testid="also-viewed"]')) return splash:html() end
然后在Scrapy请求中调用这个脚本:
def start_requests(self): url = "https://www.zappos.com/p/chaco-marshall-tartan-rust/product/8982802/color/725500" lua_script = """ -- 粘贴上面的Lua脚本内容 """ yield SplashRequest( url, callback=self.parse_product, endpoint="execute", # 使用execute端点运行自定义Lua脚本 args={ "lua_source": lua_script, "user_agent": self.settings.get("USER_AGENT"), } )
3. 模拟用户滚动(触发懒加载)
部分电商网站的“看过还看”模块需要用户滚动页面才会加载。可以在Lua脚本中加入滚动动作触发加载:
-- 在首次加载/刷新后添加滚动逻辑 splash:runjs('window.scrollTo(0, document.body.scrollHeight);') splash:wait(3) # 等待懒加载内容加载完成
4. 确保请求头和Cookie符合真实浏览器
Zappos可能会识别非浏览器样式的请求。确保Splash使用有效的User-Agent,必要时携带之前请求的Cookie。可以在Lua脚本中直接设置Cookie:
splash:set_cookie{ name = "some_cookie_name", value = "cookie_value", domain = "zappos.com", path = "/" }
5. 增加自定义重试逻辑
在Scrapy爬虫中添加重试机制,重新处理未加载目标模块的页面。检查响应中是否存在目标元素,若缺失则触发重试:
def parse_product(self, response): also_viewed = response.css('div[data-testid="also-viewed"]') if not also_viewed: retry_times = response.meta.get("retry_times", 0) if retry_times < 3: # 最多重试3次 yield response.request.replace( meta={"retry_times": retry_times + 1}, dont_filter=True # 绕过重复请求过滤 ) return # 处理“看过还看”模块数据 for item in also_viewed.css(".product-item"): # 提取商品信息...
最后提示
先从方案1(等待目标元素)开始尝试,因为它最简单直接。如果无效,再尝试模拟刷新或滚动动作。记得用浏览器开发者工具验证选择器是否正确指向目标模块。
内容的提问来源于stack exchange,提问作者jkatt

