You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy+Splash爬取Zappos时部分目标内容时隐时现的问题求助

解决Scrapy+Splash爬取Zappos“浏览过当前商品的用户还浏览过”模块加载不稳定问题

Hey there, let's break down why this module shows up inconsistently and fix it step by step. The issue is likely tied to dynamic loading triggers (like site anti-bot checks or lazy loading rules) or Splash not waiting long enough for content to render. Here are actionable solutions:

1. 等待目标元素加载,而非固定时长

Instead of using a static wait parameter, tell Splash to wait until the target module appears—this is way more reliable than guessing timings.

For example, if your target module uses a selector like div[data-testid="also-viewed"] (adjust this to match the actual element on Zappos), use the wait_for argument in your SplashRequest:

from scrapy_splash import SplashRequest

def start_requests(self):
    url = "https://www.zappos.com/p/chaco-marshall-tartan-rust/product/8982802/color/725500"
    yield SplashRequest(
        url,
        callback=self.parse_product,
        args={
            "wait_for": "div[data-testid='also-viewed']",  # 等待目标元素出现
            "timeout": 30,  # 给足加载时间
        }
    )

2. 模拟页面刷新(匹配手动刷新的行为)

Since you noticed refreshing 2-3 times makes the module appear, we can replicate this in a Splash Lua script. This helps bypass one-time anti-bot checks that block the module on first load.

Create a Lua script to reload the page multiple times before extracting content:

function main(splash, args)
    -- 模拟真实浏览器的User-Agent
    splash:set_user_agent(args.user_agent)
    
    -- 第一次加载页面
    assert(splash:go(args.url))
    splash:wait(2)
    
    -- 第一次刷新
    assert(splash:reload())
    splash:wait(2)
    
    -- 第二次刷新(可根据实际情况调整刷新次数)
    assert(splash:reload())
    splash:wait(5)
    
    -- 等待目标模块渲染完成
    assert(splash:wait_for_selector('div[data-testid="also-viewed"]'))
    
    return splash:html()
end

然后在Scrapy请求中调用这个脚本:

def start_requests(self):
    url = "https://www.zappos.com/p/chaco-marshall-tartan-rust/product/8982802/color/725500"
    lua_script = """
        -- 粘贴上面的Lua脚本内容
    """
    yield SplashRequest(
        url,
        callback=self.parse_product,
        endpoint="execute",  # 使用execute端点运行自定义Lua脚本
        args={
            "lua_source": lua_script,
            "user_agent": self.settings.get("USER_AGENT"),
        }
    )

3. 模拟用户滚动(触发懒加载)

部分电商网站的“看过还看”模块需要用户滚动页面才会加载。可以在Lua脚本中加入滚动动作触发加载:

-- 在首次加载/刷新后添加滚动逻辑
splash:runjs('window.scrollTo(0, document.body.scrollHeight);')
splash:wait(3)  # 等待懒加载内容加载完成

4. 确保请求头和Cookie符合真实浏览器

Zappos可能会识别非浏览器样式的请求。确保Splash使用有效的User-Agent,必要时携带之前请求的Cookie。可以在Lua脚本中直接设置Cookie:

splash:set_cookie{
    name = "some_cookie_name",
    value = "cookie_value",
    domain = "zappos.com",
    path = "/"
}

5. 增加自定义重试逻辑

在Scrapy爬虫中添加重试机制,重新处理未加载目标模块的页面。检查响应中是否存在目标元素,若缺失则触发重试:

def parse_product(self, response):
    also_viewed = response.css('div[data-testid="also-viewed"]')
    if not also_viewed:
        retry_times = response.meta.get("retry_times", 0)
        if retry_times < 3:  # 最多重试3次
            yield response.request.replace(
                meta={"retry_times": retry_times + 1},
                dont_filter=True  # 绕过重复请求过滤
            )
            return
    
    # 处理“看过还看”模块数据
    for item in also_viewed.css(".product-item"):
        # 提取商品信息...

最后提示

先从方案1(等待目标元素)开始尝试,因为它最简单直接。如果无效,再尝试模拟刷新或滚动动作。记得用浏览器开发者工具验证选择器是否正确指向目标模块。

内容的提问来源于stack exchange,提问作者jkatt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:16:32