You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy抓取页面href元素无响应,求排查问题原因

问题分析与解决

你遇到的问题是Scrapy爬虫未抓取到目标链接、无任何输出,核心原因集中在页面动态渲染和选择器无效两点,以下是具体排查和解决方法:

1. 核心问题:静态HTML无目标元素

StockX的商品列表是通过JavaScript动态渲染的,Scrapy默认请求获取的是服务器返回的初始静态HTML,里面并没有你指定的css-pnc6ci类对应的元素,导致LinkExtractor抓不到任何链接。

验证方法

打开Scrapy Shell测试:

scrapy shell https://www.stockx.com/sneakers

在Shell中执行XPath查询:

response.xpath("//div[@class='css-pnc6ci']/a")

如果返回空列表[],就证明初始HTML里没有目标元素,必须处理JS渲染。

2. 解决方法

方法一:用Scrapy Splash处理动态页面

Splash是Scrapy官方推荐的JS渲染工具,配置步骤:

  1. 安装依赖并启动服务:
pip install scrapy-splash
docker run -p 8050:8050 scrapinghub/splash
  1. 在项目settings.py中添加配置:
SPLASH_URL = 'http://localhost:8050'
DOWNLOADER_MIDDLEWARES = {
    'scrapy_splash.SplashCookiesMiddleware': 723,
    'scrapy_splash.SplashMiddleware': 725,
    'scrapy.downloadermiddlewares.httpcompression.HttpCompressionMiddleware': 810,
}
SPIDER_MIDDLEWARES = {
    'scrapy_splash.SplashDeduplicateArgsMiddleware': 100,
}
DUPEFILTER_CLASS = 'scrapy_splash.SplashAwareDupeFilter'
  1. 修改爬虫代码,使用SplashRequest等待渲染:
import scrapy
from scrapy.linkextractors import LinkExtractor
from scrapy.spiders import CrawlSpider, Rule
from scrapy_splash import SplashRequest

class Shoes2Spider(CrawlSpider):
    name = "shoes2"
    allowed_domains = ["stockx.com"]
    start_urls = ["https://www.stockx.com/sneakers"]

    rules = (
        Rule(
            LinkExtractor(restrict_xpaths="//div[@class='css-pnc6ci']/a"), 
            callback="parse_item", 
            follow=True
        ),
    )

    def start_requests(self):
        for url in self.start_urls:
            yield SplashRequest(url, args={'wait': 2})  # 等待2秒让JS渲染完成

    def parse_item(self, response):
        print(response.url)

方法二:直接抓取API接口(更高效)

StockX的商品数据来自后台API,可在浏览器开发者工具「网络」标签中筛选XHR请求,找到返回商品列表的API接口,直接爬取JSON数据:

import scrapy
import json

class Shoes2Spider(scrapy.Spider):
    name = "shoes2"
    allowed_domains = ["stockx.com"]
    start_urls = ["https://stockx.com/api/browse?_search=&category=sneakers&limit=40"]  # limit控制返回数量

    def parse(self, response):
        data = json.loads(response.text)
        for product in data['Products']:
            product_url = f"https://www.stockx.com{product['url']}"
            print(product_url)
            # 需跟进详情页可继续yield Request

方法三:添加User-Agent绕过基础反爬

StockX会验证请求的User-Agent,在settings.py中添加浏览器UA:

USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'

3. 其他排查点

  • 确认XPath选择器:网站可能随时更新元素class名称,可通过浏览器「检查」功能重新获取最新的XPath。
  • 爬虫规则的follow=True:只有初始页面抓到链接后,follow才会生效,优先确保初始页面的链接抓取逻辑正常。

内容的提问来源于stack exchange,提问作者The Rookie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 19:08:20