You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy脚本无法抓取stienda.uy商品数据问题排查

问题排查与修复

1. 核心问题:CSS选择器完全不匹配页面结构

(1)商品容器选择器错误

你脚本里用.grp778作为商品循环的容器,但实际目标页面的商品条目容器类是product-item,.grp778在当前页面DOM中根本不存在,导致循环直接跳过,没有任何数据产出。

(2)商品名称选择器错误

你用#catalogoProductos .tit::text获取名称,这个选择器是从全局的#catalogoProductos容器下找元素,不是当前商品卡片下的,即使找到也会重复获取同一个值。实际每个商品的名称在.product-item下的.product-name a标签内。

(3)价格选择器的兼容问题

页面中部分商品的价格可能使用不同类名,比如.price或.precioSantander,需要同时兼容,但前提是要在正确的商品容器下查找。

2. 修复后的脚本代码

import scrapy
import os
from scrapy.crawler import CrawlerProcess
from datetime import datetime

date = datetime.now().strftime("%d_%m_%Y")

class stiendaSpider(scrapy.Spider):
    name = 'stienda'
    start_urls = ['https://stienda.uy/tv']

    def parse(self, response):
        # 遍历每个商品容器,使用正确的类选择器
        for product in response.css('.product-item'):
            # 从当前商品容器下获取价格,兼容不同的价格类
            price = product.css('.precioSantander::text, .price::text').get()
            # 从当前商品容器下获取商品名称
            name = product.css('.product-name a::text').get()
            if price and name:
                yield {
                    'name': name.strip(),
                    'price': price.strip()
                }

# 路径添加r前缀避免转义字符问题
os.chdir(r'C:\Users\cabre\Desktop\scraping\stienda\data\raw')
process = CrawlerProcess(
     settings={"FEEDS": {f"stienda_{date}.csv": {"format": "csv"}}}
)
process.crawl(stiendaSpider)
process.start()

3. 额外调试技巧

可以用Scrapy自带的shell工具快速验证选择器:

scrapy shell https://stienda.uy/tv

在shell中输入response.css('.product-item')就能直观看到是否返回商品列表,方便快速排查选择器问题。

内容的提问来源于stack exchange,提问作者Bruno123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 10:56:06