You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy框架爬取电商网站产品信息遇脚本问题求助

Scrapy脚本问题修复方案

原有代码核心错误

  • 提取商品链接时使用get()方法仅能获取第一条匹配的链接,后续循环实际是遍历单个字符串的每个字符,逻辑完全错误,需替换为getall()获取全部商品链接。
  • 拼接得到完整商品链接后,错误调用字符串对象的css()方法,字符串无该属性会直接触发运行报错,直接返回拼接完成的product_link即可。
  • 评论区选择器.Natsob与目标站点实际结构不匹配,无法提取到评论内容;同时标题、价格提取结果存在多余空白字符,需要做清洗处理。

修正后的完整代码

import scrapy    

class DressSpider(scrapy.Spider):
    name = 'dress'
    allowed_domains = ['savedbythedress.com']
    start_urls = ['https://savedbythedress.com/collections/maternity-tops']

    def parse(self, response):
        domain = "https://savedbythedress.com"
        # 改为getall()获取所有商品链接
        link_products = response.css('div.product-info-inner a::attr(href)').getall()
        for link in link_products:
            product_link = domain + link   
            # 直接返回拼接好的链接,不要调用css方法
            yield{
                'product_link': product_link,
            }      
            yield scrapy.Request(url=product_link, callback=self.parse_contents)

    def parse_contents(self, response):
        yield{
            'product_title' : response.css('.sbtd-product-title ::text').get().strip(),
            'product_price' : response.css('.product-price ::text').get().strip(),
            # 修正评论选择器,适配目标站点的评论结构
            'product_review' : [rev.strip() for rev in response.css('.jdgm-rev__body ::text').getall() if rev.strip()]
        }

运行注意事项

  • 若运行时出现403拦截,可在Scrapy项目的settings.py中添加合法的请求头标识,示例配置:USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
  • 若站点开启了动态渲染,可搭配scrapy-splash或者playwright组件加载动态内容,保证评论数据完整提取。

内容的提问来源于stack exchange,提问作者Thinesh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 15:54:04