You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用单Scrapy Spider高效批量抓取Nomin商城全品类商品

全品类商品抓取Scrapy爬虫优化方案

1. 多品类批量抓取实现

原代码仅针对单个品类(ID:24175),要实现全品类抓取,先整理所有目标品类的ID与名称映射列表,在爬虫中遍历该列表,为每个品类生成对应的API请求。同时修正原代码中URL字符串拼接错误,改用json.dumps构造合法变量参数,避免语法错误。

原代码分页逻辑错误(试图从HTML页面提取分页链接,但实际调用的是GraphQL API),需通过API返回的page_info.total_pages控制分页,循环请求直到当前页超过总页数。

2. User-Agent轮换配置

为避免被反爬,可通过两种方式实现UA轮换:

  • 方式一:使用第三方库
    安装scrapy-user-agents,在项目settings.py中配置中间件:
    DOWNLOADER_MIDDLEWARES = {
        'scrapy.downloadermiddlewares.useragent.UserAgentMiddleware': None,
        'scrapy_user_agents.middlewares.RandomUserAgentMiddleware': 400,
    }
    
  • 方式二:自定义中间件
    在项目中创建自定义中间件,维护UA列表并随机选择:
    import random
    from scrapy.downloadermiddlewares.useragent import UserAgentMiddleware
    
    class RandomUserAgentMiddleware(UserAgentMiddleware):
        def __init__(self, user_agent=''):
            self.user_agent = user_agent
            self.ua_list = [
                'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
                'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/117.0.0.0 Safari/537.36',
                'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
                # 可添加更多UA字符串
            ]
    
        def process_request(self, request, spider):
            ua = random.choice(self.ua_list)
            request.headers.setdefault('User-Agent', ua)
    
    随后在settings.py中启用该中间件。

3. 原价与现价区分逻辑

从API返回数据来看,商品价格相关字段优先级为:mp_daily_deal.deal_price(每日特价)> special_price(商品特价)> price.regularPrice.amount.value(原价)。解析时按此优先级取值,同时记录价格类型:

def get_price(item):
    original_price = item['price']['regularPrice']['amount']['value']
    # 优先取每日特价
    if item.get('mp_daily_deal') and item['mp_daily_deal'].get('deal_price'):
        return {
            'price_type': 'daily_deal',
            'current_price': item['mp_daily_deal']['deal_price'],
            'original_price': original_price
        }
    # 其次取商品特价
    elif item.get('special_price') is not None:
        return {
            'price_type': 'special',
            'current_price': item['special_price'],
            'original_price': original_price
        }
    # 默认取原价
    else:
        return {
            'price_type': 'regular',
            'current_price': original_price,
            'original_price': original_price
        }

4. 完整重构代码

import scrapy
from scrapy import Request
from datetime import datetime
import json

BASE_URL = "https://eshop.nomin.mn/graphql"
QUERY = """
query category($pageSize:Int!$currentPage:Int!$filters:ProductAttributeFilterInput!$sort:ProductAttributeSortInput){
    products(pageSize:$pageSize currentPage:$currentPage filter:$filters sort:$sort){
        items{
            id name sku brand salable_qty brand_name
            mp_daily_deal{deal_price}
            special_price
            price{regularPrice{amount{value}}}
            short_description{html}
        }
        page_info{total_pages}
    }
}
"""

dt_today = datetime.now().strftime('%Y%m%d')
filename = f"{dt_today}_Nomin_All_Categories_Data.csv"

# 替换为你需要抓取的所有品类ID与名称
CATEGORIES = [
    {"id": "24175", "name": "食品"},
    {"id": "12345", "name": "家居用品"},
    {"id": "67890", "name": "电子产品"},
    # 更多品类...
]

class NominAllCategoriesSpider(scrapy.Spider):
    name = 'nomin_all_categories'
    allowed_domains = ['eshop.nomin.mn']
    custom_settings = {
        "FEEDS": {
            filename: {'format': 'csv', 'overwrite': True}
        },
        # 若使用自定义UA中间件,取消以下注释并修改项目名称
        # "DOWNLOADER_MIDDLEWARES": {
        #     'scrapy.downloadermiddlewares.useragent.UserAgentMiddleware': None,
        #     'your_project_name.middlewares.RandomUserAgentMiddleware': 400,
        # }
    }

    def start_requests(self):
        for category in CATEGORIES:
            # 从第1页开始请求
            yield self.make_category_request(category, current_page=1)

    def make_category_request(self, category, current_page):
        variables = {
            "currentPage": current_page,
            "filters": {"category_id": {"in": [category["id"]]}},
            "pageSize": 50,
            "sort": {"position": "DESC"}
        }
        payload = {
            "query": QUERY.strip(),
            "operationName": "category",
            "variables": variables
        }
        return Request(
            url=BASE_URL,
            method='POST',
            body=json.dumps(payload),
            headers={'Content-Type': 'application/json'},
            meta={
                'category': category,
                'current_page': current_page
            },
            callback=self.parse
        )

    def parse(self, response):
        data = response.json()
        category = response.meta['category']
        current_page = response.meta['current_page']
        products = data['data']['products']['items']
        total_pages = data['data']['products']['page_info']['total_pages']

        for item in products:
            price_data = self.get_price(item)
            yield {
                "category_id": category["id"],
                "category_name": category["name"],
                "product_id": item["id"],
                "name": item["name"],
                "sku": item["sku"],
                "price_type": price_data['price_type'],
                "original_price": price_data['original_price'],
                "current_price": price_data['current_price'],
                "description": item["short_description"]["html"] if item.get("short_description") else ""
            }

        # 处理分页:如果当前页小于总页数,请求下一页
        if current_page < total_pages:
            yield self.make_category_request(category, current_page + 1)

    def get_price(self, item):
        original_price = item['price']['regularPrice']['amount']['value']
        # 优先取每日特价
        if item.get('mp_daily_deal') and item['mp_daily_deal'].get('deal_price'):
            return {
                'price_type': 'daily_deal',
                'current_price': item['mp_daily_deal']['deal_price'],
                'original_price': original_price
            }
        # 其次取商品特价
        elif item.get('special_price') is not None:
            return {
                'price_type': 'special',
                'current_price': item['special_price'],
                'original_price': original_price
            }
        # 默认取原价
        else:
            return {
                'price_type': 'regular',
                'current_price': original_price,
                'original_price': original_price
            }

if __name__ == "__main__":
    from scrapy.crawler import CrawlerProcess
    process = CrawlerProcess()
    process.crawl(NominAllCategoriesSpider)
    process.start()

关键改进点说明

  • 将GraphQL查询与URL分离,提高代码可读性
  • 改用POST请求传递参数(更符合GraphQL API规范,避免GET参数过长问题)
  • 新增品类列表遍历逻辑,支持批量抓取
  • 修正分页逻辑,基于API返回的总页数自动处理多页
  • 完善价格解析逻辑,明确区分原价、特价、每日特价
  • 修复原代码中allowed_domains的错误(应为域名而非完整URL)

内容的提问来源于stack exchange,提问作者WX1505

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 03:54:55