You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Scrapy实现Nomin电商全品类动态爬取的技术求助

解决方案

1. 批量获取所有品类的信息

先从网站导航栏或品类列表页抓取所有品类的URL和对应名称,把这些信息作为后续爬取的基础:

def parse_start_url(self, response):
    # 根据页面实际结构调整选择器,提取品类链接与名称
    category_items = response.css('nav a[href$=".html"]')
    for item in category_items:
        category_name = item.css('::text').get().strip()
        category_url = response.urljoin(item.attrib['href'])
        # 将品类信息传递给下一个解析函数
        yield scrapy.Request(
            url=category_url,
            callback=self.parse_category,
            meta={'category_name': category_name}
        )

2. 提取品类API请求参数

访问单个品类页面时,从页面的JS变量或开发者工具抓取的GraphQL请求中,提取categoryId这类核心参数,用来构造全量商品的API请求。

3. 处理分页,爬取品类全量商品

利用网站支持的最大每页显示数(200条)减少请求次数,通过API返回的totalItems计算总页数,循环请求所有页面:

import json

def parse_category(self, response):
    category_name = response.meta['category_name']
    # 示例:从页面meta标签提取categoryId,实际需根据页面结构调整
    category_id = response.css('meta[name="category-id"]::attr(content)').get()
    # 初始化第一页请求,pageSize设为200
    api_payload = {
        "query": """
        query getProducts($categoryId: String!, $page: Int!, $pageSize: Int!) {
            products(categoryId: $categoryId, page: $page, pageSize: $pageSize) {
                totalItems
                items {
                    id
                    name
                    price
                    # 按需添加其他商品字段
                }
            }
        }
        """,
        "variables": {
            "categoryId": category_id,
            "page": 1,
            "pageSize": 200
        }
    }
    yield scrapy.Request(
        url='https://eshop.nomin.mn/graphql',  # 以实际API地址为准
        method='POST',
        body=json.dumps(api_payload),
        headers={'Content-Type': 'application/json'},
        callback=self.parse_products,
        meta={
            'category_name': category_name,
            'category_id': category_id,
            'current_page': 1,
            'page_size': 200
        }
    )

def parse_products(self, response):
    data = json.loads(response.text)
    category_name = response.meta['category_name']
    product_list = data['data']['products']['items']
    total_count = data['data']['products']['totalItems']
    page_size = response.meta['page_size']
    total_pages = (total_count + page_size - 1) // page_size  # 向上取整计算总页数

    # 给每个商品添加品类标识,输出结果
    for product in product_list:
        product['category'] = category_name
        yield product

    # 递归请求下一页
    current_page = response.meta['current_page']
    if current_page < total_pages:
        next_page = current_page + 1
        category_id = response.meta['category_id']
        next_payload = {
            "query": """
            query getProducts($categoryId: String!, $page: Int!, $pageSize: Int!) {
                products(categoryId: $categoryId, page: $page, pageSize: $pageSize) {
                    totalItems
                    items {
                        id
                        name
                        price
                        # 对应上面的商品字段
                    }
                }
            }
            """,
            "variables": {
                "categoryId": category_id,
                "page": next_page,
                "pageSize": page_size
            }
        }
        yield scrapy.Request(
            url='https://eshop.nomin.mn/graphql',
            method='POST',
            body=json.dumps(next_payload),
            headers={'Content-Type': 'application/json'},
            callback=self.parse_products,
            meta={
                'category_name': category_name,
                'category_id': category_id,
                'current_page': next_page,
                'page_size': page_size
            }
        )

4. 无需使用Selenium/Splash的原因

网站商品数据通过GraphQL API直接返回,只要分析清楚API的请求参数、结构和认证规则,就能直接构造请求获取全量数据,比浏览器渲染工具效率更高、稳定性更强。

注意事项

  • 核对GraphQL API地址和请求头(如User-Agent、Cookie),确保与浏览器请求一致;
  • 在Scrapy配置中设置DOWNLOAD_DELAY控制请求频率,避免触发反爬;
  • 如果自动提取categoryId困难,可先通过开发者工具收集所有品类的ID和名称,手动维护一个映射表,再批量爬取。

内容的提问来源于stack exchange,提问作者WX1505

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 09:22:33