You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多分类多页面电商数据爬取求助(Scrapy报错及JS页面问题)

在线商店数据爬取问题解决方案

一、修复代码中的已知错误

1. 解决User Agent提示问题

allowed_domains需为列表格式,且域名需匹配目标网站(去掉www前缀):

allowed_domains = ["e-shop.nomin.mn"]

2. 修复无效XPath表达式

  • 错误XPath多了右括号,修正为:
    name = card.xpath(".//a[contains(@class, 'MuiBox-root css-1efcy4n')]/text()").extract_first()
    
  • 错误XPath缺少右括号,修正为:
    skp = response.xpath(".//*[contains(@class, 'MuiBox-root css-jyp6ua')]/text()").extract_first()
    
  • 价格提取需获取文本而非HTML标签:
    price = card.xpath(".//*[contains(@class, 'MuiBox-root css-qr51gz')]/text()").extract_first().strip()
    

3. 修复面包屑分类提取逻辑

原XPath无法正确获取层级,需定位到li下的a标签文本:

list_li = response.xpath(".//*[contains(@class, 'MuiBreadcrumbs-ol css-nhb8h9')]//li/a/text()").extract()
# 避免索引越界,添加长度判断
cat1 = list_li[0].strip() if len(list_li)>=1 else ""
cat2 = list_li[1].strip() if len(list_li)>=2 else ""
cat3 = list_li[2].strip() if len(list_li)>=3 else ""
product_name = list_li[-1].strip() if list_li else ""

二、实现多分类与分页爬取

1. 遍历所有分类生成初始请求

修改start_requests,遍历预设分类字典,生成每个分类的第一页请求,并携带分类信息:

def start_requests(self):
    for cat_id, cat_info in categories.items():
        url = f"https://e-shop.nomin.mn/t/{cat_id}?page=1"
        yield scrapy.Request(
            url=url,
            errback=self.parse_error,
            meta={
                "cat_id": cat_id,
                "cat_name": cat_info["name"],
                "current_page": 1,
                "total_pages": cat_info["pages"]
            }
        )

2. 处理分页逻辑

在parse方法中,根据当前页码与总页数生成下一页请求:

# 处理分页
current_page = response.meta["current_page"]
total_pages = response.meta["total_pages"]
cat_id = response.meta["cat_id"]
cat_name = response.meta["cat_name"]

if current_page < total_pages:
    next_page_num = current_page + 1
    next_page_url = f"https://e-shop.nomin.mn/t/{cat_id}?page={next_page_num}"
    yield scrapy.Request(
        url=next_page_url,
        callback=self.parse,
        errback=self.parse_error,
        meta={
            "cat_id": cat_id,
            "cat_name": cat_name,
            "current_page": next_page_num,
            "total_pages": total_pages
        }
    )
    self.logger.info(f"正在爬取分类 {cat_name} 的第 {next_page_num} 页")

三、动态渲染判断与处理

直接查看Scrapy响应内容:若商品卡片HTML存在,无需动态渲染工具;若响应仅含骨架HTML,说明内容由JS动态加载,推荐使用Scrapy Playwright实现渲染。

安装依赖:

pip install scrapy-playwright

在custom_settings中启用:

custom_settings = {
    # ...其他配置
    "DOWNLOAD_HANDLERS": {
        "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
        "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
    },
    "PLAYWRIGHT_LAUNCH_OPTIONS": {
        "headless": True,
        "args": ["--no-sandbox"],
    },
}

并在请求中添加playwright=True参数。

四、完整修复后的代码

# -*- coding: utf-8 -*-
import scrapy
from datetime import datetime
from scrapy.crawler import CrawlerProcess
from twisted.internet.error import DNSLookupError

dt_today = datetime.now().strftime('%Y%m%d')
filename = dt_today + ' E-CPI Nomin'

categories = {
    "6011": {"pages": 60, "name": "Цахилгаан бараа"},
    "24175": {"pages": 70, "name": "Хүнс"},
    "24273": {"pages": 40, "name": "Гэр ахуй"},
    "21297": {"pages": 70, "name": "Гоо сайхан"},
    "19653": {"pages": 30, "name": "Гутал, хувцас"},
    "19451": {"pages": 10, "name": "Авто бараа"},
    "19518": {"pages": 40, "name": "Барилгын материал"},
    "19853": {"pages": 10, "name": "Аялал, Спорт бараа"},
    "19487": {"pages": 50, "name": "Ном"},
    "19767": {"pages": 20, "name": "Бичиг хэрэг"},
    "19469": {"pages": 10, "name": "Эрүүл мэнд"},
    "19545": {"pages": 20, "name": "Хүүхдийн бараа"},
}

class ecpiNominSpider(scrapy.Spider):
    name = "cpi_nomin"
    allowed_domains = ["e-shop.nomin.mn"]
    custom_settings = {
        "USER_AGENT": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/110.0.0.0 Safari/537.36",
        "FEEDS": {
            f'{filename}.csv': {
                'format': 'csv',
                'overwrite': True
            }
        },
        # 如需动态渲染,取消下方注释并安装scrapy-playwright
        # "DOWNLOAD_HANDLERS": {
        #     "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
        #     "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
        # },
        # "PLAYWRIGHT_LAUNCH_OPTIONS": {
        #     "headless": True,
        #     "args": ["--no-sandbox"],
        # },
    }

    def start_requests(self):
        for cat_id, cat_info in categories.items():
            url = f"https://e-shop.nomin.mn/t/{cat_id}?page=1"
            yield scrapy.Request(
                url=url,
                errback=self.parse_error,
                meta={
                    # "playwright": True,  # 如需动态渲染,取消注释
                    "cat_id": cat_id,
                    "cat_name": cat_info["name"],
                    "current_page": 1,
                    "total_pages": cat_info["pages"]
                }
            )

    def parse_error(self, failure):
        if failure.check(DNSLookupError):
            request = failure.request
            yield {
                'URL': request.url,
                'Status': str(failure.value)
            }

    def parse(self, response, **kwargs):
        cards = response.xpath("//*[contains(@class,'MuiBox-root css-1kmsi46')]")
        cat_name = response.meta["cat_name"]

        for card in cards:
            name = card.xpath(".//a[contains(@class, 'MuiBox-root css-1efcy4n')]/text()").extract_first()
            price = card.xpath(".//*[contains(@class, 'MuiBox-root css-qr51gz')]/text()").extract_first().strip() if card.xpath(".//*[contains(@class, 'MuiBox-root css-qr51gz')]/text()").extract_first() else ""
            link = card.xpath(".//a[contains(@href, '/p/')]/@href").get()

            if link:
                item = {
                    'main_category': cat_name,
                    'name': name.strip() if name else "",
                    'price': price,
                    'link': 'https://e-shop.nomin.mn' + link
                }
                yield response.follow(
                    link,
                    callback=self.parse_item,
                    meta={'item': item}
                    # meta={'item': item, "playwright": True}  # 如需动态渲染详情页,取消注释
                )

        # 处理分页
        current_page = response.meta["current_page"]
        total_pages = response.meta["total_pages"]
        cat_id = response.meta["cat_id"]

        if current_page < total_pages:
            next_page_num = current_page + 1
            next_page_url = f"https://e-shop.nomin.mn/t/{cat_id}?page={next_page_num}"
            yield scrapy.Request(
                url=next_page_url,
                callback=self.parse,
                errback=self.parse_error,
                meta={
                    # "playwright": True,  # 如需动态渲染分页,取消注释
                    "cat_id": cat_id,
                    "cat_name": cat_name,
                    "current_page": next_page_num,
                    "total_pages": total_pages
                }
            )
            self.logger.info(f"已发起分类 {cat_name} 第 {next_page_num} 页请求")

    def parse_item(self, response):
        item = response.meta['item']

        # 提取面包屑分类
        list_li = response.xpath(".//*[contains(@class, 'MuiBreadcrumbs-ol css-nhb8h9')]//li/a/text()").extract()
        cat1 = list_li[0].strip() if len(list_li)>=1 else ""
        cat2 = list_li[1].strip() if len(list_li)>=2 else ""
        cat3 = list_li[2].strip() if len(list_li)>=3 else ""
        product_name = list_li[-1].strip() if list_li else ""

        # 提取商品编码
        skp = response.xpath(".//*[contains(@class, 'MuiBox-root css-jyp6ua')]/text()").extract_first()
        product_code = skp.strip() if skp else ""

        item.update({
            'category_level1': cat1,
            'category_level2': cat2,
            'category_level3': cat3,
            'product_name': product_name,
            'product_code': product_code
        })

        yield item

if __name__ == "__main__":
    process = CrawlerProcess()
    process.crawl(ecpiNominSpider)
    process.start()

五、API vs Scrapy选择建议

  • 若网站更新后API结构变化且无文档,优先修复Scrapy代码:无需依赖API稳定性,直接从页面提取数据,更适配前端更新。
  • 若能通过浏览器开发者工具找到新的API接口,API爬取效率更高,可优先尝试逆向新接口。

内容的提问来源于stack exchange,提问作者WX1505

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 06:33:08