多分类多页面电商数据爬取求助(Scrapy报错及JS页面问题)
在线商店数据爬取问题解决方案
一、修复代码中的已知错误
1. 解决User Agent提示问题
allowed_domains需为列表格式,且域名需匹配目标网站(去掉www前缀):
allowed_domains = ["e-shop.nomin.mn"]
2. 修复无效XPath表达式
- 错误XPath多了右括号,修正为:
name = card.xpath(".//a[contains(@class, 'MuiBox-root css-1efcy4n')]/text()").extract_first() - 错误XPath缺少右括号,修正为:
skp = response.xpath(".//*[contains(@class, 'MuiBox-root css-jyp6ua')]/text()").extract_first() - 价格提取需获取文本而非HTML标签:
price = card.xpath(".//*[contains(@class, 'MuiBox-root css-qr51gz')]/text()").extract_first().strip()
3. 修复面包屑分类提取逻辑
原XPath无法正确获取层级,需定位到li下的a标签文本:
list_li = response.xpath(".//*[contains(@class, 'MuiBreadcrumbs-ol css-nhb8h9')]//li/a/text()").extract() # 避免索引越界,添加长度判断 cat1 = list_li[0].strip() if len(list_li)>=1 else "" cat2 = list_li[1].strip() if len(list_li)>=2 else "" cat3 = list_li[2].strip() if len(list_li)>=3 else "" product_name = list_li[-1].strip() if list_li else ""
二、实现多分类与分页爬取
1. 遍历所有分类生成初始请求
修改start_requests,遍历预设分类字典,生成每个分类的第一页请求,并携带分类信息:
def start_requests(self): for cat_id, cat_info in categories.items(): url = f"https://e-shop.nomin.mn/t/{cat_id}?page=1" yield scrapy.Request( url=url, errback=self.parse_error, meta={ "cat_id": cat_id, "cat_name": cat_info["name"], "current_page": 1, "total_pages": cat_info["pages"] } )
2. 处理分页逻辑
在parse方法中,根据当前页码与总页数生成下一页请求:
# 处理分页 current_page = response.meta["current_page"] total_pages = response.meta["total_pages"] cat_id = response.meta["cat_id"] cat_name = response.meta["cat_name"] if current_page < total_pages: next_page_num = current_page + 1 next_page_url = f"https://e-shop.nomin.mn/t/{cat_id}?page={next_page_num}" yield scrapy.Request( url=next_page_url, callback=self.parse, errback=self.parse_error, meta={ "cat_id": cat_id, "cat_name": cat_name, "current_page": next_page_num, "total_pages": total_pages } ) self.logger.info(f"正在爬取分类 {cat_name} 的第 {next_page_num} 页")
三、动态渲染判断与处理
直接查看Scrapy响应内容:若商品卡片HTML存在,无需动态渲染工具;若响应仅含骨架HTML,说明内容由JS动态加载,推荐使用Scrapy Playwright实现渲染。
安装依赖:
pip install scrapy-playwright
在custom_settings中启用:
custom_settings = { # ...其他配置 "DOWNLOAD_HANDLERS": { "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", }, "PLAYWRIGHT_LAUNCH_OPTIONS": { "headless": True, "args": ["--no-sandbox"], }, }
并在请求中添加playwright=True参数。
四、完整修复后的代码
# -*- coding: utf-8 -*- import scrapy from datetime import datetime from scrapy.crawler import CrawlerProcess from twisted.internet.error import DNSLookupError dt_today = datetime.now().strftime('%Y%m%d') filename = dt_today + ' E-CPI Nomin' categories = { "6011": {"pages": 60, "name": "Цахилгаан бараа"}, "24175": {"pages": 70, "name": "Хүнс"}, "24273": {"pages": 40, "name": "Гэр ахуй"}, "21297": {"pages": 70, "name": "Гоо сайхан"}, "19653": {"pages": 30, "name": "Гутал, хувцас"}, "19451": {"pages": 10, "name": "Авто бараа"}, "19518": {"pages": 40, "name": "Барилгын материал"}, "19853": {"pages": 10, "name": "Аялал, Спорт бараа"}, "19487": {"pages": 50, "name": "Ном"}, "19767": {"pages": 20, "name": "Бичиг хэрэг"}, "19469": {"pages": 10, "name": "Эрүүл мэнд"}, "19545": {"pages": 20, "name": "Хүүхдийн бараа"}, } class ecpiNominSpider(scrapy.Spider): name = "cpi_nomin" allowed_domains = ["e-shop.nomin.mn"] custom_settings = { "USER_AGENT": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/110.0.0.0 Safari/537.36", "FEEDS": { f'{filename}.csv': { 'format': 'csv', 'overwrite': True } }, # 如需动态渲染,取消下方注释并安装scrapy-playwright # "DOWNLOAD_HANDLERS": { # "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", # "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", # }, # "PLAYWRIGHT_LAUNCH_OPTIONS": { # "headless": True, # "args": ["--no-sandbox"], # }, } def start_requests(self): for cat_id, cat_info in categories.items(): url = f"https://e-shop.nomin.mn/t/{cat_id}?page=1" yield scrapy.Request( url=url, errback=self.parse_error, meta={ # "playwright": True, # 如需动态渲染,取消注释 "cat_id": cat_id, "cat_name": cat_info["name"], "current_page": 1, "total_pages": cat_info["pages"] } ) def parse_error(self, failure): if failure.check(DNSLookupError): request = failure.request yield { 'URL': request.url, 'Status': str(failure.value) } def parse(self, response, **kwargs): cards = response.xpath("//*[contains(@class,'MuiBox-root css-1kmsi46')]") cat_name = response.meta["cat_name"] for card in cards: name = card.xpath(".//a[contains(@class, 'MuiBox-root css-1efcy4n')]/text()").extract_first() price = card.xpath(".//*[contains(@class, 'MuiBox-root css-qr51gz')]/text()").extract_first().strip() if card.xpath(".//*[contains(@class, 'MuiBox-root css-qr51gz')]/text()").extract_first() else "" link = card.xpath(".//a[contains(@href, '/p/')]/@href").get() if link: item = { 'main_category': cat_name, 'name': name.strip() if name else "", 'price': price, 'link': 'https://e-shop.nomin.mn' + link } yield response.follow( link, callback=self.parse_item, meta={'item': item} # meta={'item': item, "playwright": True} # 如需动态渲染详情页,取消注释 ) # 处理分页 current_page = response.meta["current_page"] total_pages = response.meta["total_pages"] cat_id = response.meta["cat_id"] if current_page < total_pages: next_page_num = current_page + 1 next_page_url = f"https://e-shop.nomin.mn/t/{cat_id}?page={next_page_num}" yield scrapy.Request( url=next_page_url, callback=self.parse, errback=self.parse_error, meta={ # "playwright": True, # 如需动态渲染分页,取消注释 "cat_id": cat_id, "cat_name": cat_name, "current_page": next_page_num, "total_pages": total_pages } ) self.logger.info(f"已发起分类 {cat_name} 第 {next_page_num} 页请求") def parse_item(self, response): item = response.meta['item'] # 提取面包屑分类 list_li = response.xpath(".//*[contains(@class, 'MuiBreadcrumbs-ol css-nhb8h9')]//li/a/text()").extract() cat1 = list_li[0].strip() if len(list_li)>=1 else "" cat2 = list_li[1].strip() if len(list_li)>=2 else "" cat3 = list_li[2].strip() if len(list_li)>=3 else "" product_name = list_li[-1].strip() if list_li else "" # 提取商品编码 skp = response.xpath(".//*[contains(@class, 'MuiBox-root css-jyp6ua')]/text()").extract_first() product_code = skp.strip() if skp else "" item.update({ 'category_level1': cat1, 'category_level2': cat2, 'category_level3': cat3, 'product_name': product_name, 'product_code': product_code }) yield item if __name__ == "__main__": process = CrawlerProcess() process.crawl(ecpiNominSpider) process.start()
五、API vs Scrapy选择建议
- 若网站更新后API结构变化且无文档,优先修复Scrapy代码:无需依赖API稳定性,直接从页面提取数据,更适配前端更新。
- 若能通过浏览器开发者工具找到新的API接口,API爬取效率更高,可优先尝试逆向新接口。
内容的提问来源于stack exchange,提问作者WX1505
相关产品推荐
相关产品推荐

