如何用单Scrapy Spider高效批量抓取Nomin商城全品类商品
全品类商品抓取Scrapy爬虫优化方案
1. 多品类批量抓取实现
原代码仅针对单个品类(ID:24175),要实现全品类抓取,先整理所有目标品类的ID与名称映射列表,在爬虫中遍历该列表,为每个品类生成对应的API请求。同时修正原代码中URL字符串拼接错误,改用json.dumps构造合法变量参数,避免语法错误。
原代码分页逻辑错误(试图从HTML页面提取分页链接,但实际调用的是GraphQL API),需通过API返回的page_info.total_pages控制分页,循环请求直到当前页超过总页数。
2. User-Agent轮换配置
为避免被反爬,可通过两种方式实现UA轮换:
- 方式一:使用第三方库
安装scrapy-user-agents,在项目settings.py中配置中间件:DOWNLOADER_MIDDLEWARES = { 'scrapy.downloadermiddlewares.useragent.UserAgentMiddleware': None, 'scrapy_user_agents.middlewares.RandomUserAgentMiddleware': 400, } - 方式二:自定义中间件
在项目中创建自定义中间件,维护UA列表并随机选择:
随后在import random from scrapy.downloadermiddlewares.useragent import UserAgentMiddleware class RandomUserAgentMiddleware(UserAgentMiddleware): def __init__(self, user_agent=''): self.user_agent = user_agent self.ua_list = [ 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/117.0.0.0 Safari/537.36', 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' # 可添加更多UA字符串 ] def process_request(self, request, spider): ua = random.choice(self.ua_list) request.headers.setdefault('User-Agent', ua)settings.py中启用该中间件。
3. 原价与现价区分逻辑
从API返回数据来看,商品价格相关字段优先级为:mp_daily_deal.deal_price(每日特价)> special_price(商品特价)> price.regularPrice.amount.value(原价)。解析时按此优先级取值,同时记录价格类型:
def get_price(item): original_price = item['price']['regularPrice']['amount']['value'] # 优先取每日特价 if item.get('mp_daily_deal') and item['mp_daily_deal'].get('deal_price'): return { 'price_type': 'daily_deal', 'current_price': item['mp_daily_deal']['deal_price'], 'original_price': original_price } # 其次取商品特价 elif item.get('special_price') is not None: return { 'price_type': 'special', 'current_price': item['special_price'], 'original_price': original_price } # 默认取原价 else: return { 'price_type': 'regular', 'current_price': original_price, 'original_price': original_price }
4. 完整重构代码
import scrapy from scrapy import Request from datetime import datetime import json BASE_URL = "https://eshop.nomin.mn/graphql" QUERY = """ query category($pageSize:Int!$currentPage:Int!$filters:ProductAttributeFilterInput!$sort:ProductAttributeSortInput){ products(pageSize:$pageSize currentPage:$currentPage filter:$filters sort:$sort){ items{ id name sku brand salable_qty brand_name mp_daily_deal{deal_price} special_price price{regularPrice{amount{value}}} short_description{html} } page_info{total_pages} } } """ dt_today = datetime.now().strftime('%Y%m%d') filename = f"{dt_today}_Nomin_All_Categories_Data.csv" # 替换为你需要抓取的所有品类ID与名称 CATEGORIES = [ {"id": "24175", "name": "食品"}, {"id": "12345", "name": "家居用品"}, {"id": "67890", "name": "电子产品"}, # 更多品类... ] class NominAllCategoriesSpider(scrapy.Spider): name = 'nomin_all_categories' allowed_domains = ['eshop.nomin.mn'] custom_settings = { "FEEDS": { filename: {'format': 'csv', 'overwrite': True} }, # 若使用自定义UA中间件,取消以下注释并修改项目名称 # "DOWNLOADER_MIDDLEWARES": { # 'scrapy.downloadermiddlewares.useragent.UserAgentMiddleware': None, # 'your_project_name.middlewares.RandomUserAgentMiddleware': 400, # } } def start_requests(self): for category in CATEGORIES: # 从第1页开始请求 yield self.make_category_request(category, current_page=1) def make_category_request(self, category, current_page): variables = { "currentPage": current_page, "filters": {"category_id": {"in": [category["id"]]}}, "pageSize": 50, "sort": {"position": "DESC"} } payload = { "query": QUERY.strip(), "operationName": "category", "variables": variables } return Request( url=BASE_URL, method='POST', body=json.dumps(payload), headers={'Content-Type': 'application/json'}, meta={ 'category': category, 'current_page': current_page }, callback=self.parse ) def parse(self, response): data = response.json() category = response.meta['category'] current_page = response.meta['current_page'] products = data['data']['products']['items'] total_pages = data['data']['products']['page_info']['total_pages'] for item in products: price_data = self.get_price(item) yield { "category_id": category["id"], "category_name": category["name"], "product_id": item["id"], "name": item["name"], "sku": item["sku"], "price_type": price_data['price_type'], "original_price": price_data['original_price'], "current_price": price_data['current_price'], "description": item["short_description"]["html"] if item.get("short_description") else "" } # 处理分页:如果当前页小于总页数,请求下一页 if current_page < total_pages: yield self.make_category_request(category, current_page + 1) def get_price(self, item): original_price = item['price']['regularPrice']['amount']['value'] # 优先取每日特价 if item.get('mp_daily_deal') and item['mp_daily_deal'].get('deal_price'): return { 'price_type': 'daily_deal', 'current_price': item['mp_daily_deal']['deal_price'], 'original_price': original_price } # 其次取商品特价 elif item.get('special_price') is not None: return { 'price_type': 'special', 'current_price': item['special_price'], 'original_price': original_price } # 默认取原价 else: return { 'price_type': 'regular', 'current_price': original_price, 'original_price': original_price } if __name__ == "__main__": from scrapy.crawler import CrawlerProcess process = CrawlerProcess() process.crawl(NominAllCategoriesSpider) process.start()
关键改进点说明
- 将GraphQL查询与URL分离,提高代码可读性
- 改用POST请求传递参数(更符合GraphQL API规范,避免GET参数过长问题)
- 新增品类列表遍历逻辑,支持批量抓取
- 修正分页逻辑,基于API返回的总页数自动处理多页
- 完善价格解析逻辑,明确区分原价、特价、每日特价
- 修复原代码中
allowed_domains的错误(应为域名而非完整URL)
内容的提问来源于stack exchange,提问作者WX1505
相关产品推荐
相关产品推荐

