You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Scrapy项目中抓取多URL数据并合并Item存储至MongoDB单一文档

Hey there! Let's fix this up so you can store all your scraped seller data as a single MongoDB document with a unified timestamp. Here's a step-by-step solution tailored to your setup:

Step 1: Modify Your Spider to Collect & Merge Items

The key idea is to collect all individual ScrapingResultItem entries first, then generate the merged AllScrapedDataItem once all pages are scraped. We'll use a class variable to store intermediate results and leverage Scrapy's close_spider hook (triggered when all requests finish) to create the final merged item.

Here's the updated spider code:

import scrapy
from datetime import datetime
from your_project.items import ScrapingResultItem, AllScrapedDataItem

class SscSpider(scrapy.Spider):
    name = 'ssc'
    allowed_domains = ['store.com']
    # Make sure URLs include http/https (Scrapy requires full URLs)
    start_urls = ['https://www.store.com/seller1', 'https://www.store.com/seller2']

    def __init__(self, *args, **kwargs):
        super().__init__(*args, **kwargs)
        self.seller_results = []  # Container to hold all seller data

    def parse(self, response):
        # Extract seller name from the URL (e.g., "seller1" from the final path segment)
        seller_name = response.url.split('/')[-1]
        
        titles = response.css('.v2-listing-card__title::text').extract()
        prices = response.css('.currency-value::text').extract()
        
        scraped_items = []
        total_price = 0.0

        for title, price in zip(titles, prices):
            clean_title = title.strip()
            clean_price = float(price.strip())
            scraped_items.append({
                'title': clean_title,
                'price': clean_price
            })
            total_price += clean_price

        # Calculate average price (handle edge case where no items are found)
        price_avg = round(total_price / len(scraped_items), 2) if scraped_items else 0.0

        # Create and store the individual seller result
        seller_item = ScrapingResultItem(
            name=seller_name,  # Fixed: store the seller's name, not the spider's name
            scraped_items=scraped_items,
            price_avg=price_avg
        )
        self.seller_results.append(seller_item)

    def close_spider(self, spider):
        # Generate the merged item once all scraping is done
        merged_item = AllScrapedDataItem(
            datetime=datetime.utcnow().isoformat(),  # Use UTC timestamp for consistency
            data=self.seller_results
        )
        yield merged_item

Step 2: Adjust Your MongoDB Pipeline

Update your pipeline to only save the merged AllScrapedDataItem (and ignore individual ScrapingResultItem entries, since we've already collected them):

from pymongo import MongoClient
from your_project.items import AllScrapedDataItem

class MongoPipeline:
    def open_spider(self, spider):
        # Connect to your MongoDB instance
        self.client = MongoClient('mongodb://localhost:27017/')
        self.db = self.client['your_database_name']
        self.collection = self.db['your_collection_name']

    def process_item(self, item, spider):
        # Only save the merged item to MongoDB
        if isinstance(item, AllScrapedDataItem):
            self.collection.insert_one(dict(item))
            spider.logger.info(f"Successfully saved merged data with timestamp: {item['datetime']}")
        # Skip saving individual seller items (we already included them in the merged item)
        return item

    def close_spider(self, spider):
        self.client.close()

Key Notes to Keep in Mind

  • Seller Name Fix: I changed name=self.name to extract the seller name from the URL—your original code was storing the spider's name (ssc) instead of the actual seller identifier, which was probably unintended.
  • Full URLs: Always include http:// or https:// in your start_urls—Scrapy can't process relative URLs here.
  • UTC Timestamp: Using datetime.utcnow() ensures your timestamp is timezone-agnostic, which is better for database storage.
  • Scalability: If you add more seller URLs to start_urls, this code will automatically collect and merge all of them into a single document.

内容的提问来源于stack exchange,提问作者Freddie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 17:52:33