如何在Scrapy项目中抓取多URL数据并合并Item存储至MongoDB单一文档
Hey there! Let's fix this up so you can store all your scraped seller data as a single MongoDB document with a unified timestamp. Here's a step-by-step solution tailored to your setup:
Step 1: Modify Your Spider to Collect & Merge Items
The key idea is to collect all individual ScrapingResultItem entries first, then generate the merged AllScrapedDataItem once all pages are scraped. We'll use a class variable to store intermediate results and leverage Scrapy's close_spider hook (triggered when all requests finish) to create the final merged item.
Here's the updated spider code:
import scrapy from datetime import datetime from your_project.items import ScrapingResultItem, AllScrapedDataItem class SscSpider(scrapy.Spider): name = 'ssc' allowed_domains = ['store.com'] # Make sure URLs include http/https (Scrapy requires full URLs) start_urls = ['https://www.store.com/seller1', 'https://www.store.com/seller2'] def __init__(self, *args, **kwargs): super().__init__(*args, **kwargs) self.seller_results = [] # Container to hold all seller data def parse(self, response): # Extract seller name from the URL (e.g., "seller1" from the final path segment) seller_name = response.url.split('/')[-1] titles = response.css('.v2-listing-card__title::text').extract() prices = response.css('.currency-value::text').extract() scraped_items = [] total_price = 0.0 for title, price in zip(titles, prices): clean_title = title.strip() clean_price = float(price.strip()) scraped_items.append({ 'title': clean_title, 'price': clean_price }) total_price += clean_price # Calculate average price (handle edge case where no items are found) price_avg = round(total_price / len(scraped_items), 2) if scraped_items else 0.0 # Create and store the individual seller result seller_item = ScrapingResultItem( name=seller_name, # Fixed: store the seller's name, not the spider's name scraped_items=scraped_items, price_avg=price_avg ) self.seller_results.append(seller_item) def close_spider(self, spider): # Generate the merged item once all scraping is done merged_item = AllScrapedDataItem( datetime=datetime.utcnow().isoformat(), # Use UTC timestamp for consistency data=self.seller_results ) yield merged_item
Step 2: Adjust Your MongoDB Pipeline
Update your pipeline to only save the merged AllScrapedDataItem (and ignore individual ScrapingResultItem entries, since we've already collected them):
from pymongo import MongoClient from your_project.items import AllScrapedDataItem class MongoPipeline: def open_spider(self, spider): # Connect to your MongoDB instance self.client = MongoClient('mongodb://localhost:27017/') self.db = self.client['your_database_name'] self.collection = self.db['your_collection_name'] def process_item(self, item, spider): # Only save the merged item to MongoDB if isinstance(item, AllScrapedDataItem): self.collection.insert_one(dict(item)) spider.logger.info(f"Successfully saved merged data with timestamp: {item['datetime']}") # Skip saving individual seller items (we already included them in the merged item) return item def close_spider(self, spider): self.client.close()
Key Notes to Keep in Mind
- Seller Name Fix: I changed
name=self.nameto extract the seller name from the URL—your original code was storing the spider's name (ssc) instead of the actual seller identifier, which was probably unintended. - Full URLs: Always include
http://orhttps://in yourstart_urls—Scrapy can't process relative URLs here. - UTC Timestamp: Using
datetime.utcnow()ensures your timestamp is timezone-agnostic, which is better for database storage. - Scalability: If you add more seller URLs to
start_urls, this code will automatically collect and merge all of them into a single document.
内容的提问来源于stack exchange,提问作者Freddie

