You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy多爬虫数据存储至MongoDB不同集合问题求助

Solution: Route Scrapy Spiders to Different MongoDB Collections

To keep your whole.py spider storing data in the courses collection while directing admissionReq.py to its own dedicated collection, you can adjust your Scrapy pipeline to dynamically choose the target collection based on the spider. Here are two practical approaches:

This method is flexible and scalable—you can easily add more spiders with unique collections later without modifying the pipeline code.

Step 1: Add Collection Attribute to Your Spiders

  • In whole.py, add a class attribute to confirm its target collection:
    class WholeSpider(scrapy.Spider):
        name = 'whole'
        collection_name = 'courses'  # Matches your existing collection
        # ... rest of your spider code
    
  • In admissionReq.py, add a similar attribute for its dedicated collection (pick a name like admission_requirements that fits your data):
    class AdmissionReqSpider(scrapy.Spider):
        name = 'admissionReq'
        collection_name = 'admission_requirements'  # Your new dedicated collection
        # ... rest of your spider code
    

Step 2: Update the MongoDB Pipeline

Modify your pipelines.py to use the spider's collection_name attribute when saving items. If a spider doesn't define this attribute, it will fall back to your default courses collection:

class MongoDBPipeline:
    def __init__(self, mongo_uri, mongo_db, default_collection):
        self.mongo_uri = mongo_uri
        self.mongo_db = mongo_db
        self.default_collection = default_collection

    @classmethod
    def from_crawler(cls, crawler):
        return cls(
            mongo_uri=crawler.settings.get('MONGO_URI'),
            mongo_db=crawler.settings.get('MONGO_DATABASE', 'items'),
            default_collection=crawler.settings.get('MONGO_COLLECTION', 'courses')
        )

    def open_spider(self, spider):
        self.client = pymongo.MongoClient(self.mongo_uri)
        self.db = self.client[self.mongo_db]

    def close_spider(self, spider):
        self.client.close()

    def process_item(self, item, spider):
        # Use the spider's specified collection, or default if not set
        target_collection = getattr(spider, 'collection_name', self.default_collection)
        self.db[target_collection].insert_one(dict(item))
        return item

Approach 2: Route by Spider Name (Simpler for Two Spiders)

If you only have these two spiders and don't plan to add more, you can directly check the spider's name in the pipeline:

def process_item(self, item, spider):
    if spider.name == 'admissionReq':
        collection_name = 'admission_requirements'
    else:
        collection_name = 'courses'  # Default for whole.py and any other spiders
    self.db[collection_name].insert_one(dict(item))
    return item

Final Check

  • Ensure your settings.py has the MongoDB pipeline enabled in ITEM_PIPELINES:
    ITEM_PIPELINES = {
        'your_project_name.pipelines.MongoDBPipeline': 300,
    }
    
  • Run your spiders and verify your MongoDB database—you should see data in both the courses collection and your new admission_requirements collection.

内容的提问来源于stack exchange,提问作者x_P0seid0N_x

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:30:10