Scrapy多爬虫数据存储至MongoDB不同集合问题求助
To keep your whole.py spider storing data in the courses collection while directing admissionReq.py to its own dedicated collection, you can adjust your Scrapy pipeline to dynamically choose the target collection based on the spider. Here are two practical approaches:
Approach 1: Define Collection Name in Each Spider (Recommended)
This method is flexible and scalable—you can easily add more spiders with unique collections later without modifying the pipeline code.
Step 1: Add Collection Attribute to Your Spiders
- In
whole.py, add a class attribute to confirm its target collection:class WholeSpider(scrapy.Spider): name = 'whole' collection_name = 'courses' # Matches your existing collection # ... rest of your spider code - In
admissionReq.py, add a similar attribute for its dedicated collection (pick a name likeadmission_requirementsthat fits your data):class AdmissionReqSpider(scrapy.Spider): name = 'admissionReq' collection_name = 'admission_requirements' # Your new dedicated collection # ... rest of your spider code
Step 2: Update the MongoDB Pipeline
Modify your pipelines.py to use the spider's collection_name attribute when saving items. If a spider doesn't define this attribute, it will fall back to your default courses collection:
class MongoDBPipeline: def __init__(self, mongo_uri, mongo_db, default_collection): self.mongo_uri = mongo_uri self.mongo_db = mongo_db self.default_collection = default_collection @classmethod def from_crawler(cls, crawler): return cls( mongo_uri=crawler.settings.get('MONGO_URI'), mongo_db=crawler.settings.get('MONGO_DATABASE', 'items'), default_collection=crawler.settings.get('MONGO_COLLECTION', 'courses') ) def open_spider(self, spider): self.client = pymongo.MongoClient(self.mongo_uri) self.db = self.client[self.mongo_db] def close_spider(self, spider): self.client.close() def process_item(self, item, spider): # Use the spider's specified collection, or default if not set target_collection = getattr(spider, 'collection_name', self.default_collection) self.db[target_collection].insert_one(dict(item)) return item
Approach 2: Route by Spider Name (Simpler for Two Spiders)
If you only have these two spiders and don't plan to add more, you can directly check the spider's name in the pipeline:
def process_item(self, item, spider): if spider.name == 'admissionReq': collection_name = 'admission_requirements' else: collection_name = 'courses' # Default for whole.py and any other spiders self.db[collection_name].insert_one(dict(item)) return item
Final Check
- Ensure your
settings.pyhas the MongoDB pipeline enabled inITEM_PIPELINES:ITEM_PIPELINES = { 'your_project_name.pipelines.MongoDBPipeline': 300, } - Run your spiders and verify your MongoDB database—you should see data in both the
coursescollection and your newadmission_requirementscollection.
内容的提问来源于stack exchange,提问作者x_P0seid0N_x

