Scrapy多场景爬取技术咨询:论坛爬取项目执行逻辑问题
Hey there! Let's work through refining your Scrapy forum crawler, focusing on fixing the work_executor logic and resolving common pitfalls with the MongoDB-backed parse_topics method.
1. Fix the work_executor Flag Handling
First up, your current code references parse_only_categories without defining it—let's fix that to properly pass the flag:
Option 1: Pass the flag as a method parameter
def work_executor(self, response, parse_only_categories=False): yield from self.parse_categories(response) if not parse_only_categories: yield from self.parse_topics()
Option 2: Accept the flag via command-line arguments (more flexible)
If you want to toggle the behavior when starting the crawler, add it to your spider's __init__:
def __init__(self, parse_only_categories=False, *args, **kwargs): super().__init__(*args, **kwargs) # Convert string input from command line to boolean self.parse_only_categories = parse_only_categories.lower() == 'true' if isinstance(parse_only_categories, str) else parse_only_categories def work_executor(self, response): yield from self.parse_categories(response) if not self.parse_only_categories: yield from self.parse_topics()
Run it with: scrapy crawl forum_spider -a parse_only_categories=true
2. Implement parse_topics with MongoDB Integration
Here's a robust implementation to fetch categories from MongoDB, generate requests, and parse topics/posts. We'll handle MongoDB setup, request generation, and common edge cases:
import scrapy import pymongo class ForumSpider(scrapy.Spider): name = "forum_spider" start_urls = ["https://your-forum-url.com"] def __init__(self, *args, **kwargs): super().__init__(*args, **kwargs) # Initialize MongoDB connection (adjust credentials as needed) self.mongo_client = pymongo.MongoClient("mongodb://localhost:27017/") self.db = self.mongo_client["forum_crawler"] self.categories_collection = self.db["categories"] def parse(self, response): # Kick off the workflow yield from self.work_executor(response) def work_executor(self, response): yield from self.parse_categories(response) if not self.parse_only_categories: yield from self.parse_topics() def parse_categories(self, response): # Scrape category data and save to MongoDB for category in response.css("div.category-item"): category_data = { "name": category.css("h3.category-name::text").get().strip(), "url": response.urljoin(category.css("a.category-link::attr(href)").get()), "scraped": False # Flag to track if we've processed this category } # Avoid duplicate entries if not self.categories_collection.find_one({"url": category_data["url"]}): self.categories_collection.insert_one(category_data) yield category_data def parse_topics(self): # Fetch unprocessed categories from MongoDB unprocessed_cats = self.categories_collection.find({"scraped": False}) for cat in unprocessed_cats: # Yield request to scrape topics in this category yield scrapy.Request( url=cat["url"], callback=self.parse_topic_list, meta={"category_id": cat["_id"]} # Pass category ID to mark as processed later ) def parse_topic_list(self, response): # Parse all topic links in the category for topic in response.css("div.topic-thread"): topic_url = response.urljoin(topic.css("a.topic-title::attr(href)").get()) yield scrapy.Request( url=topic_url, callback=self.parse_posts, meta=response.meta # Pass category ID through to post parsing ) # Handle pagination (if the forum has multiple pages of topics) next_page = response.css("a.next-page::attr(href)").get() if next_page: yield scrapy.Request( url=response.urljoin(next_page), callback=self.parse_topic_list, meta=response.meta ) else: # Mark category as processed once all pages are scraped self.categories_collection.update_one( {"_id": response.meta["category_id"]}, {"$set": {"scraped": True}} ) def parse_posts(self, response): # Parse individual posts in a topic category = self.categories_collection.find_one({"_id": response.meta["category_id"]}) topic_title = response.css("h1.topic-header::text").get().strip() for post in response.css("div.post-container"): yield { "category": category["name"], "topic_title": topic_title, "author": post.css("span.post-author::text").get().strip(), "content": "\n".join(post.css("div.post-content::text").getall()).strip(), "post_date": post.css("span.post-timestamp::text").get().strip() }
3. Common Pitfalls to Fix
- MongoDB Connection Issues: Ensure your MongoDB service is running, and the connection string matches your setup (add username/password if using authenticated clusters).
- Duplicate Entries: We added a check in
parse_categoriesto avoid saving duplicate categories—adjust the unique field (likeurl) based on your forum's structure. - Blocking Operations: Pymongo is synchronous, but for most forum crawlers this won't be an issue. If you're dealing with thousands of categories, consider using an async MongoDB driver like
motorto keep the crawler fast. - Pagination Handling: Don't forget to mark categories as processed only after scraping all topic pages (we do this in
parse_topic_listwhen there's no next page).
内容的提问来源于stack exchange,提问作者chodi
相关产品推荐
相关产品推荐

