You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy多场景爬取技术咨询:论坛爬取项目执行逻辑问题

Hey there! Let's work through refining your Scrapy forum crawler, focusing on fixing the work_executor logic and resolving common pitfalls with the MongoDB-backed parse_topics method.

1. Fix the work_executor Flag Handling

First up, your current code references parse_only_categories without defining it—let's fix that to properly pass the flag:

Option 1: Pass the flag as a method parameter

def work_executor(self, response, parse_only_categories=False):
    yield from self.parse_categories(response)
    if not parse_only_categories:
        yield from self.parse_topics()

Option 2: Accept the flag via command-line arguments (more flexible)

If you want to toggle the behavior when starting the crawler, add it to your spider's __init__:

def __init__(self, parse_only_categories=False, *args, **kwargs):
    super().__init__(*args, **kwargs)
    # Convert string input from command line to boolean
    self.parse_only_categories = parse_only_categories.lower() == 'true' if isinstance(parse_only_categories, str) else parse_only_categories

def work_executor(self, response):
    yield from self.parse_categories(response)
    if not self.parse_only_categories:
        yield from self.parse_topics()

Run it with: scrapy crawl forum_spider -a parse_only_categories=true


2. Implement parse_topics with MongoDB Integration

Here's a robust implementation to fetch categories from MongoDB, generate requests, and parse topics/posts. We'll handle MongoDB setup, request generation, and common edge cases:

import scrapy
import pymongo

class ForumSpider(scrapy.Spider):
    name = "forum_spider"
    start_urls = ["https://your-forum-url.com"]

    def __init__(self, *args, **kwargs):
        super().__init__(*args, **kwargs)
        # Initialize MongoDB connection (adjust credentials as needed)
        self.mongo_client = pymongo.MongoClient("mongodb://localhost:27017/")
        self.db = self.mongo_client["forum_crawler"]
        self.categories_collection = self.db["categories"]

    def parse(self, response):
        # Kick off the workflow
        yield from self.work_executor(response)

    def work_executor(self, response):
        yield from self.parse_categories(response)
        if not self.parse_only_categories:
            yield from self.parse_topics()

    def parse_categories(self, response):
        # Scrape category data and save to MongoDB
        for category in response.css("div.category-item"):
            category_data = {
                "name": category.css("h3.category-name::text").get().strip(),
                "url": response.urljoin(category.css("a.category-link::attr(href)").get()),
                "scraped": False  # Flag to track if we've processed this category
            }
            # Avoid duplicate entries
            if not self.categories_collection.find_one({"url": category_data["url"]}):
                self.categories_collection.insert_one(category_data)
            yield category_data

    def parse_topics(self):
        # Fetch unprocessed categories from MongoDB
        unprocessed_cats = self.categories_collection.find({"scraped": False})
        
        for cat in unprocessed_cats:
            # Yield request to scrape topics in this category
            yield scrapy.Request(
                url=cat["url"],
                callback=self.parse_topic_list,
                meta={"category_id": cat["_id"]}  # Pass category ID to mark as processed later
            )

    def parse_topic_list(self, response):
        # Parse all topic links in the category
        for topic in response.css("div.topic-thread"):
            topic_url = response.urljoin(topic.css("a.topic-title::attr(href)").get())
            yield scrapy.Request(
                url=topic_url,
                callback=self.parse_posts,
                meta=response.meta  # Pass category ID through to post parsing
            )

        # Handle pagination (if the forum has multiple pages of topics)
        next_page = response.css("a.next-page::attr(href)").get()
        if next_page:
            yield scrapy.Request(
                url=response.urljoin(next_page),
                callback=self.parse_topic_list,
                meta=response.meta
            )
        else:
            # Mark category as processed once all pages are scraped
            self.categories_collection.update_one(
                {"_id": response.meta["category_id"]},
                {"$set": {"scraped": True}}
            )

    def parse_posts(self, response):
        # Parse individual posts in a topic
        category = self.categories_collection.find_one({"_id": response.meta["category_id"]})
        topic_title = response.css("h1.topic-header::text").get().strip()
        
        for post in response.css("div.post-container"):
            yield {
                "category": category["name"],
                "topic_title": topic_title,
                "author": post.css("span.post-author::text").get().strip(),
                "content": "\n".join(post.css("div.post-content::text").getall()).strip(),
                "post_date": post.css("span.post-timestamp::text").get().strip()
            }

3. Common Pitfalls to Fix

  • MongoDB Connection Issues: Ensure your MongoDB service is running, and the connection string matches your setup (add username/password if using authenticated clusters).
  • Duplicate Entries: We added a check in parse_categories to avoid saving duplicate categories—adjust the unique field (like url) based on your forum's structure.
  • Blocking Operations: Pymongo is synchronous, but for most forum crawlers this won't be an issue. If you're dealing with thousands of categories, consider using an async MongoDB driver like motor to keep the crawler fast.
  • Pagination Handling: Don't forget to mark categories as processed only after scraping all topic pages (we do this in parse_topic_list when there's no next page).

内容的提问来源于stack exchange,提问作者chodi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:17:13