You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Scrapy不同函数中向JSON写入页面标题与内容

Fixing Your Scrapy Project to Capture Title + Div Content in JSON

Hey there! Since you're new to Scrapy and working with JSON outputs, let's walk through how to adjust your spider to get exactly the structure you want. The key issue most new folks run into here is passing the captured title from your title function to the parse function, and structuring your data correctly so the final JSON includes both the title and the list of content from those m1/m2 divs.

Using a Scrapy Item helps organize your data fields and makes your code cleaner. If you haven't already, add this to your project's items.py:

import scrapy
from scrapy.item import Item, Field

class PageContentItem(Item):
    title = Field()
    content_list = Field()

Corrected Spider Code

Here's how to rewrite your spider to capture the title first, then pass that data to the parse function to grab the div content:

import scrapy
from your_project_name.items import PageContentItem  # Replace with your actual project name

class MyContentSpider(scrapy.Spider):
    name = 'content_spider'
    start_urls = ['https://your-target-page.com']  # Replace with your target URL

    def start_requests(self):
        # Initiate the first request to handle the page title
        for url in self.start_urls:
            yield scrapy.Request(url, callback=self.title)

    def title(self, response):
        # Initialize the item to store the page title
        item = PageContentItem()
        # Grab the page title - adjust the selector to match your page's structure
        page_title = response.xpath('//title/text()').get()
        if page_title:
            item['title'] = page_title.strip()
        
        # Pass the item to the parse function using meta
        yield scrapy.Request(
            response.url,
            callback=self.parse,
            meta={'page_item': item}
        )

    def parse(self, response):
        # Retrieve the item with the title from meta
        item = response.meta['page_item']
        content_list = []

        # Capture content from all divs with class "m1"
        for m1_div in response.css('div.m1'):
            # Extract and clean text from the div
            raw_text = m1_div.xpath('.//text()').getall()
            clean_text = ' '.join([txt.strip() for txt in raw_text if txt.strip()])
            if clean_text:
                content_list.append(clean_text)
        
        # Capture content from all divs with class "m2"
        for m2_div in response.css('div.m2'):
            raw_text = m2_div.xpath('.//text()').getall()
            clean_text = ' '.join([txt.strip() for txt in raw_text if txt.strip()])
            if clean_text:
                content_list.append(clean_text)
        
        # Attach the content list to the item
        item['content_list'] = content_list
        # Yield the item - Scrapy will convert this to JSON automatically
        yield item

How to Generate the JSON Output

Run your spider with this command to save the results to a JSON file:

scrapy crawl content_spider -o desired_output.json

Expected JSON Structure

You'll get output that matches exactly what you're looking for:

[
    {
        "title": "Your Page Title Here",
        "content_list": [
            "Content from first m1 div",
            "Content from second m1 div",
            "Content from first m2 div"
        ]
    }
]

Quick Adjustment Tips

  • Tweak Selectors: If your page title isn't in the <title> tag (e.g., it's in an <h1>), update the selector to something like response.css('h1.page-header::text').get()
  • Simplify Text Cleaning: If you don't need to remove extra whitespace, skip the join step and use m1_div.xpath('.//text()').get() for single-line content
  • Multiple Pages: Add more URLs to start_urls and each page will generate its own entry in the JSON file

内容的提问来源于stack exchange,提问作者niloofar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 04:11:58