如何在Scrapy不同函数中向JSON写入页面标题与内容
Hey there! Since you're new to Scrapy and working with JSON outputs, let's walk through how to adjust your spider to get exactly the structure you want. The key issue most new folks run into here is passing the captured title from your title function to the parse function, and structuring your data correctly so the final JSON includes both the title and the list of content from those m1/m2 divs.
First, Define a Scrapy Item (Optional but Recommended)
Using a Scrapy Item helps organize your data fields and makes your code cleaner. If you haven't already, add this to your project's items.py:
import scrapy from scrapy.item import Item, Field class PageContentItem(Item): title = Field() content_list = Field()
Corrected Spider Code
Here's how to rewrite your spider to capture the title first, then pass that data to the parse function to grab the div content:
import scrapy from your_project_name.items import PageContentItem # Replace with your actual project name class MyContentSpider(scrapy.Spider): name = 'content_spider' start_urls = ['https://your-target-page.com'] # Replace with your target URL def start_requests(self): # Initiate the first request to handle the page title for url in self.start_urls: yield scrapy.Request(url, callback=self.title) def title(self, response): # Initialize the item to store the page title item = PageContentItem() # Grab the page title - adjust the selector to match your page's structure page_title = response.xpath('//title/text()').get() if page_title: item['title'] = page_title.strip() # Pass the item to the parse function using meta yield scrapy.Request( response.url, callback=self.parse, meta={'page_item': item} ) def parse(self, response): # Retrieve the item with the title from meta item = response.meta['page_item'] content_list = [] # Capture content from all divs with class "m1" for m1_div in response.css('div.m1'): # Extract and clean text from the div raw_text = m1_div.xpath('.//text()').getall() clean_text = ' '.join([txt.strip() for txt in raw_text if txt.strip()]) if clean_text: content_list.append(clean_text) # Capture content from all divs with class "m2" for m2_div in response.css('div.m2'): raw_text = m2_div.xpath('.//text()').getall() clean_text = ' '.join([txt.strip() for txt in raw_text if txt.strip()]) if clean_text: content_list.append(clean_text) # Attach the content list to the item item['content_list'] = content_list # Yield the item - Scrapy will convert this to JSON automatically yield item
How to Generate the JSON Output
Run your spider with this command to save the results to a JSON file:
scrapy crawl content_spider -o desired_output.json
Expected JSON Structure
You'll get output that matches exactly what you're looking for:
[ { "title": "Your Page Title Here", "content_list": [ "Content from first m1 div", "Content from second m1 div", "Content from first m2 div" ] } ]
Quick Adjustment Tips
- Tweak Selectors: If your page title isn't in the
<title>tag (e.g., it's in an<h1>), update the selector to something likeresponse.css('h1.page-header::text').get() - Simplify Text Cleaning: If you don't need to remove extra whitespace, skip the
joinstep and usem1_div.xpath('.//text()').get()for single-line content - Multiple Pages: Add more URLs to
start_urlsand each page will generate its own entry in the JSON file
内容的提问来源于stack exchange,提问作者niloofar

