You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy中如何将birth_date合并到主parse函数的输出结果?

问题:Scrapy爬虫无法将birth_date字段与主数据一同输出

作为Python及Scrapy库的初学者,我尝试将爬取到的作者birth_date字段,与包含title、author、url_author等的主输出数据一同yield,但输出结果中始终缺少birth_date字段。

原爬虫代码

import scrapy

class QuotespiderSpider(scrapy.Spider):
    name = "quotespider"
    allowed_domains = ["quotes.toscrape.com"]
    start_urls = ["https://quotes.toscrape.com"]

    def parse(self, response):
        quotes = response.css(".quote")
        all_url_tags = quotes.css(".tags a.tag::attr(href)").getall()

        for quote in quotes:
            url_tags = []
            tags = (quote.css(".tags meta.keywords::attr(content)").get()).split(',')
            url_tags = [j for i in tags for j in all_url_tags if i in j]

            url_author = quote.css("span a::attr(href)").get()
            author_page = "https://quotes.toscrape.com/" + url_author

            items = {
                'title' : quote.css("span.text::text").get(),
                'author' : quote.css("span small.author::text").get(),
                'url_author' : quote.css("span a::attr(href)").get(),
                'tags' : tags,
                'url_tags' : url_tags
                }

            yield scrapy.Request(url=author_page, callback=self.parse_author_page, meta={'items':items})
            yield items

        next_page = response.css("li.next a::attr(href)").get()

        if next_page is not None:
            next_page_url = "https://quotes.toscrape.com/" + next_page
            yield response.follow(next_page_url, callback=self.parse)

    def parse_author_page(self, response):
        items = response.meta['items']
        birth_date = response.css(".author-details p span.author-born-date::text").get()
        items['birth_date'] = birth_date
        return items

当前输出示例(缺少birth_date)

{"title": "“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”", "author": "Albert Einstein", "url_author": "/author/Albert-Einstein", "tags": ["change", "deep-thoughts", "thinking", "world"], "url_tags": ["/tag/change/page/1/", "/tag/deep-thoughts/page/1/", "/tag/thinking/page/1/", "/tag/world/page/1/"]}

解决方法

问题核心:你在parse函数里先yield了未包含birth_date的items,而parse_author_page中仅返回补充完birth_date的items,但未将其yield出去,导致最终输出只有原始的不完整数据。

修改方案:去掉parse里的yield items,转而在parse_author_page处理完birth_date后,统一yield完整的items。

修改后的代码

import scrapy

class QuotespiderSpider(scrapy.Spider):
    name = "quotespider"
    allowed_domains = ["quotes.toscrape.com"]
    start_urls = ["https://quotes.toscrape.com"]

    def parse(self, response):
        quotes = response.css(".quote")
        all_url_tags = quotes.css(".tags a.tag::attr(href)").getall()

        for quote in quotes:
            url_tags = []
            tags = (quote.css(".tags meta.keywords::attr(content)").get()).split(',')
            url_tags = [j for i in tags for j in all_url_tags if i in j]

            url_author = quote.css("span a::attr(href)").get()
            author_page = "https://quotes.toscrape.com/" + url_author

            items = {
                'title' : quote.css("span.text::text").get(),
                'author' : quote.css("span small.author::text").get(),
                'url_author' : quote.css("span a::attr(href)").get(),
                'tags' : tags,
                'url_tags' : url_tags
                }

            # 仅发送请求到作者页,不在此处输出原始items
            yield scrapy.Request(url=author_page, callback=self.parse_author_page, meta={'items':items})

        next_page = response.css("li.next a::attr(href)").get()

        if next_page is not None:
            next_page_url = "https://quotes.toscrape.com/" + next_page
            yield response.follow(next_page_url, callback=self.parse)

    def parse_author_page(self, response):
        items = response.meta['items']
        birth_date = response.css(".author-details p span.author-born-date::text").get()
        items['birth_date'] = birth_date
        # 此处输出包含birth_date的完整数据
        yield items

修改后输出示例(包含birth_date)

{"title": "“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”", "author": "Albert Einstein", "url_author": "/author/Albert-Einstein", "tags": ["change", "deep-thoughts", "thinking", "world"], "url_tags": ["/tag/change/page/1/", "/tag/deep-thoughts/page/1/", "/tag/thinking/page/1/", "/tag/world/page/1/"], "birth_date": "March 14, 1879"}

内容的提问来源于stack exchange,提问作者Siroj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 04:57:28