You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy多URL数据抓取问题:如何将两URL数据存入同一ScrapyItem

Scrapy多页面抓取数据关联问题

我是Scrapy新手,正在抓取目标网站,需要获取以下数据:

  • quote(名言)
  • author(作者)
  • date of birth(出生日期)
  • local of birth(出生地)

其中quote和author可从首页直接抓取;出生日期和出生地需要进入作者详情页(URL格式为https://quotes.toscrape.com/author/[NAME OF THE AUTHOR])获取。

我的items.py代码如下:

import scrapy

class QuotesItem(scrapy.Item):
    quote = scrapy.Field()
    author = scrapy.Field()
    date_birth = scrapy.Field()
    local_birth = scrapy.Field()

quotespider.py代码如下:

import scrapy
from ..items import QuotesItem  


class QuotespiderSpider(scrapy.Spider):
    name = "quotespider"
    allowed_domains = ["quotes.toscrape.com"]
    start_urls = ["https://quotes.toscrape.com/"]

    def parse(self, response):
        all_items= QuotesItem()

        quotes = response.xpath("//div[@class='row']/div[@class='col-md-8']/div")

        for quote in quotes:
            all_items['quote'] = quote.xpath("./span[@class='text']/text()").get()
            all_items['author'] = quote.xpath("./span[2]/small/text()").get()
            # 此处获取前两类数据

            about = quote.xpath("./span[2]/small/following-sibling::a/@href").get()
            url_about = 'https://quotes.toscrape.com' + about  # 跳转至作者详情页的URL

            yield response.follow(url_about, callback=self.about_autor,
                                  cb_kwargs={'items': all_items})
            
            yield item


    def about_autor(self, response, items):  # 应获取另外两类数据(date_birth, local_bith)

        item['date_birth '] = response.xpath("/html/body/div/div[2]/p[1]/span[1]/text()").get()
        item['local_bith '] = response.xpath("/html/body/div/div[2]/p[1]/span[2]/text()").get()

        yield item

遇到的问题

实际输出结果:

[
{"quote": "quote1", 
"autor": "author1",  
"date_birth": "",
 "local_birth": ""}, # 前10条数据对应字段为空
...
{"quote":"quote10", 
"autor": "author10", 
"date_birth": "", 
"local_birth": ""}, # 第10条数据对应字段同样为空

{"quote":"quote10", 
"autor": "author10", 
"date_birth": "December 16, 1775", 
"local_birth": "in Steventon"},  # 第10条数据重复,且携带错误的出生日期和出生地

{"quote": "quote10", 
"autor": "author10", 
"date_birth": "June 01, 1926", 
"local_birth": "United States"}, # 第10条数据重复,携带另一错误的出生日期和出生地
]

前10条数据的local_birth和date_birth字段均为空,且最后一条数据重复出现并携带所有作者的出生日期和出生地信息。

期望输出

[{'quote': 'quote1',
'author': 'author1',
'date_birth': 'date_birth1',
'local_birth': 'local_birth1'},

{'quote': 'quote2',
'author': 'author2',
'date_birth': 'date_birth2',
'local_birth': 'local_birth2'},

{'quote': 'quote3',
'author': 'author3',
'date_birth': 'date_birth3',
'local_birth': 'local_birth'},
]

解决方案

你的代码存在几个关键问题,修改如下:

  1. Item实例创建位置错误:循环外创建全局all_items会导致所有循环复用同一个Item,数据互相覆盖。需要把Item创建放到for quote in quotes循环内部。
  2. 错误变量名与多余yield:parse方法里的yield item是错误的(变量应为all_items),且此时Item无作者详情数据,需删除该语句,仅在详情页解析完成后yield完整Item。
  3. 字段名拼写错误:about_autor方法里的local_bith应为local_birth,date_birth 后多余空格需删除,确保与Item字段名完全一致。
  4. XPath优化:用类名定位替代绝对路径,提升稳定性。

修改后的quotespider.py代码:

import scrapy
from ..items import QuotesItem  


class QuotespiderSpider(scrapy.Spider):
    name = "quotespider"
    allowed_domains = ["quotes.toscrape.com"]
    start_urls = ["https://quotes.toscrape.com/"]

    def parse(self, response):
        quotes = response.xpath("//div[@class='row']/div[@class='col-md-8']/div")

        for quote in quotes:
            # 每个quote创建独立Item
            item = QuotesItem()
            item['quote'] = quote.xpath("./span[@class='text']/text()").get()
            item['author'] = quote.xpath("./span[2]/small/text()").get()

            about_url = quote.xpath("./span[2]/a/@href").get()
            # 用response.follow自动补全域名
            yield response.follow(about_url, callback=self.parse_author, cb_kwargs={'item': item})

    def parse_author(self, response, item):
        # 提取作者详情数据,字段名与Item保持一致
        item['date_birth'] = response.xpath("//span[@class='author-born-date']/text()").get()
        item['local_birth'] = response.xpath("//span[@class='author-born-location']/text()").get()
        
        # 数据完整后输出Item
        yield item

说明

  • 每个名言对应独立Item,避免数据覆盖。
  • response.follow直接传入相对路径,Scrapy自动拼接域名,更简洁。
  • 作者详情页用类名定位XPath,适配页面小改动。
  • 仅在详情页解析完成后yield Item,确保每条数据完整。

内容的提问来源于stack exchange,提问作者Gabriel Lucena

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 20:57:12