Scrapy多URL数据抓取问题:如何将两URL数据存入同一ScrapyItem
Scrapy多页面抓取数据关联问题
我是Scrapy新手,正在抓取目标网站,需要获取以下数据:
- quote(名言)
- author(作者)
- date of birth(出生日期)
- local of birth(出生地)
其中quote和author可从首页直接抓取;出生日期和出生地需要进入作者详情页(URL格式为https://quotes.toscrape.com/author/[NAME OF THE AUTHOR])获取。
我的items.py代码如下:
import scrapy class QuotesItem(scrapy.Item): quote = scrapy.Field() author = scrapy.Field() date_birth = scrapy.Field() local_birth = scrapy.Field()
quotespider.py代码如下:
import scrapy from ..items import QuotesItem class QuotespiderSpider(scrapy.Spider): name = "quotespider" allowed_domains = ["quotes.toscrape.com"] start_urls = ["https://quotes.toscrape.com/"] def parse(self, response): all_items= QuotesItem() quotes = response.xpath("//div[@class='row']/div[@class='col-md-8']/div") for quote in quotes: all_items['quote'] = quote.xpath("./span[@class='text']/text()").get() all_items['author'] = quote.xpath("./span[2]/small/text()").get() # 此处获取前两类数据 about = quote.xpath("./span[2]/small/following-sibling::a/@href").get() url_about = 'https://quotes.toscrape.com' + about # 跳转至作者详情页的URL yield response.follow(url_about, callback=self.about_autor, cb_kwargs={'items': all_items}) yield item def about_autor(self, response, items): # 应获取另外两类数据(date_birth, local_bith) item['date_birth '] = response.xpath("/html/body/div/div[2]/p[1]/span[1]/text()").get() item['local_bith '] = response.xpath("/html/body/div/div[2]/p[1]/span[2]/text()").get() yield item
遇到的问题
实际输出结果:
[ {"quote": "quote1", "autor": "author1", "date_birth": "", "local_birth": ""}, # 前10条数据对应字段为空 ... {"quote":"quote10", "autor": "author10", "date_birth": "", "local_birth": ""}, # 第10条数据对应字段同样为空 {"quote":"quote10", "autor": "author10", "date_birth": "December 16, 1775", "local_birth": "in Steventon"}, # 第10条数据重复,且携带错误的出生日期和出生地 {"quote": "quote10", "autor": "author10", "date_birth": "June 01, 1926", "local_birth": "United States"}, # 第10条数据重复,携带另一错误的出生日期和出生地 ]
前10条数据的local_birth和date_birth字段均为空,且最后一条数据重复出现并携带所有作者的出生日期和出生地信息。
期望输出
[{'quote': 'quote1', 'author': 'author1', 'date_birth': 'date_birth1', 'local_birth': 'local_birth1'}, {'quote': 'quote2', 'author': 'author2', 'date_birth': 'date_birth2', 'local_birth': 'local_birth2'}, {'quote': 'quote3', 'author': 'author3', 'date_birth': 'date_birth3', 'local_birth': 'local_birth'}, ]
解决方案
你的代码存在几个关键问题,修改如下:
- Item实例创建位置错误:循环外创建全局
all_items会导致所有循环复用同一个Item,数据互相覆盖。需要把Item创建放到for quote in quotes循环内部。 - 错误变量名与多余yield:
parse方法里的yield item是错误的(变量应为all_items),且此时Item无作者详情数据,需删除该语句,仅在详情页解析完成后yield完整Item。 - 字段名拼写错误:
about_autor方法里的local_bith应为local_birth,date_birth后多余空格需删除,确保与Item字段名完全一致。 - XPath优化:用类名定位替代绝对路径,提升稳定性。
修改后的quotespider.py代码:
import scrapy from ..items import QuotesItem class QuotespiderSpider(scrapy.Spider): name = "quotespider" allowed_domains = ["quotes.toscrape.com"] start_urls = ["https://quotes.toscrape.com/"] def parse(self, response): quotes = response.xpath("//div[@class='row']/div[@class='col-md-8']/div") for quote in quotes: # 每个quote创建独立Item item = QuotesItem() item['quote'] = quote.xpath("./span[@class='text']/text()").get() item['author'] = quote.xpath("./span[2]/small/text()").get() about_url = quote.xpath("./span[2]/a/@href").get() # 用response.follow自动补全域名 yield response.follow(about_url, callback=self.parse_author, cb_kwargs={'item': item}) def parse_author(self, response, item): # 提取作者详情数据,字段名与Item保持一致 item['date_birth'] = response.xpath("//span[@class='author-born-date']/text()").get() item['local_birth'] = response.xpath("//span[@class='author-born-location']/text()").get() # 数据完整后输出Item yield item
说明
- 每个名言对应独立Item,避免数据覆盖。
response.follow直接传入相对路径,Scrapy自动拼接域名,更简洁。- 作者详情页用类名定位XPath,适配页面小改动。
- 仅在详情页解析完成后yield Item,确保每条数据完整。
内容的提问来源于stack exchange,提问作者Gabriel Lucena
相关产品推荐
相关产品推荐

