Scrapy多页面采集问题:如何将作者详情合并至同一条Item
解决Scrapy中合并作者详情到同一条Item且不生成新行的问题
你的代码存在两个核心问题导致字段缺失和多生成行:
- 复用同一个类属性
self.quote_item,循环中Item对象被重复覆盖,且异步请求执行时会导致数据混乱 - 在
parse中提前yield Item,此时parse_page还未完成作者信息的填充,所以输出缺少birth_date和birth_location;若在parse_page单独yield又会生成新行
修改后的代码
def parse(self, response): qs = response.css('.quote') for q in qs: # 每个quote创建独立的Item,避免复用覆盖 quote_item = QuotesItem() # 提取当前quote的字段,去掉末尾逗号避免生成元组 quote_item['quote'] = q.css('.text ::text').get() # 用getall()简化tags提取 quote_item['tag'] = q.css('.tag ::text').getall() quote_item['author'] = q.css('span .author ::text').get() # 获取作者详情页链接,传给parse_page author_url = q.css('span a::attr(href)').get() # 用response.follow自动拼接域名,更简洁 yield response.follow( author_url, callback=self.parse_page, cb_kwargs={'item': quote_item} ) def parse_page(self, response, item): # 提取作者详情字段 author_details = response.css('.author-details') item['birth_date'] = author_details.css('.author-born-date ::text').get() item['birth_location'] = author_details.css('.author-born-location ::text').get() # 填充完所有字段后统一yield,确保一条数据包含所有内容 yield item
关键修改说明
- 每个Quote独立实例化Item:循环内每次创建新的
QuotesItem,彻底避免数据覆盖和异步请求的线程安全问题 - 简化字段提取逻辑:用
getall()替代手动循环提取标签,代码更高效简洁 - 延迟yield时机:不在
parse中提前输出Item,而是将未完成的Item传递给parse_page,待填充完作者信息后,在parse_page中统一yield完整的Item,确保一条数据包含所有字段且仅生成一行 - 优化链接处理:使用
response.follow自动拼接域名,无需手动拼接URL,降低出错概率
预期输出示例
{"quote": "“One good thing about music, when it hits you, you feel no pain.”", "tag": ["music"], "author": "Bob Marley", "birth_date": "February 06, 1945", "birth_location": "Nine Mile, Saint Ann, Jamaica"}
内容的提问来源于stack exchange,提问作者Low LiFe
相关产品推荐
相关产品推荐

