Scrapy爬取Audible网站结果异常:存在大量空白及错误数据
问题描述
使用Scrapy框架爬取Audible网站时,整体逻辑框架没问题,但返回结果中存在大量空白数据(如title、length为None,author为空列表)及部分错误数据,期望获取每一页中每本有声书的标题(title)、作者(author)和时长(length),但结果不符合预期。
爬取代码如下:
import scrapy class AudibleSpider(scrapy.Spider): name = 'audible' allowed_domains = ['www.audible.com'] start_urls = ['https://www.audible.com/search/'] def parse(self, response): # Getting the box that contains all the info we want (title, author, length) product_container = response.xpath('//div[@class="adbl-impression-container "]//ul') # Looping through each product listed in the product_container box for product in product_container: book_title = product.xpath('.//h3[contains(@class, "bc-heading")]/a/text()').get() book_author = product.xpath('.//li[contains(@class, "authorLabel")]/span/a/text()').getall() book_length = product.xpath('.//li[contains(@class, "runtimeLabel")]/span/text()').get() # Return data extracted yield { 'title': book_title, 'author': book_author, 'length': book_length, } pagination = response.xpath('//ul[contains(@class, "pagingElements")]') next_page_url = pagination.xpath('.//span[contains(@class, "nextButton")]/a/@href').get() if next_page_url: yield response.follow(url=next_page_url, callback=self.parse)
问题分析与修复方案
核心问题
- 产品容器XPath定位错误:原代码中
product_container匹配的是adbl-impression-container下的所有ul标签,这些ul是包含多个产品的列表容器,而非单个有声书的节点,循环时会拿到整个列表而非单本书的信息,导致后续数据提取失败。 - 字段提取XPath精度不足:部分字段的XPath未准确匹配到目标文本节点,容易因页面结构小变化出现提取空白的情况。
修复后的代码
import scrapy class AudibleSpider(scrapy.Spider): name = 'audible' allowed_domains = ['www.audible.com'] start_urls = ['https://www.audible.com/search/'] def parse(self, response): # 定位到每一本有声书的单独节点 product_items = response.xpath('//li[contains(@class, "productListItem")]') for item in product_items: # 提取标题,处理空白并设置默认值 book_title = item.xpath('.//h3[contains(@class, "bc-heading")]/a/text()').get(default='').strip() # 提取作者,合并多个作者为字符串 book_author = ', '.join([author.strip() for author in item.xpath('.//li[contains(@class, "authorLabel")]//span[contains(@class, "bc-text")]/text()').getall()]) # 提取时长,处理空白并设置默认值 book_length = item.xpath('.//li[contains(@class, "runtimeLabel")]/span[contains(@class, "bc-text")]/text()').get(default='').strip() yield { 'title': book_title, 'author': book_author, 'length': book_length, } # 处理分页 next_page_url = response.xpath('//span[contains(@class, "nextButton")]/a/@href').get() if next_page_url: yield response.follow(url=next_page_url, callback=self.parse)
关键修改点
- 修正产品节点定位:改用
//li[contains(@class, "productListItem")]直接定位每一本有声书的独立节点,确保循环时处理的是单本书的信息。 - 优化字段提取逻辑:
- 标题和时长添加
strip()去除多余空格,并用get(default='')避免返回None。 - 作者提取时匹配更准确的文本节点,同时将多个作者合并为字符串,避免返回空列表。
- 标题和时长添加
- 简化分页XPath:直接定位下一页按钮的链接,无需先匹配父级
ul,减少冗余。
内容的提问来源于stack exchange,提问作者user22027271
相关产品推荐
相关产品推荐

