You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬取Audible网站结果异常:存在大量空白及错误数据

问题描述

使用Scrapy框架爬取Audible网站时,整体逻辑框架没问题,但返回结果中存在大量空白数据(如title、length为None,author为空列表)及部分错误数据,期望获取每一页中每本有声书的标题(title)、作者(author)和时长(length),但结果不符合预期。

爬取代码如下:

import scrapy

class AudibleSpider(scrapy.Spider):
    name = 'audible'
    allowed_domains = ['www.audible.com']
    start_urls = ['https://www.audible.com/search/']

    def parse(self, response):
        # Getting the box that contains all the info we want (title, author, length)
        product_container = response.xpath('//div[@class="adbl-impression-container "]//ul')

        # Looping through each product listed in the product_container box
        for product in product_container:
            book_title = product.xpath('.//h3[contains(@class, "bc-heading")]/a/text()').get()
            book_author = product.xpath('.//li[contains(@class, "authorLabel")]/span/a/text()').getall()
            book_length = product.xpath('.//li[contains(@class, "runtimeLabel")]/span/text()').get()

            # Return data extracted
            yield {
                'title': book_title,
                'author': book_author,
                'length': book_length,
            }
            

        pagination = response.xpath('//ul[contains(@class, "pagingElements")]')
        next_page_url = pagination.xpath('.//span[contains(@class, "nextButton")]/a/@href').get()
        if next_page_url:
            yield response.follow(url=next_page_url, callback=self.parse)
问题分析与修复方案

核心问题

  1. 产品容器XPath定位错误:原代码中product_container匹配的是adbl-impression-container下的所有ul标签,这些ul是包含多个产品的列表容器,而非单个有声书的节点,循环时会拿到整个列表而非单本书的信息,导致后续数据提取失败。
  2. 字段提取XPath精度不足:部分字段的XPath未准确匹配到目标文本节点,容易因页面结构小变化出现提取空白的情况。

修复后的代码

import scrapy

class AudibleSpider(scrapy.Spider):
    name = 'audible'
    allowed_domains = ['www.audible.com']
    start_urls = ['https://www.audible.com/search/']

    def parse(self, response):
        # 定位到每一本有声书的单独节点
        product_items = response.xpath('//li[contains(@class, "productListItem")]')

        for item in product_items:
            # 提取标题,处理空白并设置默认值
            book_title = item.xpath('.//h3[contains(@class, "bc-heading")]/a/text()').get(default='').strip()
            # 提取作者,合并多个作者为字符串
            book_author = ', '.join([author.strip() for author in item.xpath('.//li[contains(@class, "authorLabel")]//span[contains(@class, "bc-text")]/text()').getall()])
            # 提取时长,处理空白并设置默认值
            book_length = item.xpath('.//li[contains(@class, "runtimeLabel")]/span[contains(@class, "bc-text")]/text()').get(default='').strip()

            yield {
                'title': book_title,
                'author': book_author,
                'length': book_length,
            }

        # 处理分页
        next_page_url = response.xpath('//span[contains(@class, "nextButton")]/a/@href').get()
        if next_page_url:
            yield response.follow(url=next_page_url, callback=self.parse)

关键修改点

  • 修正产品节点定位:改用//li[contains(@class, "productListItem")]直接定位每一本有声书的独立节点,确保循环时处理的是单本书的信息。
  • 优化字段提取逻辑:
    • 标题和时长添加strip()去除多余空格,并用get(default='')避免返回None。
    • 作者提取时匹配更准确的文本节点,同时将多个作者合并为字符串,避免返回空列表。
  • 简化分页XPath:直接定位下一页按钮的链接,无需先匹配父级ul,减少冗余。

内容的提问来源于stack exchange,提问作者user22027271

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 15:57:02