You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Selenium爬取Audible搜索页CSV为空,求问题排查

问题根源及修复方案

1. 页面未完成加载就执行元素定位

audible的搜索页面包含动态加载内容,代码在打开页面后立即获取元素,此时DOM中目标元素尚未渲染完成,导致container或products为空列表,循环无法执行,最终三个数据列表都是空的。

2. CLASS_NAME参数包含多余空格

By.CLASS_NAME仅支持单个类名,你写的'adbl-impression-container '末尾带有空格,会导致无法匹配到正确元素。

3. 元素定位路径不准确

原代码中products = container.find_elements(By.XPATH, './li')的路径不符合audible实际页面结构,搜索结果项并非直接是容器的li子元素;同时标题、作者、时长的XPATH也未精准匹配页面元素的层级结构。

4. 无异常处理,定位失败直接中断

一旦某个元素定位失败,代码会抛出异常并停止执行,导致已抓取的数据也无法保存到CSV中。


修正后的完整代码

from selenium import webdriver
from selenium.webdriver.chrome.service import Service as ChromeService
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from webdriver_manager.chrome import ChromeDriverManager
import pandas as pd

driver = webdriver.Chrome(service=ChromeService(ChromeDriverManager().install()))
driver.get('https://www.audible.com/search')
driver.maximize_window()

# 显式等待容器加载完成,最长等待10秒
wait = WebDriverWait(driver, 10)
# 用CSS_SELECTOR定位容器,避免CLASS_NAME的空格问题
container = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, '.adbl-impression-container')))
# 匹配audible实际的搜索结果项结构
products = container.find_elements(By.XPATH, './/li[contains(@class, "bc-list-item")]')

book_title = []
book_author = []
book_length = []

for product in products:
    try:
        # 精准定位标题元素
        title = product.find_element(By.XPATH, './/h3[contains(@class, "bc-heading")]/a').text
        book_title.append(title)
        # 定位作者文本(跳过标签前缀)
        author = product.find_element(By.XPATH, './/li[contains(@class, "authorLabel")]/span[2]').text
        book_author.append(author)
        # 定位时长文本(跳过标签前缀)
        length = product.find_element(By.XPATH, './/li[contains(@class, "runtimeLabel")]/span[2]').text
        book_length.append(length)
    except Exception as e:
        # 捕获异常避免循环中断,缺失数据填充空值
        print(f"处理商品时出错: {e}")
        book_title.append("")
        book_author.append("")
        book_length.append("")

driver.quit()
df = pd.DataFrame({'title': book_title, 'Author': book_author, 'Length': book_length})
# 过滤空行(排除定位失败的无效数据)
df = df[df['title'].str.strip() != '']
df.to_csv('books_amazon.csv', index=False)

内容的提问来源于stack exchange,提问作者Moniruzzaman Monir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 09:13:16