Python Selenium爬取Audible搜索页CSV为空,求问题排查
问题根源及修复方案
1. 页面未完成加载就执行元素定位
audible的搜索页面包含动态加载内容,代码在打开页面后立即获取元素,此时DOM中目标元素尚未渲染完成,导致container或products为空列表,循环无法执行,最终三个数据列表都是空的。
2. CLASS_NAME参数包含多余空格
By.CLASS_NAME仅支持单个类名,你写的'adbl-impression-container '末尾带有空格,会导致无法匹配到正确元素。
3. 元素定位路径不准确
原代码中products = container.find_elements(By.XPATH, './li')的路径不符合audible实际页面结构,搜索结果项并非直接是容器的li子元素;同时标题、作者、时长的XPATH也未精准匹配页面元素的层级结构。
4. 无异常处理,定位失败直接中断
一旦某个元素定位失败,代码会抛出异常并停止执行,导致已抓取的数据也无法保存到CSV中。
修正后的完整代码
from selenium import webdriver from selenium.webdriver.chrome.service import Service as ChromeService from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from webdriver_manager.chrome import ChromeDriverManager import pandas as pd driver = webdriver.Chrome(service=ChromeService(ChromeDriverManager().install())) driver.get('https://www.audible.com/search') driver.maximize_window() # 显式等待容器加载完成,最长等待10秒 wait = WebDriverWait(driver, 10) # 用CSS_SELECTOR定位容器,避免CLASS_NAME的空格问题 container = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, '.adbl-impression-container'))) # 匹配audible实际的搜索结果项结构 products = container.find_elements(By.XPATH, './/li[contains(@class, "bc-list-item")]') book_title = [] book_author = [] book_length = [] for product in products: try: # 精准定位标题元素 title = product.find_element(By.XPATH, './/h3[contains(@class, "bc-heading")]/a').text book_title.append(title) # 定位作者文本(跳过标签前缀) author = product.find_element(By.XPATH, './/li[contains(@class, "authorLabel")]/span[2]').text book_author.append(author) # 定位时长文本(跳过标签前缀) length = product.find_element(By.XPATH, './/li[contains(@class, "runtimeLabel")]/span[2]').text book_length.append(length) except Exception as e: # 捕获异常避免循环中断,缺失数据填充空值 print(f"处理商品时出错: {e}") book_title.append("") book_author.append("") book_length.append("") driver.quit() df = pd.DataFrame({'title': book_title, 'Author': book_author, 'Length': book_length}) # 过滤空行(排除定位失败的无效数据) df = df[df['title'].str.strip() != ''] df.to_csv('books_amazon.csv', index=False)
内容的提问来源于stack exchange,提问作者Moniruzzaman Monir
相关产品推荐
相关产品推荐

