You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium+Python爬取亚马逊图书:循环重复提取数据问题求助

问题分析与解决方案

问题根源

内层循环中提取图书信息的XPath使用了绝对路径(以//开头),这会让Selenium在整个页面范围内查找匹配元素,而非当前图书条目(result元素)下的子元素。因此每次循环都会提取页面中第一本图书的信息,导致重复数据。

修正方案

  1. 将提取标题、价格、评分、作者的XPath改为相对路径(以./开头),限定在当前result元素内查找子元素。
  2. 增加异常处理,应对部分图书缺失价格、评分或作者的情况,避免程序中断。
  3. 优化页面切换后的等待逻辑,确保新页面完全加载后再提取数据。

修正后的代码

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium import webdriver
from selenium.common import NoSuchElementException
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
from webdriver_manager.chrome import ChromeDriverManager

driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()))
driver.get("https://amazon.com")
driver.maximize_window()

# 搜索儿童书籍
driver.implicitly_wait(2)
searchBox = driver.find_element(By.ID, "twotabsearchtextbox")
searchBox.send_keys("children Books")
searchBtn = driver.find_element(By.ID, "nav-search-submit-button")
searchBtn.click()

booksList = []
resultsXpath = "//div[@data-component-type='s-search-result']"

# 爬取7页数据
for i in range(1, 8):
    # 等待所有搜索结果加载完成
    WebDriverWait(driver, 25).until(EC.visibility_of_all_elements_located((By.XPATH, resultsXpath)))
    results = driver.find_elements(By.XPATH, resultsXpath)

    for result in results:
        book_data = {}
        # 提取标题(相对路径)
        try:
            title = result.find_element(By.XPATH, "./div/div/div/div[2]/div[2]/div[1]/div[1]/h2/a/span").text.strip()
            book_data['book title'] = title
        except NoSuchElementException:
            book_data['book title'] = "N/A"
        
        # 提取价格(相对路径)
        try:
            price = result.find_element(By.XPATH, "./div/div/div/div[2]/div[2]/div[3]/div[1]/span[1]/span[2]").text
            book_data['price'] = price
        except NoSuchElementException:
            book_data['price'] = "N/A"
        
        # 提取评分(相对路径)
        try:
            rating = result.find_element(By.XPATH, "./div/div/div/div[2]/div[2]/div[2]/div[1]/span[1]/span[1]").text
            book_data['rating'] = rating
        except NoSuchElementException:
            book_data['rating'] = "N/A"
        
        # 提取作者(相对路径)
        try:
            auth = result.find_element(By.XPATH, "./div/div/div/div[2]/div[2]/div[1]/div[2]/div[1]/span").text.strip()
            book_data['authors'] = auth
        except NoSuchElementException:
            book_data['authors'] = "N/A"
        
        booksList.append(book_data)

    # 点击下一页并等待页面切换完成
    try:
        nextBtn = driver.find_element(By.XPATH, "//a[@class='s-pagination-item s-pagination-next s-pagination-button s-pagination-separator']")
        nextBtn.click()
        WebDriverWait(driver, 20).until(EC.staleness_of(results[0]))
    except NoSuchElementException:
        break

# 打印提取结果
for book in booksList:
    print(book)

print(f"\n共提取{len(booksList)}条图书数据")
driver.quit()

关键说明

  • 相对路径:使用./开头的XPath,确保仅在当前图书条目节点下查找子元素,避免全局匹配导致的重复数据。
  • 异常处理:针对预售图书、新书等可能缺失字段的情况添加try-except,保证程序稳定运行。
  • 页面切换等待:点击下一页后,通过staleness_of等待旧页面元素失效,确保新页面完全加载后再执行后续逻辑。

内容的提问来源于stack exchange,提问作者SadiqHussain

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 16:20:31