You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python爬取亚马逊搜索结果全页面?现有代码仅爬第一页

亚马逊全页面搜索结果爬取修复方案

原代码问题分析

  1. 分页按钮选择器错误:原代码中定位“Next”按钮的class参数写法有误,'s-pagination-item' 's-pagination-button'缺少逗号分隔,导致BeautifulSoup无法识别正确的类名组合,根本找不到分页按钮。
  2. 未执行页面跳转:即使找到Next按钮,代码也没有触发跳转操作,始终停留在第一页循环。

修改后的完整代码

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import pandas as pd
import time
import random

def get_url(search_term):
    template = 'https://www.amazon.com/s?k={}'
    search_term = search_term.replace(' ', '+')
    url = template.format(search_term)
    return url

def scrape_records(item):
    atag = item.h2.a
    description = atag.text.strip()
    url = 'https://amazon.com' + atag.get('href')

    price_parent = item.find('span', 'a-price')
    price = price_parent.find('span', 'a-offscreen').text.strip() if price_parent and price_parent.find('span', 'a-offscreen') else ''

    rating_element = item.find('span', {'class': 'a-icon-alt'})
    rating = rating_element.text.strip() if rating_element else ''

    review_count_element = item.find('span', {'class': 'a-size-base', 'dir': 'auto'})
    review_count = review_count_element.text.strip() if review_count_element else ''

    return (description, price, rating, review_count, url)

def scrape_amazon(search_term):
    driver = webdriver.Firefox()
    records = []
    page = 1

    url = get_url(search_term)
    driver.get(url)
    time.sleep(random.uniform(2, 4))  # 随机延迟,降低反爬风险

    while True:
        # 滚动到底部加载所有商品
        driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
        time.sleep(random.uniform(2, 3))
        
        # 再次滚动确保加载完成(亚马逊有时候需要多次滚动)
        driver.execute_script("window.scrollTo(0, document.body.scrollHeight - 500);")
        time.sleep(random.uniform(1, 2))

        soup = BeautifulSoup(driver.page_source, 'html.parser')
        results = soup.find_all('div', {'data-component-type': 's-search-result'})

        for item in results:
            try:
                record = scrape_records(item)
                records.append(record)
            except Exception as e:
                print(f"爬取商品出错: {e}")

        # 定位并点击Next按钮
        try:
            # 使用WebDriverWait等待Next按钮可点击,避免页面未加载完成
            next_button = WebDriverWait(driver, 10).until(
                EC.element_to_be_clickable((By.CSS_SELECTOR, 'a.s-pagination-item.s-pagination-button.s-pagination-next'))
            )
            next_button.click()
            page += 1
            print(f"正在爬取第 {page} 页")
            time.sleep(random.uniform(3, 5))  # 页面跳转后等待加载
        except Exception as e:
            print(f"无更多页面或跳转失败: {e}")
            break

    driver.quit()  # 用quit()代替close(),彻底关闭浏览器进程

    # 保存为Excel
    df = pd.DataFrame(records, columns=['商品描述', '价格', '评分', '评论数', '商品链接'])
    return df

# 搜索关键词
search_term = 'ultrawide monitor'

# 执行爬取
df = scrape_amazon(search_term)
df.to_excel('亚马逊搜索结果.xlsx', index=False)
print(f"爬取完成,共获取 {len(df)} 条商品数据")

关键修改说明

  • 修复分页选择器:使用CSS选择器a.s-pagination-item.s-pagination-button.s-pagination-next精准定位Next按钮,避免类名匹配错误。
  • 添加显式等待:用WebDriverWait等待按钮可点击,解决页面加载延迟导致的元素找不到问题。
  • 随机延迟替代固定等待:降低被亚马逊反爬机制检测的概率。
  • 优化滚动逻辑:多次滚动确保页面商品全部加载完成。
  • 替换close()为quit():彻底关闭浏览器进程,避免残留资源。

内容的提问来源于stack exchange,提问作者John Mark Johnson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 06:37:27