如何用Python爬取亚马逊搜索结果全页面?现有代码仅爬第一页
亚马逊全页面搜索结果爬取修复方案
原代码问题分析
- 分页按钮选择器错误:原代码中定位“Next”按钮的class参数写法有误,
's-pagination-item' 's-pagination-button'缺少逗号分隔,导致BeautifulSoup无法识别正确的类名组合,根本找不到分页按钮。 - 未执行页面跳转:即使找到Next按钮,代码也没有触发跳转操作,始终停留在第一页循环。
修改后的完整代码
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import pandas as pd import time import random def get_url(search_term): template = 'https://www.amazon.com/s?k={}' search_term = search_term.replace(' ', '+') url = template.format(search_term) return url def scrape_records(item): atag = item.h2.a description = atag.text.strip() url = 'https://amazon.com' + atag.get('href') price_parent = item.find('span', 'a-price') price = price_parent.find('span', 'a-offscreen').text.strip() if price_parent and price_parent.find('span', 'a-offscreen') else '' rating_element = item.find('span', {'class': 'a-icon-alt'}) rating = rating_element.text.strip() if rating_element else '' review_count_element = item.find('span', {'class': 'a-size-base', 'dir': 'auto'}) review_count = review_count_element.text.strip() if review_count_element else '' return (description, price, rating, review_count, url) def scrape_amazon(search_term): driver = webdriver.Firefox() records = [] page = 1 url = get_url(search_term) driver.get(url) time.sleep(random.uniform(2, 4)) # 随机延迟,降低反爬风险 while True: # 滚动到底部加载所有商品 driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(random.uniform(2, 3)) # 再次滚动确保加载完成(亚马逊有时候需要多次滚动) driver.execute_script("window.scrollTo(0, document.body.scrollHeight - 500);") time.sleep(random.uniform(1, 2)) soup = BeautifulSoup(driver.page_source, 'html.parser') results = soup.find_all('div', {'data-component-type': 's-search-result'}) for item in results: try: record = scrape_records(item) records.append(record) except Exception as e: print(f"爬取商品出错: {e}") # 定位并点击Next按钮 try: # 使用WebDriverWait等待Next按钮可点击,避免页面未加载完成 next_button = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.CSS_SELECTOR, 'a.s-pagination-item.s-pagination-button.s-pagination-next')) ) next_button.click() page += 1 print(f"正在爬取第 {page} 页") time.sleep(random.uniform(3, 5)) # 页面跳转后等待加载 except Exception as e: print(f"无更多页面或跳转失败: {e}") break driver.quit() # 用quit()代替close(),彻底关闭浏览器进程 # 保存为Excel df = pd.DataFrame(records, columns=['商品描述', '价格', '评分', '评论数', '商品链接']) return df # 搜索关键词 search_term = 'ultrawide monitor' # 执行爬取 df = scrape_amazon(search_term) df.to_excel('亚马逊搜索结果.xlsx', index=False) print(f"爬取完成,共获取 {len(df)} 条商品数据")
关键修改说明
- 修复分页选择器:使用CSS选择器
a.s-pagination-item.s-pagination-button.s-pagination-next精准定位Next按钮,避免类名匹配错误。 - 添加显式等待:用
WebDriverWait等待按钮可点击,解决页面加载延迟导致的元素找不到问题。 - 随机延迟替代固定等待:降低被亚马逊反爬机制检测的概率。
- 优化滚动逻辑:多次滚动确保页面商品全部加载完成。
- 替换
close()为quit():彻底关闭浏览器进程,避免残留资源。
内容的提问来源于stack exchange,提问作者John Mark Johnson
相关产品推荐
相关产品推荐

