You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取亚马逊评论时翻页无输出问题求助

亚马逊评论爬取翻页异常解决方法

核心问题定位

原代码中reviewsHtml函数构造了分页参数params,但调用requests.get()时没有传入该参数,导致所有请求都指向第一页,自然出现内容重复或无新数据的情况。

修复后的代码

import requests
import pandas as pd
from bs4 import BeautifulSoup
import time
import random

# 完善请求头,模拟真实浏览器行为
headers = {
    'accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9',
    'accept-language': 'en-US,en;q=0.9',
    'sec-ch-ua': '"Not A;Brand";v="99", "Chromium";v="106", "Google Chrome";v="106"',
    'sec-ch-ua-mobile': '?0',
    'sec-ch-ua-platform': '"Windows"',
    'sec-fetch-dest': 'document',
    'sec-fetch-mode': 'navigate',
    'sec-fetch-site': 'same-origin',
    'sec-fetch-user': '?1',
    'upgrade-insecure-requests': '1',
    'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/106.0.0.0 Safari/537.36'
}

reviews_url = 'https://www.amazon.com/Legendary-Whitetails-Journeyman-Jacket-Tarmac/product-reviews/B013KW38RQ/'
len_page = 3

def reviewsHtml(url, len_page):
    soups = []
    # 使用Session维持会话,保留cookies
    session = requests.Session()
    session.headers.update(headers)
    
    # 先请求首页获取必要会话信息
    session.get(url)
    
    for page_no in range(1, len_page + 1):
        params = {
            'ie': 'UTF8',
            'reviewerType': 'all_reviews',
            'filterByStar': 'critical',
            'pageNumber': page_no,
            'sortBy': 'recent'  # 加入排序参数,避免缓存重复内容
        }
        # 关键:传入分页参数
        response = session.get(url, params=params)
        # 检查请求状态
        if response.status_code != 200:
            print(f"第{page_no}页请求失败,状态码:{response.status_code}")
            continue
        soup = BeautifulSoup(response.text, 'lxml')
        soups.append(soup)
        # 随机延迟,降低反爬风险
        time.sleep(random.uniform(1.5, 3))
    return soups

def getReviews(html_data):
    data_dicts = []
    # 修正选择器引号问题,避免转义错误
    boxes = html_data.select('div[data-hook="review"]')
    if not boxes:
        print("未找到评论区块,可能触发反爬或页面结构变更")
    for box in boxes:
        title = box.select_one('[data-hook="review-title"]').text.strip() if box.select_one('[data-hook="review-title"]') else 'N/A'
        description = box.select_one('[data-hook="review-body"]').text.strip() if box.select_one('[data-hook="review-body"]') else 'N/A'
        data_dicts.append({
            'Title': title,
            'Description': description
        })
    return data_dicts

reviews = []
html_datas = reviewsHtml(reviews_url, len_page)
for html_data in html_datas:
    review_batch = getReviews(html_data)
    reviews += review_batch

df_reviews = pd.DataFrame(reviews)
print(df_reviews)

额外优化说明

  • 会话维持:用requests.Session()保持会话,保留亚马逊设置的cookies,降低被反爬识别的概率。
  • 反爬应对:加入随机延迟,避免请求频率过高;完善请求头字段,让请求更接近真实浏览器行为。
  • 错误排查:添加状态码检查、评论区块空值判断,方便快速定位问题。
  • 参数适配:亚马逊的分页参数和页面结构可能随时调整,若再次出现异常,需检查pageNumber参数是否有效,或评论区块的选择器是否需要更新。

内容的提问来源于stack exchange,提问作者Indirakanth

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 19:07:03