You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取亚马逊家具产品遭封禁的问题求助

亚马逊爬虫封禁问题求助与解决方案

问题描述

我使用BeautifulSoup爬取亚马逊家具产品,最初未添加请求头时直接被封禁,返回提示:

如需讨论亚马逊数据的自动化访问,请联系api-services-support@amazon.com

添加请求头HEADERS(内容为{'User-Agent':'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko)Chrome/44.0.2403.157 Safari/537.36','Accept-Language': 'en-US, en;q=0.5'})后,前2-3次爬取成功并将数据写入CSV,但第3次之后再次被封禁,爬取无结果且CSV为空。尝试过Stack Overflow上的解决方案,添加代理也无效,作为爬虫新手请求解决。

我的代码

from bs4 import BeautifulSoup
import requests

# 提取商品标题
def get_title(soup):
    try:
        title = soup.find("span", attrs={"id":"productTitle"})
        title_value = title.string
        title_string = title_value.strip()
    except AttributeError:
        title_string = ""   
    return title_string

# 提取商品价格
def get_price(soup):
    try:
        price = soup.find("span", attrs={'id':'priceblock_ourprice'}).string.strip()
    except AttributeError:
        try:
            # 处理促销价格
            price = soup.find("span", attrs={'id':'priceblock_dealprice'}).string.strip()
        except:     
            price = ""  
    return price

# 提取商品评分
def get_rating(soup):
    try:
        rating = soup.find("i", attrs={'class':'a-icon a-icon-star a-star-4-5'}).string.strip()
    except AttributeError:
        try:
            rating = soup.find("span", attrs={'class':'a-icon-alt'}).string.strip()
        except:
            rating = "" 
    return rating

# 提取评论数量
def get_review_count(soup):
    try:
        review_count = soup.find("span", attrs={'id':'acrCustomerReviewText'}).string.strip()
    except AttributeError:
        review_count = ""   
    return review_count

# 提取库存状态
def get_availability(soup):
    try:
        available = soup.find("div", attrs={'id':'availability'})
        available = available.find("span").string.strip()
    except AttributeError:
        available = "Not Available" 
    return available    

if __name__ == '__main__':
    # 请求头
    HEADERS = ({'User-Agent':'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko)Chrome/44.0.2403.157 Safari/537.36',
                                'Accept-Language': 'en-US, en;q=0.5'})
    # 搜索页URL
    URL = "https://www.amazon.com/s?k=furniture&crid=3C1AP0SFA5J8Y&sprefix=furniture%2Caps%2C389&ref=nb_sb_noss_1s"
    # 发起请求
    webpage = requests.get(URL, headers=HEADERS)
    # 解析页面
    soup = BeautifulSoup(webpage.content, "lxml")
    # 获取商品链接
    links = soup.find_all("a", attrs={'class':'a-link-normal s-no-outline'})
    links_list = []
    for link in links:
        links_list.append(link.get('href'))

    # 遍历商品链接提取详情
    for link in links_list:
        File = open("product_record.csv", "a")
        new_webpage = requests.get("https://www.amazon.com" + link, headers=HEADERS)
        new_soup = BeautifulSoup(new_webpage.content, "lxml")
        
        # 输出并写入数据
        print("Product Title =", get_title(new_soup))
        File.write(f"{get_title(new_soup)},")
        print("Product Price =", get_price(new_soup))
        File.write(f"{get_price(new_soup)},")
        print("Product Rating =", get_rating(new_soup))
        File.write(f"{get_rating(new_soup)},")
        print("Number of Product Reviews =", get_review_count(new_soup))
        File.write(f"{get_review_count(new_soup)},")
        print("Availability =", get_availability(new_soup))
        File.write(f"{get_availability(new_soup)},")
        File.close()
        print()
        print()

解决建议

1. 更新请求头,模拟真实浏览器

你的User-Agent版本过于老旧(Chrome 44),亚马逊极易识别。换成当前主流浏览器的User-Agent,并补充更多字段让请求更接近真人操作:

HEADERS = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36',
    'Accept-Language': 'en-US,en;q=0.9',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8',
    'Referer': 'https://www.amazon.com/',
    'DNT': '1'
}

2. 添加随机请求延迟

短时间内高频请求是触发反爬的核心原因,在每次爬取商品链接后添加2-5秒的随机延迟:

import time
import random

# 放在每个商品请求完成后
time.sleep(random.uniform(2, 5))

3. 优化代理使用(若需)

确保使用高匿代理,且每次请求轮换代理池中的IP,避免单一代理被封禁:

proxies_list = [
    {'http': 'http://proxy1:port', 'https': 'https://proxy1:port'},
    {'http': 'http://proxy2:port', 'https': 'https://proxy2:port'},
    # 补充更多有效代理
]

# 每次请求随机选择代理
random_proxy = random.choice(proxies_list)
new_webpage = requests.get("https://www.amazon.com" + link, headers=HEADERS, proxies=random_proxy)

4. 使用会话维持Cookie

亚马逊通过Cookie识别会话,用requests.Session()维持会话,避免每次请求都是新的无状态访问:

session = requests.Session()
session.headers.update(HEADERS)

# 先访问亚马逊首页获取有效Cookie
session.get("https://www.amazon.com/")

# 后续所有请求都使用session
webpage = session.get(URL)
new_webpage = session.get("https://www.amazon.com" + link)

5. 修正代码错误与CSV写入优化

  • 原代码中get_review_count和get_availability函数的缩进错误已修正。
  • 每次循环打开关闭文件效率低下且易出错,改用csv模块规范写入:
import csv

with open("product_record.csv", "a", newline='', encoding='utf-8') as file:
    writer = csv.writer(file)
    # 首次运行时可写入表头
    # writer.writerow(["商品标题", "价格", "评分", "评论数", "库存状态"])
    for link in links_list:
        new_webpage = session.get("https://www.amazon.com" + link)
        new_soup = BeautifulSoup(new_webpage.content, "lxml")
        # 写入数据
        writer.writerow([
            get_title(new_soup),
            get_price(new_soup),
            get_rating(new_soup),
            get_review_count(new_soup),
            get_availability(new_soup)
        ])
        time.sleep(random.uniform(2, 5))

内容的提问来源于stack exchange,提问作者Syed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 18:57:07