You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python亚马逊价格爬虫可靠性问题求助:多数无法获取价格元素

问题描述

我是Python新手,正在开发一款简易价格爬虫:从电子表格读取ASIN列表,查询亚马逊对应商品价格后写入表格下一列,新旧价格不匹配时用颜色标记。

当前代码能运行,但99%的情况无法获取price_element,仅偶尔成功,成功时可正常写入价格并标记颜色。代码里保留了print(price_element)用于排查。

附代码如下:

import requests
import openpyxl
import time
from datetime import datetime
from bs4 import BeautifulSoup

# User agent for the requests
user_agent = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/116.0.0.0 Safari/537.36"

# Load the Excel file
excel_file = 'products_demo.xlsx'
workbook = openpyxl.load_workbook(excel_file)
sheet = workbook.active

# Get the number of rows in the sheet
num_rows = sheet.max_row

# Set the first row of column C with the current date and time
sheet['C1'] = 'Price as of ' + datetime.now().strftime('%Y-%m-%d %H:%M:%S')

# Loop through ASINs and scrape prices
for row in range(2, num_rows + 1):
    asin = sheet.cell(row=row, column=1).value
    url = f"https://www.amazon.com/dp/{asin}"
    
    headers = {
        'User-Agent': user_agent
    }
    
    response = requests.get(url, headers=headers)
    print(response)
    
    # Print response code in column D
    sheet.cell(row=row, column=4, value=response.status_code)
    
    if response.status_code == 200:
        # Find the <span> with class "a-offscreen"
        soup = BeautifulSoup(response.content, 'html.parser')
        price_element = soup.find('span', {'class': 'a-offscreen'})
        print(price_element)
        
        if price_element:
            current_price = price_element.get_text().strip()
            sheet.cell(row=row, column=3, value=current_price)
            
            # Compare with the previous price in column B and highlight if different
            previous_price = sheet.cell(row=row, column=2).value
            if previous_price != current_price:
                sheet.cell(row=row, column=3).fill = openpyxl.styles.PatternFill(start_color='FFFF00', end_color='FFFF00', fill_type='solid')
    else:
        print(f"Error: Unable to fetch data for ASIN {asin}")
    
    # Save the changes and wait for 5 seconds before the next request
    workbook.save('products_demo_updated.xlsx')
    time.sleep(5)

print("Scraping complete.")
解决方案

1. 强化请求头,规避亚马逊反爬检测

亚马逊会通过请求头识别非浏览器请求,仅用User-Agent不够,需要补充更多真实浏览器的请求头字段:

headers = {
    'User-Agent': user_agent,
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8',
    'Accept-Language': 'en-US,en;q=0.5',
    'Referer': 'https://www.amazon.com/',
    'Upgrade-Insecure-Requests': '1',
    'Connection': 'keep-alive'
}

2. 优化价格元素选择器

页面中可能存在多个a-offscreen类的span,直接用soup.find()可能拿到非目标元素,或者亚马逊页面结构变更,建议用更精准的选择器定位商品价格:

# 尝试定位商品主价格的父容器,再找a-offscreen
price_container = soup.find('div', {'id': 'corePriceDisplay_desktop_feature_div'})
if price_container:
    price_element = price_container.find('span', {'class': 'a-offscreen'})
# 备选:如果上面的容器找不到,用另一个常见价格id
else:
    price_element = soup.find('span', {'id': 'priceblock_ourprice'}) or soup.find('span', {'id': 'priceblock_dealprice'})

3. 处理动态加载/反爬跳转

如果请求返回200但页面是亚马逊的反验证页面(比如需要人机验证),requests无法处理动态验证,此时可以:

  • 加入随机延迟(比如3-10秒,不要固定5秒)
  • 使用代理IP轮换
  • 改用selenium模拟真实浏览器操作(适合新手快速绕开验证,但需要安装浏览器驱动)

4. 调试建议

在解析前打印响应文本的前1000字符,确认是否拿到了真实商品页面:

print(response.text[:1000])

如果出现验证码或重定向内容,说明被反爬命中,需要调整请求策略。

修改后的完整代码示例

import requests
import openpyxl
import time
import random
from datetime import datetime
from bs4 import BeautifulSoup

# User agent for the requests
user_agent = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/116.0.0.0 Safari/537.36"

# Load the Excel file
excel_file = 'products_demo.xlsx'
workbook = openpyxl.load_workbook(excel_file)
sheet = workbook.active

# Get the number of rows in the sheet
num_rows = sheet.max_row

# Set the first row of column C with the current date and time
sheet['C1'] = 'Price as of ' + datetime.now().strftime('%Y-%m-%d %H:%M:%S')

# Loop through ASINs and scrape prices
for row in range(2, num_rows + 1):
    asin = sheet.cell(row=row, column=1).value
    url = f"https://www.amazon.com/dp/{asin}"
    
    headers = {
        'User-Agent': user_agent,
        'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8',
        'Accept-Language': 'en-US,en;q=0.5',
        'Referer': 'https://www.amazon.com/',
        'Upgrade-Insecure-Requests': '1',
        'Connection': 'keep-alive'
    }
    
    response = requests.get(url, headers=headers)
    print(f"ASIN {asin} - 响应状态码: {response.status_code}")
    
    # Print response code in column D
    sheet.cell(row=row, column=4, value=response.status_code)
    
    if response.status_code == 200:
        # 调试用:打印页面开头确认是否为商品页
        # print(response.text[:1000])
        
        soup = BeautifulSoup(response.content, 'html.parser')
        # 精准定位价格元素
        price_container = soup.find('div', {'id': 'corePriceDisplay_desktop_feature_div'})
        if price_container:
            price_element = price_container.find('span', {'class': 'a-offscreen'})
        else:
            # 备选选择器
            price_element = soup.find('span', {'id': 'priceblock_ourprice'}) or soup.find('span', {'id': 'priceblock_dealprice'})
        
        print(f"ASIN {asin} - 价格元素: {price_element}")
        
        if price_element:
            current_price = price_element.get_text().strip()
            sheet.cell(row=row, column=3, value=current_price)
            
            # 对比旧价格并标记颜色
            previous_price = sheet.cell(row=row, column=2).value
            if previous_price != current_price:
                sheet.cell(row=row, column=3).fill = openpyxl.styles.PatternFill(start_color='FFFF00', end_color='FFFF00', fill_type='solid')
        else:
            print(f"ASIN {asin} - 未找到价格元素")
    else:
        print(f"Error: Unable to fetch data for ASIN {asin}")
    
    # 保存并设置随机延迟
    workbook.save('products_demo_updated.xlsx')
    time.sleep(random.randint(3, 10))

print("Scraping complete.")

内容的提问来源于stack exchange,提问作者boogaboogapickle

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 04:17:23