You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python爬取Amazon商品时元素未找到的问题求助

无法爬取Amazon商品元素,所有信息返回None的排查方案

运行输出

未找到图片元素
未找到标题元素。
未找到价格元素。
未找到尺寸元素。
未找到销售方元素。
Item 1 is None.
Item 2 is None.
Item 3 is None.
Item 4 is None.
Item 5 is None.
已保存到urunler.pptx

爬取代码

import requests
from bs4 import BeautifulSoup
from pptx import Presentation
from pptx.util import Inches, Pt

# 通过URL爬取数据
def get_product_info(url):
    headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/88.0.4324.190 Safari/537.36"}
    r = requests.get(url, headers=headers)
    soup = BeautifulSoup(r.content, "html.parser")

    # 商品信息
    image_element = soup.find("img", {"id": "landingImage"})
    if image_element is not None:
        image = image_element["src"]
    else:
        print("未找到图片元素。")
        image = None

    title_element = soup.find("span", {"id": "productTitle"})
    if title_element is not None:
        title = title_element.string.strip()
    else:
        print("未找到标题元素。")
        title = None

    price_element = soup.find("span", {"id": "priceblock_ourprice"})
    if price_element is not None:
        price = price_element.string
    else:
        print("未找到价格元素。")
        price = None

    dimensions_element = soup.find("span", {"class": "a-list-item"})  # 这通常是一个合理的猜测。
    if dimensions_element is not None:
        dimensions = dimensions_element.string
    else:
        print("未找到尺寸元素。")
        dimensions = None

    sold_by_element = soup.find("a", {"id": "sellerProfileTriggerId"})
    if sold_by_element is not None:
        sold_by = sold_by_element.string
    else:
        print("未找到销售方元素。")
        sold_by = None

    # 返回信息
    return image, title, price, dimensions, sold_by

# 将信息保存到PPTX
def save_to_pptx(info, pptx_filename):
    prs = Presentation()
    slide_layout = prs.slide_layouts[0]

    # 将信息添加到幻灯片
    for i, item in enumerate(info):
        if item is not None:
            slide = prs.slides.add_slide(slide_layout)
            left = top = Inches(i)
            txBox = slide.shapes.add_textbox(left, top, Inches(6), Inches(1))
            tf = txBox.text_frame
            p = tf.add_paragraph()
            p.text = str(item)
        else:
            print(f"Item {i+1} is None.")

    # 保存文件
    try:
        prs.save(pptx_filename)
        print(f"已保存到 {pptx_filename}")
    except Exception as e:
        print(f"保存PPTX失败: {e}")

# 请求URL
url = "https://www.amazon.com/dp/B08NDJZ2JZ"  # 用目标URL替换此URL

# 保存商品详情到PPTX
info = get_product_info(url)
save_to_pptx(info, "urunler.pptx")

问题排查与解决

1. 核心原因

反爬机制拦截

Amazon会对非浏览器请求做严格检测,仅设置User-Agent不足以绕过拦截,可能返回无真实商品数据的伪装页面,导致BeautifulSoup无法匹配到目标元素。

元素选择器过时

Amazon页面结构频繁更新,你使用的选择器已失效:

  • priceblock_ourprice 这类ID已被替换,当前价格通常使用.a-price .a-offscreen等选择器
  • a-list-item 过于宽泛,无法精准定位尺寸信息
  • sellerProfileTriggerId 仅在特定场景存在,第三方卖家信息的位置已变更

动态渲染限制

部分商品内容由JavaScript动态加载,requests只能获取静态HTML,无法获取JS渲染后的真实数据。

2. 解决步骤

验证请求内容

先打印返回页面的前1000字符,确认是否拿到真实商品页面:

r = requests.get(url, headers=headers)
print(r.text[:1000])

如果输出登录提示、验证码或无关内容,说明被反爬拦截。

完善请求头

添加更多浏览器常用请求头,提升请求合法性:

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
    "Accept-Language": "en-US,en;q=0.9",
    "Accept-Encoding": "gzip, deflate, br",
    "Referer": "https://www.amazon.com/",
    "DNT": "1"
}

若仍被拦截,可从浏览器开发者工具复制有效Cookie添加到请求头(注意Cookie有有效期)。

更新元素选择器

针对当前Amazon页面结构调整选择器:

  • 标题:soup.find("span", {"id": "productTitle"}).get_text(strip=True)
  • 价格:soup.find("span", class_="a-offscreen").get_text(strip=True)
  • 尺寸:先定位详情区域再精准匹配:
    details_div = soup.find("div", id="productDetails_feature_div")
    if details_div:
        dimensions = details_div.find("span", string=lambda t: t and "Product Dimensions" in t).find_next("span").get_text(strip=True)
    
  • 销售方:soup.find("div", id="merchant-info").get_text(strip=True)

处理动态渲染

若静态请求无法获取数据,改用selenium模拟浏览器加载:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options

options = Options()
options.add_argument("--headless=new")
driver = webdriver.Chrome(options=options)
driver.get(url)

title = driver.find_element(By.ID, "productTitle").text.strip()
price = driver.find_element(By.CSS_SELECTOR, ".a-price .a-offscreen").text
driver.quit()

合规提醒

Amazon的Robots协议禁止未经授权的商品数据爬取,建议使用官方Amazon Product Advertising API获取合规数据。

内容的提问来源于stack exchange,提问作者Emre Sural

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 16:25:32