使用Python爬取Amazon商品时元素未找到的问题求助
无法爬取Amazon商品元素,所有信息返回None的排查方案
运行输出
未找到图片元素 未找到标题元素。 未找到价格元素。 未找到尺寸元素。 未找到销售方元素。 Item 1 is None. Item 2 is None. Item 3 is None. Item 4 is None. Item 5 is None. 已保存到urunler.pptx
爬取代码
import requests from bs4 import BeautifulSoup from pptx import Presentation from pptx.util import Inches, Pt # 通过URL爬取数据 def get_product_info(url): headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/88.0.4324.190 Safari/537.36"} r = requests.get(url, headers=headers) soup = BeautifulSoup(r.content, "html.parser") # 商品信息 image_element = soup.find("img", {"id": "landingImage"}) if image_element is not None: image = image_element["src"] else: print("未找到图片元素。") image = None title_element = soup.find("span", {"id": "productTitle"}) if title_element is not None: title = title_element.string.strip() else: print("未找到标题元素。") title = None price_element = soup.find("span", {"id": "priceblock_ourprice"}) if price_element is not None: price = price_element.string else: print("未找到价格元素。") price = None dimensions_element = soup.find("span", {"class": "a-list-item"}) # 这通常是一个合理的猜测。 if dimensions_element is not None: dimensions = dimensions_element.string else: print("未找到尺寸元素。") dimensions = None sold_by_element = soup.find("a", {"id": "sellerProfileTriggerId"}) if sold_by_element is not None: sold_by = sold_by_element.string else: print("未找到销售方元素。") sold_by = None # 返回信息 return image, title, price, dimensions, sold_by # 将信息保存到PPTX def save_to_pptx(info, pptx_filename): prs = Presentation() slide_layout = prs.slide_layouts[0] # 将信息添加到幻灯片 for i, item in enumerate(info): if item is not None: slide = prs.slides.add_slide(slide_layout) left = top = Inches(i) txBox = slide.shapes.add_textbox(left, top, Inches(6), Inches(1)) tf = txBox.text_frame p = tf.add_paragraph() p.text = str(item) else: print(f"Item {i+1} is None.") # 保存文件 try: prs.save(pptx_filename) print(f"已保存到 {pptx_filename}") except Exception as e: print(f"保存PPTX失败: {e}") # 请求URL url = "https://www.amazon.com/dp/B08NDJZ2JZ" # 用目标URL替换此URL # 保存商品详情到PPTX info = get_product_info(url) save_to_pptx(info, "urunler.pptx")
问题排查与解决
1. 核心原因
反爬机制拦截
Amazon会对非浏览器请求做严格检测,仅设置User-Agent不足以绕过拦截,可能返回无真实商品数据的伪装页面,导致BeautifulSoup无法匹配到目标元素。
元素选择器过时
Amazon页面结构频繁更新,你使用的选择器已失效:
priceblock_ourprice这类ID已被替换,当前价格通常使用.a-price .a-offscreen等选择器a-list-item过于宽泛,无法精准定位尺寸信息sellerProfileTriggerId仅在特定场景存在,第三方卖家信息的位置已变更
动态渲染限制
部分商品内容由JavaScript动态加载,requests只能获取静态HTML,无法获取JS渲染后的真实数据。
2. 解决步骤
验证请求内容
先打印返回页面的前1000字符,确认是否拿到真实商品页面:
r = requests.get(url, headers=headers) print(r.text[:1000])
如果输出登录提示、验证码或无关内容,说明被反爬拦截。
完善请求头
添加更多浏览器常用请求头,提升请求合法性:
headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Accept-Language": "en-US,en;q=0.9", "Accept-Encoding": "gzip, deflate, br", "Referer": "https://www.amazon.com/", "DNT": "1" }
若仍被拦截,可从浏览器开发者工具复制有效Cookie添加到请求头(注意Cookie有有效期)。
更新元素选择器
针对当前Amazon页面结构调整选择器:
- 标题:
soup.find("span", {"id": "productTitle"}).get_text(strip=True) - 价格:
soup.find("span", class_="a-offscreen").get_text(strip=True) - 尺寸:先定位详情区域再精准匹配:
details_div = soup.find("div", id="productDetails_feature_div") if details_div: dimensions = details_div.find("span", string=lambda t: t and "Product Dimensions" in t).find_next("span").get_text(strip=True) - 销售方:
soup.find("div", id="merchant-info").get_text(strip=True)
处理动态渲染
若静态请求无法获取数据,改用selenium模拟浏览器加载:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options options = Options() options.add_argument("--headless=new") driver = webdriver.Chrome(options=options) driver.get(url) title = driver.find_element(By.ID, "productTitle").text.strip() price = driver.find_element(By.CSS_SELECTOR, ".a-price .a-offscreen").text driver.quit()
合规提醒
Amazon的Robots协议禁止未经授权的商品数据爬取,建议使用官方Amazon Product Advertising API获取合规数据。
内容的提问来源于stack exchange,提问作者Emre Sural
相关产品推荐
相关产品推荐

