如何用Beautiful Soup获取亚马逊首个搜索结果的商品标题与价格?
解决亚马逊搜索结果首个商品名称与价格爬取问题
你的问题主要出在三个地方:亚马逊的反爬机制拦截了请求、选择器不够精准、价格处理逻辑存在漏洞。下面是具体的修复方案:
问题分析
- 反爬拦截:直接用
requests.get发送请求时,没有携带浏览器标识的请求头,亚马逊会返回验证页面而非真实搜索结果,导致无法找到目标元素。 - 选择器不精准:
soup.find(class_='a-price')可能匹配到页面其他非商品的价格元素,或者因商品容器的层级关系,无法定位到首个搜索结果的价格。 - 价格处理逻辑缺陷:直接截取
text[1:]可能遇到换行、空格等多余字符,导致无法正常转成浮点数。
修复后的代码
import requests from bs4 import BeautifulSoup def submit_search(): product_name = product_entry.get() # 添加请求头模拟浏览器访问 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } # 发送带请求头的请求 product_page = requests.get(f"https://www.amazon.com/s?k={product_name}", headers=headers) soup = BeautifulSoup(product_page.content, 'html.parser') # 定位首个搜索结果的商品容器 first_product = soup.find('div', {'data-component-type': 's-search-result'}) if not first_product: print("未找到搜索结果或被反爬拦截") return # 获取商品名称 product_title = first_product.find('span', class_='a-size-medium a-color-base a-text-normal') title = product_title.text.strip() if product_title else "无法获取商品名称" # 获取商品价格 price_whole = first_product.find('span', class_='a-price-whole') price_fraction = first_product.find('span', class_='a-price-fraction') if price_whole and price_fraction: price = float(f"{price_whole.text.strip()}.{price_fraction.text.strip()}") print(f"首个商品名称:{title}\n价格:${price}") else: print(f"首个商品名称:{title}\n无法获取商品价格")
关键修复点说明
- 请求头模拟:通过
User-Agent伪装成浏览器,避免被亚马逊的反爬机制直接拦截。 - 精准定位商品容器:利用
data-component-type属性定位首个搜索结果的容器,再在容器内部查找名称和价格,避免匹配到无关元素。 - 价格拆分处理:分别获取价格的整数部分和小数部分,拼接后转成浮点数,避免直接处理文本时的格式问题。
内容的提问来源于stack exchange,提问作者Amogh_Spam
相关产品推荐
相关产品推荐

