Python亚马逊价格爬虫可靠性问题求助:多数无法获取价格元素
问题描述
我是Python新手,正在开发一款简易价格爬虫:从电子表格读取ASIN列表,查询亚马逊对应商品价格后写入表格下一列,新旧价格不匹配时用颜色标记。
当前代码能运行,但99%的情况无法获取price_element,仅偶尔成功,成功时可正常写入价格并标记颜色。代码里保留了print(price_element)用于排查。
附代码如下:
import requests import openpyxl import time from datetime import datetime from bs4 import BeautifulSoup # User agent for the requests user_agent = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/116.0.0.0 Safari/537.36" # Load the Excel file excel_file = 'products_demo.xlsx' workbook = openpyxl.load_workbook(excel_file) sheet = workbook.active # Get the number of rows in the sheet num_rows = sheet.max_row # Set the first row of column C with the current date and time sheet['C1'] = 'Price as of ' + datetime.now().strftime('%Y-%m-%d %H:%M:%S') # Loop through ASINs and scrape prices for row in range(2, num_rows + 1): asin = sheet.cell(row=row, column=1).value url = f"https://www.amazon.com/dp/{asin}" headers = { 'User-Agent': user_agent } response = requests.get(url, headers=headers) print(response) # Print response code in column D sheet.cell(row=row, column=4, value=response.status_code) if response.status_code == 200: # Find the <span> with class "a-offscreen" soup = BeautifulSoup(response.content, 'html.parser') price_element = soup.find('span', {'class': 'a-offscreen'}) print(price_element) if price_element: current_price = price_element.get_text().strip() sheet.cell(row=row, column=3, value=current_price) # Compare with the previous price in column B and highlight if different previous_price = sheet.cell(row=row, column=2).value if previous_price != current_price: sheet.cell(row=row, column=3).fill = openpyxl.styles.PatternFill(start_color='FFFF00', end_color='FFFF00', fill_type='solid') else: print(f"Error: Unable to fetch data for ASIN {asin}") # Save the changes and wait for 5 seconds before the next request workbook.save('products_demo_updated.xlsx') time.sleep(5) print("Scraping complete.")
解决方案
1. 强化请求头,规避亚马逊反爬检测
亚马逊会通过请求头识别非浏览器请求,仅用User-Agent不够,需要补充更多真实浏览器的请求头字段:
headers = { 'User-Agent': user_agent, 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Referer': 'https://www.amazon.com/', 'Upgrade-Insecure-Requests': '1', 'Connection': 'keep-alive' }
2. 优化价格元素选择器
页面中可能存在多个a-offscreen类的span,直接用soup.find()可能拿到非目标元素,或者亚马逊页面结构变更,建议用更精准的选择器定位商品价格:
# 尝试定位商品主价格的父容器,再找a-offscreen price_container = soup.find('div', {'id': 'corePriceDisplay_desktop_feature_div'}) if price_container: price_element = price_container.find('span', {'class': 'a-offscreen'}) # 备选:如果上面的容器找不到,用另一个常见价格id else: price_element = soup.find('span', {'id': 'priceblock_ourprice'}) or soup.find('span', {'id': 'priceblock_dealprice'})
3. 处理动态加载/反爬跳转
如果请求返回200但页面是亚马逊的反验证页面(比如需要人机验证),requests无法处理动态验证,此时可以:
- 加入随机延迟(比如3-10秒,不要固定5秒)
- 使用代理IP轮换
- 改用
selenium模拟真实浏览器操作(适合新手快速绕开验证,但需要安装浏览器驱动)
4. 调试建议
在解析前打印响应文本的前1000字符,确认是否拿到了真实商品页面:
print(response.text[:1000])
如果出现验证码或重定向内容,说明被反爬命中,需要调整请求策略。
修改后的完整代码示例
import requests import openpyxl import time import random from datetime import datetime from bs4 import BeautifulSoup # User agent for the requests user_agent = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/116.0.0.0 Safari/537.36" # Load the Excel file excel_file = 'products_demo.xlsx' workbook = openpyxl.load_workbook(excel_file) sheet = workbook.active # Get the number of rows in the sheet num_rows = sheet.max_row # Set the first row of column C with the current date and time sheet['C1'] = 'Price as of ' + datetime.now().strftime('%Y-%m-%d %H:%M:%S') # Loop through ASINs and scrape prices for row in range(2, num_rows + 1): asin = sheet.cell(row=row, column=1).value url = f"https://www.amazon.com/dp/{asin}" headers = { 'User-Agent': user_agent, 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Referer': 'https://www.amazon.com/', 'Upgrade-Insecure-Requests': '1', 'Connection': 'keep-alive' } response = requests.get(url, headers=headers) print(f"ASIN {asin} - 响应状态码: {response.status_code}") # Print response code in column D sheet.cell(row=row, column=4, value=response.status_code) if response.status_code == 200: # 调试用:打印页面开头确认是否为商品页 # print(response.text[:1000]) soup = BeautifulSoup(response.content, 'html.parser') # 精准定位价格元素 price_container = soup.find('div', {'id': 'corePriceDisplay_desktop_feature_div'}) if price_container: price_element = price_container.find('span', {'class': 'a-offscreen'}) else: # 备选选择器 price_element = soup.find('span', {'id': 'priceblock_ourprice'}) or soup.find('span', {'id': 'priceblock_dealprice'}) print(f"ASIN {asin} - 价格元素: {price_element}") if price_element: current_price = price_element.get_text().strip() sheet.cell(row=row, column=3, value=current_price) # 对比旧价格并标记颜色 previous_price = sheet.cell(row=row, column=2).value if previous_price != current_price: sheet.cell(row=row, column=3).fill = openpyxl.styles.PatternFill(start_color='FFFF00', end_color='FFFF00', fill_type='solid') else: print(f"ASIN {asin} - 未找到价格元素") else: print(f"Error: Unable to fetch data for ASIN {asin}") # 保存并设置随机延迟 workbook.save('products_demo_updated.xlsx') time.sleep(random.randint(3, 10)) print("Scraping complete.")
内容的提问来源于stack exchange,提问作者boogaboogapickle
相关产品推荐
相关产品推荐

