网页抓取时同一商品动态标签问题及脚本修复求助
问题分析与修正方案
原脚本核心问题是折扣判断逻辑完全颠倒,且未处理元素不存在的异常情况,导致无法正确提取商品信息。以下是具体问题点和修正后的解决方案:
原脚本的问题
- 折扣判断逻辑错误:把无折扣的价格类
atg_store_newPrice当成了有折扣标识,反而将有折扣的price_discount放到了无折扣分支 - 原价提取错误:有折扣时错误查找了和当前价格相同的类,完全没定位到正确的原价标签
- 未处理空元素:当找不到价格或描述元素时,直接调用
.text会触发AttributeError - 缺少请求头:部分网站会拦截无标识的爬虫请求,导致返回空内容或错误状态码
修正后的代码
import requests from bs4 import BeautifulSoup def scrape_product_info(product_element): # 初始化默认值,避免未找到元素时报错 current_price = None previous_price = None description = None # 优先检查是否有折扣(通过折扣价格标签判断) discount_price_elem = product_element.find('span', class_='price_discount') if discount_price_elem: # 有折扣的情况 current_price = discount_price_elem.text.strip() # 提取原价 old_price_elem = product_element.find('span', class_='atg_store_oldPrice price_regular') if old_price_elem: previous_price = old_price_elem.text.strip() # 提取折扣状态下的描述 desc_elem = product_element.find('div', class_='descrip_truncated') if desc_elem: description = desc_elem.text.strip() else: # 无折扣的情况 regular_price_elem = product_element.find('span', class_='atg_store_newPrice') if regular_price_elem: current_price = regular_price_elem.text.strip() # 提取无折扣状态下的描述 desc_elem = product_element.find('div', class_='descrip_full') if desc_elem: description = desc_elem.text.strip() # 打印提取结果(可替换为存储逻辑) print("商品描述:", description if description else "未获取到") print("当前价格:", current_price if current_price else "未获取到") print("原价:", previous_price if previous_price else "无折扣") print("=" * 50) # 目标URL url = 'https://www.cotodigital3.com.ar/sitios/cdigi/browse/catalogo-textil-indumentaria-ni%C3%B1os-remeras/_/N-1vjzjgf' # 添加请求头模拟浏览器访问,避免被反爬拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } # 发送请求 response = requests.get(url, headers=headers) if response.status_code == 200: soup = BeautifulSoup(response.text, 'html.parser') # 定位商品元素(使用更精准的选择器,避免抓取非商品项) product_elements = soup.find_all('li', class_='clearfix product') if not product_elements: print("未找到商品元素,请检查页面结构是否变化") else: for product in product_elements: scrape_product_info(product) else: print(f'请求失败,状态码: {response.status_code}')
关键优化点
- 调整折扣判断逻辑:先检查是否存在折扣价格标签,再区分两种状态
- 增加空值处理:每个元素提取前先判断是否存在,避免报错
- 添加请求头:模拟浏览器行为,降低被反爬拦截的概率
- 优化商品选择器:使用更精准的类名组合,避免抓取到非商品元素
内容的提问来源于stack exchange,提问作者znah
相关产品推荐
相关产品推荐

