You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup提取HTML同类名下折扣价格及爬虫代码优化

网页解析价格提取问题解答

1. 同类名标签下折扣价提取方法

你当前只提取了带删除线的<s>标签包裹的原价,折扣价和原价虽然类名相同,但标签类型不同:折扣价是普通<span>标签包裹,没有删除线样式。你可以通过标签类型区分两者,提取逻辑如下:

  • 先检查商品容器中是否存在类为product-card__regular-price的<s>标签:存在说明该商品有折扣,<s>标签内容为原价,同容器内类为product-card__regular-price的<span>标签内容就是折扣价
  • 若不存在<s>标签,再检查是否有同类名的<span>标签:存在说明是正价商品,内容即为售价
  • 若两种标签都不存在,说明商品售罄

2. 代码优化方案与实现思路

现有代码可优化点

  • 查找单个元素时不需要用findAll返回列表后再取下标0,直接用find方法即可,代码更简洁
  • 不要使用裸except捕获所有异常,会掩盖代码本身的错误,建议明确捕获索引错误等目标异常
  • urllib库操作繁琐,新手更推荐使用requests库发起请求,语法更简单,也更容易配置请求头避免被网站反爬拦截
  • 现有价格提取逻辑不完善,仅覆盖了原价和售罄两种场景,缺少折扣价、正价商品的判断

优化后参考代码

import requests
from bs4 import BeautifulSoup

# 配置请求头,模拟浏览器访问避免被拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}
base_url = "https://www.parkoutlet.com.ph/collections/mens-footwear?page={}"

# 可修改页数范围爬取多页
for page in range(1, 3):
    url = base_url.format(page)
    html = requests.get(url, headers=headers).text
    soup = BeautifulSoup(html, "html.parser")
    item_containers = soup.find_all("div", class_="grid__item small--one-half medium-up--one-fifth")
    
    for container in item_containers:
        # 提取品牌和商品名,用strip()去除多余空白字符
        brand = container.find("div", class_="product-card__brand").text.strip()
        name = container.find("div", class_="product-card__name").text.strip()
        
        # 价格提取逻辑
        original_price = container.find("s", class_="product-card__regular-price")
        sale_price = container.find("span", class_="product-card__regular-price")
        sold_out = container.find("div", class_="product-card__availability")
        
        print(f"品牌:{brand}")
        print(f"商品名:{name}")
        if sold_out:
            print("状态:SOLD OUT")
        elif original_price:
            print(f"原价:{original_price.text.strip()}")
            print(f"折扣价:{sale_price.text.strip()}")
        else:
            print(f"售价:{sale_price.text.strip()}")
        print("-"*30)

进阶实现思路

如果需要长期爬取该网站,可以进一步优化:

  • 可以把提取到的数据存入CSV、Excel或者数据库,方便后续分析
  • 增加异常重试逻辑,遇到网络请求失败时自动重试几次
  • 增加爬取间隔,避免请求频率太高被网站封禁IP

内容的提问来源于stack exchange,提问作者Raphael Ramirez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 03:18:02