You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Beautiful Soup获取<small>标签产品名并排除价格数量项

提取标签中的产品名称(排除价格/数量标签)

方法1:利用标签结构筛选

观察HTML结构可知,包含产品名称的<small>标签内部带有<span class="pull-right">元素,而价格、数量相关的<small>标签没有该元素。可通过此特征精准定位目标标签:

from bs4 import BeautifulSoup

# 替换为你的HTML内容
html = """<tbody>
       <tr>
        <td class="no-border" colspan="2">
         <small>
          1:   EFURIX CREM 15 GR S VALEA
          <span class="pull-right">
          </span>
         </small>
         <small>
          1,00 x R$22,50
         </small>
         <td class="no-border text-right">
          <small>
           R$22,50
          </small>
         </td>
        </td>
       </tr>
       <tr>
        <td class="no-border" colspan="2">
         <small>
          2:   ASDRON XPE FR 100ML
          <span class="pull-right">
          </span>
         </small>
         <small>
          1,00 x R$50,32
         </small>
         <td class="no-border text-right">
          <small>
           R$50,32
          </small>
         </td>
        </td>
       </tr>
       <tr>
        <td class="no-border" colspan="2">
         <small>
          3:   DIAD  0,75MGC/ 2 COMP
          <span class="pull-right">
          </span>
         </small>
         <small>
          1,00 x R$5,00
         </small>
         <td class="no-border text-right">
          <small>
           R$5,00
          </small>
         </td>
        </td>
       </tr>
      </tbody>"""

soup = BeautifulSoup(html, 'html.parser')

# 选择包含指定span的small标签
product_tags = soup.select('small:has(span.pull-right)')

# 提取并清理产品名称
product_names = []
for tag in product_tags:
    raw_text = tag.get_text(strip=True)
    # 去除开头的序号和冒号,保留纯产品名
    clean_name = raw_text.split(':', 1)[1].strip()
    product_names.append(clean_name)

print(product_names)
# 输出: ['EFURIX CREM 15 GR S VALEA', 'ASDRON XPE FR 100ML', 'DIAD  0,75MGC/ 2 COMP']

方法2:通过文本内容过滤

如果标签结构存在变动风险,可通过文本特征排除价格、数量标签(包含R$或数字,数字 x格式的文本):

from bs4 import BeautifulSoup
import re

# 替换为你的HTML内容
html = """[你的HTML内容]"""
soup = BeautifulSoup(html, 'html.parser')

product_names = []
for small_tag in soup.find_all('small'):
    text = small_tag.get_text(strip=True)
    # 排除含价格标识或数量价格组合的文本
    if not ('R$' in text or re.match(r'\d+,\d+ x', text)):
        clean_name = text.split(':', 1)[1].strip()
        product_names.append(clean_name)

print(product_names)
# 输出与方法1一致

内容的提问来源于stack exchange,提问作者Rafael R Souza

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 22:50:07