如何用Beautiful Soup获取<small>标签产品名并排除价格数量项
提取标签中的产品名称(排除价格/数量标签)
方法1:利用标签结构筛选
观察HTML结构可知,包含产品名称的<small>标签内部带有<span class="pull-right">元素,而价格、数量相关的<small>标签没有该元素。可通过此特征精准定位目标标签:
from bs4 import BeautifulSoup # 替换为你的HTML内容 html = """<tbody> <tr> <td class="no-border" colspan="2"> <small> 1: EFURIX CREM 15 GR S VALEA <span class="pull-right"> </span> </small> <small> 1,00 x R$22,50 </small> <td class="no-border text-right"> <small> R$22,50 </small> </td> </td> </tr> <tr> <td class="no-border" colspan="2"> <small> 2: ASDRON XPE FR 100ML <span class="pull-right"> </span> </small> <small> 1,00 x R$50,32 </small> <td class="no-border text-right"> <small> R$50,32 </small> </td> </td> </tr> <tr> <td class="no-border" colspan="2"> <small> 3: DIAD 0,75MGC/ 2 COMP <span class="pull-right"> </span> </small> <small> 1,00 x R$5,00 </small> <td class="no-border text-right"> <small> R$5,00 </small> </td> </td> </tr> </tbody>""" soup = BeautifulSoup(html, 'html.parser') # 选择包含指定span的small标签 product_tags = soup.select('small:has(span.pull-right)') # 提取并清理产品名称 product_names = [] for tag in product_tags: raw_text = tag.get_text(strip=True) # 去除开头的序号和冒号,保留纯产品名 clean_name = raw_text.split(':', 1)[1].strip() product_names.append(clean_name) print(product_names) # 输出: ['EFURIX CREM 15 GR S VALEA', 'ASDRON XPE FR 100ML', 'DIAD 0,75MGC/ 2 COMP']
方法2:通过文本内容过滤
如果标签结构存在变动风险,可通过文本特征排除价格、数量标签(包含R$或数字,数字 x格式的文本):
from bs4 import BeautifulSoup import re # 替换为你的HTML内容 html = """[你的HTML内容]""" soup = BeautifulSoup(html, 'html.parser') product_names = [] for small_tag in soup.find_all('small'): text = small_tag.get_text(strip=True) # 排除含价格标识或数量价格组合的文本 if not ('R$' in text or re.match(r'\d+,\d+ x', text)): clean_name = text.split(':', 1)[1].strip() product_names.append(clean_name) print(product_names) # 输出与方法1一致
内容的提问来源于stack exchange,提问作者Rafael R Souza
相关产品推荐
相关产品推荐

