You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python提取含sup标签的指定span标签后的文本?

解决HTML结构更新后提取Price Target价格的问题

问题分析

之前的HTML结构中,Price Target对应的<span>标签直接包含价格文本:

<span>$167.00</span>

网站更新后,<span>内新增了<sup>(2)</sup>子标签,结构变为:

<span>$167.00<sup>(2)</sup></span>

如果原代码依赖.string(当标签有多个子节点时会返回None)或仅提取第一个子节点的文本逻辑,就会失效。

解决方案

以下是几种可行的修改方案,基于BeautifulSoup实现:

方案1:获取全部文本后切割清理

直接提取<span>的所有文本内容,再通过切割去掉标注部分,适合标注格式固定的场景:

from bs4 import BeautifulSoup

# 模拟更新后的HTML
html = '<span>$167.00<sup>(2)</sup></span>'
soup = BeautifulSoup(html, 'html.parser')

# 获取span标签的全部文本并去除空白
full_text = soup.find('span').get_text(strip=True)
# 按左括号分割,取第一部分即为价格
price = full_text.split('(')[0].strip()
print(price)  # 输出: $167.00

方案2:精准定位价格文本节点

遍历<span>的子节点,直接提取第一个纯文本节点(价格部分),适合结构固定、价格始终是第一个子内容的场景:

from bs4 import BeautifulSoup

html = '<span>$167.00<sup>(2)</sup></span>'
soup = BeautifulSoup(html, 'html.parser')
span_tag = soup.find('span')

# 筛选出第一个纯文本子节点
price_text = next(child for child in span_tag.children if isinstance(child, str))
price = price_text.strip()
print(price)  # 输出: $167.00

方案3:正则匹配价格格式

通过正则表达式匹配价格的固定格式($+数字+小数点后两位),鲁棒性最强,不受标注内容变化影响:

import re
from bs4 import BeautifulSoup

html = '<span>$167.00<sup>(2)</sup></span>'
soup = BeautifulSoup(html, 'html.parser')
full_text = soup.find('span').get_text(strip=True)

# 匹配价格格式
price_match = re.search(r'\$\d+\.\d{2}', full_text)
if price_match:
    price = price_match.group()
    print(price)  # 输出: $167.00

内容的提问来源于stack exchange,提问作者Newbie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 15:38:30