如何用Python提取含sup标签的指定span标签后的文本?
解决HTML结构更新后提取Price Target价格的问题
问题分析
之前的HTML结构中,Price Target对应的<span>标签直接包含价格文本:
<span>$167.00</span>
网站更新后,<span>内新增了<sup>(2)</sup>子标签,结构变为:
<span>$167.00<sup>(2)</sup></span>
如果原代码依赖.string(当标签有多个子节点时会返回None)或仅提取第一个子节点的文本逻辑,就会失效。
解决方案
以下是几种可行的修改方案,基于BeautifulSoup实现:
方案1:获取全部文本后切割清理
直接提取<span>的所有文本内容,再通过切割去掉标注部分,适合标注格式固定的场景:
from bs4 import BeautifulSoup # 模拟更新后的HTML html = '<span>$167.00<sup>(2)</sup></span>' soup = BeautifulSoup(html, 'html.parser') # 获取span标签的全部文本并去除空白 full_text = soup.find('span').get_text(strip=True) # 按左括号分割,取第一部分即为价格 price = full_text.split('(')[0].strip() print(price) # 输出: $167.00
方案2:精准定位价格文本节点
遍历<span>的子节点,直接提取第一个纯文本节点(价格部分),适合结构固定、价格始终是第一个子内容的场景:
from bs4 import BeautifulSoup html = '<span>$167.00<sup>(2)</sup></span>' soup = BeautifulSoup(html, 'html.parser') span_tag = soup.find('span') # 筛选出第一个纯文本子节点 price_text = next(child for child in span_tag.children if isinstance(child, str)) price = price_text.strip() print(price) # 输出: $167.00
方案3:正则匹配价格格式
通过正则表达式匹配价格的固定格式($+数字+小数点后两位),鲁棒性最强,不受标注内容变化影响:
import re from bs4 import BeautifulSoup html = '<span>$167.00<sup>(2)</sup></span>' soup = BeautifulSoup(html, 'html.parser') full_text = soup.find('span').get_text(strip=True) # 匹配价格格式 price_match = re.search(r'\$\d+\.\d{2}', full_text) if price_match: price = price_match.group() print(price) # 输出: $167.00
内容的提问来源于stack exchange,提问作者Newbie
相关产品推荐
相关产品推荐

