如何从含HTML span元素的列表中提取并转换价格数据?
从HTML列表中提取价格的解决方案
原始输入的HTML列表:
[<span class="primary-price">$61,989</span>, <span class="primary-price">$21,905</span>, <span class="primary-price">$20,595</span>]
方案1:提取为数字格式(如 [61.989, 21.905, 20.595])
使用Python的BeautifulSoup解析HTML,提取价格后转换为目标数字格式:
from bs4 import BeautifulSoup # 原始HTML内容 html = '''[<span class="primary-price">$61,989</span>, <span class="primary-price">$21,905</span>, <span class="primary-price">$20,595</span>]''' # 解析并提取价格文本 soup = BeautifulSoup(html, 'html.parser') price_strings = [span.text for span in soup.find_all('span', class_='primary-price')] # 转换为带小数点的数字格式(替换逗号为小数点,去掉$符号) prices = [float(price.replace('$', '').replace(',', '.')) for price in price_strings] print(f"prices = {prices}")
执行结果:
prices = [61.989, 21.905, 20.595]
方案2:保留原始价格格式(如 [$61,989, $21,905, $20,595])
如果只需要提取原始格式的价格字符串,直接提取文本即可:
from bs4 import BeautifulSoup html = '''[<span class="primary-price">$61,989</span>, <span class="primary-price">$21,905</span>, <span class="primary-price">$20,595</span>]''' soup = BeautifulSoup(html, 'html.parser') prices = [span.text for span in soup.find_all('span', class_='primary-price')] print(f"prices = {prices}")
执行结果:
prices = ['$61,989', '$21,905', '$20,595']
内容的提问来源于stack exchange,提问作者deagle
相关产品推荐
相关产品推荐

