如何使用Python BeautifulSoup提取HTML中的商品价格13990?
解决BeautifulSoup提取商品价格并转为数字的问题
你的问题出在用.string获取嵌套标签内的文本——因为crwActualPrice这个span里面还有多层子标签,.string只能获取当前标签的直接文本(这里是空的),没法拿到嵌套在里面的价格文本。咱们可以这样修改:
步骤1:正确提取价格文本
改用.get_text()来获取该标签下的所有文本内容,然后清理掉多余的空格和隐藏的符号:
from bs4 import BeautifulSoup # 假设你的HTML内容存在html变量里 html = '''<span class="crwActualPrice"> <span style="text-decoration: inherit; white-space: nowrap;"> <span class="currencyINR"> </span> <span class="currencyINRFallback" style="display:none"> Rs. </span> 13,990.00 </span> </span>''' soup = BeautifulSoup(html, 'html.parser') dprice = soup.find_all("span", class_="crwActualPrice") for each_price in dprice: # 获取所有文本并自动清理多余空格 money_str = each_price.get_text(strip=True) print(money_str) # 输出: 13,990.00
步骤2:将文本转为纯数字13990
接下来需要去掉逗号和小数点,再转换为整数:
# 接上面的代码 # 先去掉逗号,再切割小数点取整数部分 cleaned_price = money_str.replace(',', '').split('.')[0] # 转为整数类型 final_price = int(cleaned_price) print(final_price) # 输出: 13990
合并后的完整代码
from bs4 import BeautifulSoup html = '''<span class="crwActualPrice"> <span style="text-decoration: inherit; white-space: nowrap;"> <span class="currencyINR"> </span> <span class="currencyINRFallback" style="display:none"> Rs. </span> 13,990.00 </span> </span>''' soup = BeautifulSoup(html, 'html.parser') dprice = soup.find_all("span", class_="crwActualPrice") for each_price in dprice: money_str = each_price.get_text(strip=True) # 清理格式并转数字 final_price = int(money_str.replace(',', '').split('.')[0]) print(final_price)
简单说,核心就是用.get_text(strip=True)替代.string来获取嵌套标签的文本,再通过字符串处理去掉非数字部分,最后转为整数即可。
内容的提问来源于stack exchange,提问作者mohinuddeen
相关产品推荐
相关产品推荐

