如何用Python BeautifulSoup从指定HTML的span标签中提取数据?
使用BeautifulSoup提取products区块中的价格数据
我来帮你一步步拆解怎么用Python的BeautifulSoup库从这段HTML的products区块里提取价格数据,步骤很清晰:
第一步:准备工作
首先确保你已经安装了BeautifulSoup和(如果需要从网页获取HTML的话)requests库,用pip安装:
pip install beautifulsoup4 requests
第二步:解析HTML
先把你提供的HTML内容存成字符串(如果是从网页爬的,就用requests.get获取响应文本),然后用BeautifulSoup解析:
from bs4 import BeautifulSoup # 你的HTML内容 html_content = '''<section class = "products"> <span class="price-box ri"> <span class="price "> <span data-currency-iso="PKR">Rs.</span> <span dir="ltr" data-price="5999"> 5,999</span> </span> <span class="price -old "> <span data-currency-iso="PKR">Rs.</span> <span dir="ltr" data-price="9999"> 9,999</span> </span> </span> </section>''' # 解析HTML soup = BeautifulSoup(html_content, 'html.parser')
第三步:定位products区块
先找到目标的<section>标签,通过它的class属性定位:
products_section = soup.find('section', class_='products')
这里用class_而不是class,因为class是Python的关键字,避免冲突。
第四步:提取价格数据
接下来我们可以提取当前价格和原价,有两种常用方式:提取文本内容,或者直接获取data-price属性(更精准,因为文本可能有空格或格式符)。
方式一:提取文本内容
# 获取当前价格的span current_price_span = products_section.find('span', class_='price') # 提取价格文本(去掉多余空格) current_price = current_price_span.get_text(strip=True) # 输出:Rs.5,999 # 获取原价的span old_price_span = products_section.find('span', class_='price -old') # 提取原价文本 old_price = old_price_span.get_text(strip=True) # 输出:Rs.9,999
方式二:提取data-price属性(更推荐)
这种方式直接拿到数字值,方便后续计算或存储:
# 获取当前价格的data-price值 current_price_num = current_price_span.find('span', {'data-price': True})['data-price'] # 输出:5999 # 获取原价的data-price值 old_price_num = old_price_span.find('span', {'data-price': True})['data-price'] # 输出:9999
完整示例代码
把上面的步骤整合起来:
from bs4 import BeautifulSoup html_content = '''<section class = "products"> <span class="price-box ri"> <span class="price "> <span data-currency-iso="PKR">Rs.</span> <span dir="ltr" data-price="5999"> 5,999</span> </span> <span class="price -old "> <span data-currency-iso="PKR">Rs.</span> <span dir="ltr" data-price="9999"> 9,999</span> </span> </span> </section>''' soup = BeautifulSoup(html_content, 'html.parser') products_section = soup.find('section', class_='products') # 提取带格式的价格文本 current_price_text = products_section.find('span', class_='price').get_text(strip=True) old_price_text = products_section.find('span', class_='price -old').get_text(strip=True) # 提取纯数字价格 current_price_num = products_section.find('span', class_='price').find('span', {'data-price': True})['data-price'] old_price_num = products_section.find('span', class_='price -old').find('span', {'data-price': True})['data-price'] print(f"当前价格(文本):{current_price_text}") print(f"原价(文本):{old_price_text}") print(f"当前价格(数字):{current_price_num}") print(f"原价(数字):{old_price_num}")
运行这段代码后,你就能得到两种格式的价格数据啦,根据你的需求选择就行~
内容的提问来源于stack exchange,提问作者Aftab
相关产品推荐
相关产品推荐

