如何将Beautiful Soup提取的价格文本转换为浮点型(float)并生成浮点列表,解决循环外仅输出最后一个价格的问题
问题:将网页提取的价格文本转换为浮点型列表
我尝试从HTML网站提取文本创建一个浮点型(float)列表,以便检查列表中是否存在符合或低于我预算的价格。我使用Beautiful Soup 4的find_all()方法提取不带HTML标签的文本,代码如下:
import requests from bs4 import BeautifulSoup url = "https://starpets.gg/adopt-me/shop/pet/firefly/15292" result = requests.get(url) doc = BeautifulSoup(result.text, "html.parser") tag = doc.find("span", itemprop = "price") for price1 in tag: print("Best price", price1.text) print("-----") tag = doc.find_all("div", class_ = "_text_j98bt_1 _text__size_m_j98bt_40 _text__weight_bold_j98bt_83 _text__style_normal_j98bt_95 _text__decoration_normal_j98bt_104 _content-price_1gow4_85") for price2 in tag: print(price2.text)
运行结果:
Best price 1.36 ----- 1.35 $ 1.7 $ 3.25 $
由于该网站根据位置使用两种不同的标签来展示价格,因此我设置了两个独立的价格变量,其中类标签对应的价格数量占多数且会不断变化。但我发现,如果在for price2 in tag:循环外部单独打印price2.text,只会输出最后一个price2的内容。请问如何将提取的价格文本转换为浮点型并生成对应的列表,同时解决循环取值的问题?
解决方案
没问题,我来帮你搞定这个问题。核心思路是创建一个空列表,把所有提取到的价格处理后都添加进去,这样就能避免循环后只保留最后一个值的问题,同时完成文本到浮点型的转换。
1. 写一个通用的价格解析函数
首先我们需要一个小工具函数,用来处理各种格式的价格字符串——不管是带$符号还是纯数字的,都能提取出有效数字并转成float类型:
def parse_price(price_str): # 移除字符串里所有非数字和小数点的字符,比如$符号、空格 cleaned_price = ''.join(char for char in price_str if char.isdigit() or char == '.') # 转换为浮点型并返回 return float(cleaned_price)
2. 收集所有价格到列表中
接下来我们修改原代码,把两个来源的价格都收集到同一个列表里,而不是只打印出来:
import requests from bs4 import BeautifulSoup def parse_price(price_str): cleaned_price = ''.join(char for char in price_str if char.isdigit() or char == '.') return float(cleaned_price) url = "https://starpets.gg/adopt-me/shop/pet/firefly/15292" result = requests.get(url) doc = BeautifulSoup(result.text, "html.parser") # 创建空列表存储所有浮点型价格 price_list = [] # 处理第一个来源的「Best price」 best_price_tag = doc.find("span", itemprop="price") if best_price_tag: # 原代码里的for循环其实没必要,因为find返回的是单个标签,直接取文本即可 best_price_text = best_price_tag.get_text(strip=True) best_price = parse_price(best_price_text) price_list.append(best_price) print(f"已添加最优价格:{best_price}") print("-----") # 处理第二个来源的多个价格 price_tags = doc.find_all("div", class_="_text_j98bt_1 _text__size_m_j98bt_40 _text__weight_bold_j98bt_83 _text__style_normal_j98bt_95 _text__decoration_normal_j98bt_104 _content-price_1gow4_85") for tag in price_tags: price_text = tag.get_text(strip=True) price = parse_price(price_text) price_list.append(price) print(f"已添加价格:{price}") # 输出最终的浮点型价格列表 print("\n最终的价格列表:") print(price_list)
3. 检查预算(可选)
有了浮点型列表后,你可以轻松筛选出符合预算的价格,比如你的预算是2.0:
budget = 2.0 affordable_prices = [price for price in price_list if price <= budget] print(f"\n符合预算(≤{budget})的价格:") print(affordable_prices)
为什么原代码循环后只保留最后一个值?
你之前遇到的问题很常见:在for price2 in tag:循环里,price2每次都会被赋值为当前的标签对象,循环结束后它自然就指向最后一个元素。而用列表收集的方式,每处理一个价格就把它添加到列表里,就能保留所有提取到的价格啦。
内容的提问来源于stack exchange,提问作者Rosie
相关产品推荐
相关产品推荐

