如何在BeautifulSoup中匹配多个class抓取Steam游戏原价与折扣价
解决Steam畅销榜价格与标题对应抓取的方案
直接分开调用两次find_all获取原价、折扣价列表的方式会出现索引错位,因为两类价格的条目是混排在榜单中的,无法和你提前抓取的标题列表按顺序一一匹配。
正确实现思路
- 先定位每个游戏的独立根容器:Steam搜索结果页每个游戏条目都对应一个class为
search_result_row的div节点,遍历该节点列表,每个节点对应一个独立游戏,从根节点内分别抓取标题、价格,就能天然保证对应关系不会错乱。 - 单条目内匹配价格时合并两个class规则,不需要分开查询,用自定义class匹配规则即可同时匹配原价、折扣价两类节点。
示例代码
from bs4 import BeautifulSoup import requests # 抓取页面内容 url = "https://store.steampowered.com/search/?filter=topsellers&snr=1_241_4_action_12" resp = requests.get(url) soup = BeautifulSoup(resp.text, "html.parser") # 遍历所有游戏条目 game_list = [] for item in soup.find_all("div", class_="search_result_row"): # 抓取标题,可直接替换为你已实现的标题抓取逻辑 title = item.find("span", class_="title").text.strip() # 合并匹配两类价格节点:同时匹配带/不带discounted的价格div price_elem = item.find("div", class_=lambda c: c and all(key in c for key in ["search_price", "responsive_secondrow"])) # 处理价格文本,去除多余空白和换行 price = price_elem.text.strip() if price_elem else "无价格信息" game_list.append({"title": title, "price": price}) # 输出验证结果 for game in game_list: print(f"游戏:{game['title']},价格:{game['price']}")
可选优化:折扣价拆分
如果需要把折扣价中的原价、实付价分开提取,可以对拿到的price字符串做拆分处理,折扣价的文本默认是原价在前、实付价在后,用换行符分割即可:
if "\n" in price: origin_price, discount_price = [i.strip() for i in price.split("\n") if i.strip()]
内容的提问来源于stack exchange,提问作者Tranquil Oshan
相关产品推荐
相关产品推荐

