如何用BeautifulSoup提取重复class下的Steam发行商字符串?
提取Steam游戏页面指定发行商名称的BeautifulSoup方案
问题背景
需要从Steam游戏页面的HTML中提取「FromSoftware, Inc.」和「Bandai Namco Entertainment」两个发行商名称,但页面中dev_row类的元素重复出现,原脚本用soup.find('div', class_='dev_row')只能获取第一个dev_row的内容,无法精准定位到发行商所在行。
解决方案
核心思路是通过「Publisher:」这个标识文本,精准定位到发行商对应的dev_row,再提取其中的发行商名称。以下提供两种可行实现:
方法一:通过子元素定位父容器
直接找到文本为「Publisher:」的副标题元素,再回溯到对应的dev_row父容器,最后提取所有<a>标签的文本:
from bs4 import BeautifulSoup # 假设soup已完成HTML解析 subtitle = soup.find('div', class_='subtitle column', string='Publisher:') if subtitle: publisher_row = subtitle.find_parent('div', class_='dev_row') if publisher_row: publishers = [a.text.strip() for a in publisher_row.find_all('a')] print(publishers) # 输出: ['FromSoftware, Inc.', 'Bandai Namco Entertainment'] else: print("N/A")
方法二:遍历所有dev_row匹配标识
遍历页面中所有dev_row元素,检查其中的副标题文本是否为「Publisher:」,匹配成功后提取发行商名称:
from bs4 import BeautifulSoup publishers = [] for dev_row in soup.find_all('div', class_='dev_row'): subtitle = dev_row.find('div', class_='subtitle column') if subtitle and subtitle.text.strip() == 'Publisher:': publishers = [a.text.strip() for a in dev_row.find_all('a')] break print(publishers if publishers else "N/A")
说明
两种方法均能避开dev_row重复的问题,精准定位到发行商所在行。方法一效率更高,方法二可兼容副标题文本存在多余空格的场景。
内容的提问来源于stack exchange,提问作者Walter Paleari
相关产品推荐
相关产品推荐

