如何用BeautifulSoup4仅获取class为right而非right gamelink的元素文本?
问题:如何在BeautifulSoup4中仅匹配仅含"right"类的元素?
我是BeautifulSoup4(BS4)新手,尝试抓取特定HTML类的数据。我的HTML片段如下:
<td class="right">31</td> <td class="right gamelink"> <a href="/boxscores/20220908ram.htm"> "F" <span class="no_mobile">inal</span> </a> </td>
当我使用findAll()方法查找class为"right"的元素时,会同时获取到class为"right gamelink"的元素内容。我的代码如下:
from bs4 import BeautifulSoup import requests weekNumber = 1 url = "https://www.pro-football-reference.com/years/2022/week_"+str(weekNumber)+".htm" print(url) req = requests.get(url) webpage = BeautifulSoup(req.text, 'html.parser') scores = webpage.findAll("td", attrs={'class': 'right'}) for score in scores: current_score = score.text.strip() print(current_score)
输出结果为:
31 Final
请问是否有办法指定仅返回仅含"right"类的元素文本,而非包含"right gamelink"类的内容?
解决方案
以下几种方法都可以实现仅匹配仅包含"right"单个类的元素:
方法1:使用CSS选择器精确匹配
利用CSS选择器的:not()伪类,筛选出class属性值中不含空格的元素(即只有单个类的元素):scores = webpage.select("td.right:not([class*=' '])")方法2:过滤元素的class列表长度
先获取所有带有"right"类的元素,再过滤掉class列表长度大于1的元素:scores = webpage.findAll("td", class_="right") filtered_scores = [score for score in scores if len(score.get('class', [])) == 1] # 后续遍历filtered_scores即可 for score in filtered_scores: current_score = score.text.strip() print(current_score)这里
score.get('class', [])会返回元素的类名列表,长度为1说明只有"right"这一个类。方法3:直接精确匹配class属性值
通过lambda表达式判断class属性值是否完全等于"right":scores = webpage.findAll("td", attrs={'class': lambda x: x == 'right'})也可以用CSS选择器的精确属性匹配:
scores = webpage.select("td[class='right']")
内容的提问来源于stack exchange,提问作者acegay27
相关产品推荐
相关产品推荐

