You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup4仅获取class为right而非right gamelink的元素文本?

问题:如何在BeautifulSoup4中仅匹配仅含"right"类的元素?

我是BeautifulSoup4(BS4)新手,尝试抓取特定HTML类的数据。我的HTML片段如下:

<td class="right">31</td>
<td class="right gamelink">
   <a href="/boxscores/20220908ram.htm">
      "F"
      <span class="no_mobile">inal</span>
   </a>
</td>

当我使用findAll()方法查找class为"right"的元素时,会同时获取到class为"right gamelink"的元素内容。我的代码如下:

from bs4 import BeautifulSoup
import requests

weekNumber = 1
url = "https://www.pro-football-reference.com/years/2022/week_"+str(weekNumber)+".htm"

print(url)

req = requests.get(url)
webpage = BeautifulSoup(req.text, 'html.parser')

scores = webpage.findAll("td", attrs={'class': 'right'})

for score in scores:
    current_score = score.text.strip()
    print(current_score)

输出结果为:

31
Final

请问是否有办法指定仅返回仅含"right"类的元素文本,而非包含"right gamelink"类的内容?


解决方案

以下几种方法都可以实现仅匹配仅包含"right"单个类的元素:

  • 方法1:使用CSS选择器精确匹配
    利用CSS选择器的:not()伪类,筛选出class属性值中不含空格的元素(即只有单个类的元素):

    scores = webpage.select("td.right:not([class*=' '])")
    
  • 方法2:过滤元素的class列表长度
    先获取所有带有"right"类的元素,再过滤掉class列表长度大于1的元素:

    scores = webpage.findAll("td", class_="right")
    filtered_scores = [score for score in scores if len(score.get('class', [])) == 1]
    
    # 后续遍历filtered_scores即可
    for score in filtered_scores:
        current_score = score.text.strip()
        print(current_score)
    

    这里score.get('class', [])会返回元素的类名列表,长度为1说明只有"right"这一个类。

  • 方法3:直接精确匹配class属性值
    通过lambda表达式判断class属性值是否完全等于"right":

    scores = webpage.findAll("td", attrs={'class': lambda x: x == 'right'})
    

    也可以用CSS选择器的精确属性匹配:

    scores = webpage.select("td[class='right']")
    

内容的提问来源于stack exchange,提问作者acegay27

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 07:35:22