You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup抓取带有指定class属性的href标签?

错误原因
  • 类名绑定对象匹配错误
    你观察到的sidearm-sports-file-link-read sidearm-sports-file-link-processed类属于h3标签内部的a标签,并非h3标签本身的属性。你在代码中直接将该类作为h3标签的筛选条件,没有匹配到符合要求的元素,所以返回空列表。
  • 可选优化:部分站点会拦截无特征的爬虫请求,添加User-Agent请求头模拟浏览器访问可以降低被拦截的概率。
修正后可运行代码
import requests
from bs4 import BeautifulSoup

# 模拟浏览器请求头
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36"
}
page = requests.get('https://athletics.baruch.cuny.edu/sports/mens-swimming-and-diving/roster', headers=headers)
soup = BeautifulSoup(page.content, 'html.parser')
# 筛选携带指定类名的a标签,提取文本即为人员姓名
name_tags = soup.find_all('a', class_="sidearm-sports-file-link-read sidearm-sports-file-link-processed")
names = [tag.get_text(strip=True) for tag in name_tags]
print(names)

内容的提问来源于stack exchange,提问作者user17003245

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 13:45:08