如何使用Python BeautifulSoup抓取自行车赛事页面的选手完赛时间数据
自行车赛事完赛时间提取实现方案
1. 安装依赖库
需要用到requests发起网络请求,beautifulsoup4解析HTML,执行以下命令安装:
pip install requests beautifulsoup4
2. 完整实现代码
import requests from bs4 import BeautifulSoup # 目标赛事页面地址 url = "https://www.datasport.com/live/ranking/?lang=en&racenr=23466#1_5584A2" # 添加请求头避免被反爬拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } # 发起请求获取页面内容 response = requests.get(url, headers=headers) response.encoding = "utf-8" # 解析HTML soup = BeautifulSoup(response.text, "html.parser") # 按提供的特征匹配所有选手行 race_rows = soup.find_all("tr", class_="Hover LastRecordLine", style="cursor: pointer;") # 存储完赛时间的数组 finish_times = [] for row in race_rows: # 匹配时间所在的td标签 time_tag = row.find("td", style="font-weight: bold; text-align: left;") if time_tag: # 提取时间文本并去除多余空格后加入数组 time_str = time_tag.get_text(strip=True) finish_times.append(time_str) # 输出结果测试 print(f"共提取到{len(finish_times)}条完赛时间:") print(finish_times)
注意事项
- 如果直接请求返回的HTML中没有对应数据,说明页面数据是JavaScript动态加载的,可以改用Selenium、Playwright等工具模拟浏览器加载完整页面后再用BeautifulSoup解析
- 如果后续页面结构有调整,只需要对应修改
find_all和find方法里的匹配规则即可
内容的提问来源于stack exchange,提问作者chris_tri
相关产品推荐
相关产品推荐

