求助:使用Python+BeautifulSoup无法获取F2实时车手积分榜更新
F2实时车手积分榜爬取解决方案
问题根源
F2实时计时页面的实时数据是通过JavaScript动态渲染生成的,而BeautifulSoup仅能解析静态HTML内容,无法捕获JS加载后的动态数据,这是你无法获取实时更新的核心原因。
可行解决方案
方法一:用Selenium模拟浏览器加载动态内容
Selenium可以模拟真实浏览器的行为,自动执行页面JS,从而获取到实时渲染的数据。
- 安装依赖
pip install selenium beautifulsoup4
- 核心代码实现
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import time # 获取赛前静态积分榜 def get_pre_race_standings(): standings = {} driver = webdriver.Chrome() driver.get("https://www.fiaformula2.com/standings/drivers") # 等待页面加载完成 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, "table.standings-table")) ) soup = BeautifulSoup(driver.page_source, "html.parser") for row in soup.select("table.standings-table tr")[1:]: # 跳过表头行 cols = row.find_all("td") if len(cols) >= 4: driver_name = cols[1].text.strip() points = int(cols[-1].text.strip()) standings[driver_name] = points driver.quit() return standings # 实时获取比赛名次 def get_live_positions(): driver = webdriver.Chrome() driver.get("https://www.fiaformula2.com/livetiming/index.html") # 等待实时数据表格加载 try: WebDriverWait(driver, 15).until( EC.presence_of_element_located((By.CSS_SELECTOR, ".live-timing__table tr")) ) except: print("实时页面加载超时") driver.quit() return while True: soup = BeautifulSoup(driver.page_source, "html.parser") positions = [] for row in soup.select(".live-timing__table tr")[1:]: cols = row.find_all("td") if len(cols) >= 2: try: pos = int(cols[0].text.strip()) driver_name = cols[1].text.strip() positions.append((pos, driver_name)) except ValueError: continue # 跳过非数字的名次(如退赛车手) yield positions time.sleep(5) # 每5秒刷新一次数据 # 计算实时积分榜 def calculate_live_standings(): pre_standings = get_pre_race_standings() # F2积分规则(正赛,冲刺赛规则需调整) points_system = {1:25, 2:18, 3:15, 4:12, 5:10, 6:8, 7:6, 8:4, 9:2, 10:1} for live_positions in get_live_positions(): live_standings = pre_standings.copy() for pos, driver in live_positions: if pos in points_system: live_standings[driver] = pre_standings.get(driver, 0) + points_system[pos] # 打印排序后的实时积分榜 print("\n=== 实时车手积分榜 ===") for idx, (driver, points) in enumerate(sorted(live_standings.items(), key=lambda x: x[1], reverse=True), 1): print(f"{idx}. {driver}: {points}分") if __name__ == "__main__": calculate_live_standings()
- 注意事项
- 需下载对应浏览器的驱动(如ChromeDriver),并确保其路径可被Python识别,或使用
webdriver-manager自动管理驱动版本。 - 页面元素选择器可能随网站更新变化,需根据实际页面结构调整CSS选择器。
- 可根据比赛类型(正赛/冲刺赛)修改积分规则,同时注意最快圈速额外积分的处理。
方法二:抓包获取实时数据API(更高效)
通过浏览器开发者工具(F12)的网络面板,捕获实时页面的XHR请求,找到返回车手名次的API接口,直接用requests请求JSON数据,无需模拟浏览器。
示例代码:
import requests import time def get_live_positions_via_api(): # 替换为抓包获取的真实API地址 api_url = "https://api.example.com/f2/live/positions" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } while True: resp = requests.get(api_url, headers=headers) if resp.status_code == 200: data = resp.json() positions = [(item["position"], item["driver"]["fullName"]) for item in data["drivers"]] yield positions time.sleep(5) # 后续积分计算逻辑同方法一
这种方法资源占用更低,但需要自行抓包分析API接口,部分接口可能需要处理认证或签名。
内容的提问来源于stack exchange,提问作者frd
相关产品推荐
相关产品推荐

