You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬取Flashscore:无法获取赛事ID以获取H2H数据

解决Flashscore赛事ID获取与H2H数据爬取问题

核心思路

Flashscore的赛事ID并非藏得很深,要么在首页HTML的赛事链接属性里,要么在动态加载的API数据中。用requests替代Selenium的关键是模拟浏览器请求头+精准定位数据来源。

方法一:解析首页HTML提取赛事ID

Flashscore首页的赛事条目会包含指向单场赛事的链接,从中可以直接提取赛事ID:

import requests
from bs4 import BeautifulSoup

# 模拟浏览器请求头,避免被反爬拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
    "Referer": "https://www.flashscore.com/"
}

# 明日赛事页面URL
tomorrow_matches_url = "https://www.flashscore.com/football/tomorrow/"
response = requests.get(tomorrow_matches_url, headers=headers)
soup = BeautifulSoup(response.text, "html.parser")

# 遍历所有赛事链接,提取ID
match_ids = []
for match_link in soup.find_all("a", class_="event__link"):
    href = match_link.get("href")
    if "/match/" in href:
        # 从链接中拆分出赛事ID,例如"/match/abc123/" -> "abc123"
        match_id = href.split("/")[2]
        match_ids.append(match_id)
        print(f"已获取赛事ID: {match_id}")

# 构造H2H页面URL并请求
for match_id in match_ids[:2]:  # 先测试前2个赛事
    h2h_url = f"https://www.flashscore.com/match/{match_id}/#/match-head-to-head/overall"
    h2h_response = requests.get(h2h_url, headers=headers)
    # 这里可以继续用BeautifulSoup解析H2H页面的统计数据
    print(f"H2H页面请求状态: {h2h_response.status_code}")

方法二:直接调用赛事列表API(更稳定)

如果首页HTML是动态渲染的(requests拿不到赛事数据),可以通过浏览器开发者工具抓包,找到Flashscore加载赛事列表的API接口,直接从接口JSON中提取赛事ID:

import requests
import json

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
    "Referer": "https://www.flashscore.com/"
}

# 替换为你抓包得到的赛事列表API URL(Flashscore的API通常以d.flashscore.com开头)
api_url = "https://d.flashscore.com/x/feed/df_m_1"
response = requests.get(api_url, headers=headers)

# Flashscore的API返回是JSONP格式,需要去掉前后的包裹字符
raw_data = response.text.lstrip("fsCallback(").rstrip(");")
match_data = json.loads(raw_data)

# 从API数据中提取赛事ID
for match in match_data["doc"]["matches"]:
    match_id = match["id"]
    print(f"赛事ID: {match_id}")
    # 同样可以构造H2H的API或页面URL进行后续请求

注意事项

  • 反爬处理:如果requests请求被拦截,复制Selenium请求中的Cookie字段加入headers,或者使用requests.Session()保持会话。
  • 数据解析:H2H页面的统计数据可以用BeautifulSoup定位对应class提取,或者抓H2H数据的API接口直接拿结构化JSON。
  • 接口更新:Flashscore的API路径可能会变动,若出现请求失败,重新用浏览器抓包更新API URL即可。

内容的提问来源于stack exchange,提问作者Sergey

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 15:20:25