求助:用Python爬取SofaScore意甲球队首发阵容、评分等数据
解决SofaScore意甲数据爬取问题
你的代码为什么找不到目标内容?
- 动态内容加载:SofaScore的比赛数据是通过JavaScript动态渲染的,
requests获取的只是初始HTML框架,没有包含你要的首发、评分等核心数据。 - 动态类名:你用的
sc-fqkvVR eeeBnr sc-d8bc48b6-2 cUcAWg这类带sc-前缀的类名,是前端框架自动生成的,每次页面加载都可能变化,完全不能作为定位依据。
更靠谱的方案:调用SofaScore公开API
SofaScore有一套未公开但稳定可用的API接口,直接请求API比解析HTML高效10倍,还不用处理动态渲染问题。
1. 单场比赛数据爬取示例
以你提供的萨索洛vs亚特兰大的比赛为例:
- 先从浏览器开发者工具的Network面板找到比赛的数字
eventId(这场的ID是11096647),用以下接口获取数据:- 首发阵容+球员评分:
https://api.sofascore.com/api/v1/event/{eventId}/lineups - 进阶统计数据:
https://api.sofascore.com/api/v1/event/{eventId}/statistics
- 首发阵容+球员评分:
代码示例:
import requests import time # 模拟浏览器请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/116.0.0.0 Safari/537.36' } def get_single_match_data(event_id): # 获取首发阵容和评分 lineups_url = f'https://api.sofascore.com/api/v1/event/{event_id}/lineups' lineups_resp = requests.get(lineups_url, headers=headers) lineups_data = lineups_resp.json() # 提取主队首发信息 home_starting = [] for player in lineups_data['home']['startingLineup']: home_starting.append({ 'name': player['player']['name'], 'position': player['position'], 'rating': player.get('rating', 'N/A') }) # 提取客队首发信息 away_starting = [] for player in lineups_data['away']['startingLineup']: away_starting.append({ 'name': player['player']['name'], 'position': player['position'], 'rating': player.get('rating', 'N/A') }) # 获取进阶统计数据 stats_url = f'https://api.sofascore.com/api/v1/event/{event_id}/statistics' stats_resp = requests.get(stats_url, headers=headers) stats_data = stats_resp.json() # 整理统计数据(如控球率、射门数等) match_stats = {} for group in stats_data: for item in group['statisticsItems']: match_stats[item['name']] = { 'home': item['home'], 'away': item['away'] } return { 'home_starting': home_starting, 'away_starting': away_starting, 'stats': match_stats } # 测试单场比赛数据爬取 match_data = get_single_match_data(11096647) print("主队首发:", match_data['home_starting']) print("\n客队首发:", match_data['away_starting']) print("\n比赛统计:", match_data['stats']) # 避免频繁请求,加延迟 time.sleep(2)
2. 扩展到意甲所有轮次和球队
要爬取整个赛季的所有比赛,按以下步骤操作:
- 获取意甲赛季ID:意甲的唯一联赛ID是
11,请求https://api.sofascore.com/api/v1/unique-tournament/11/seasons,找到当前赛季的id(比如2024/2025赛季ID为51505)。 - 获取赛季所有轮次:请求
https://api.sofascore.com/api/v1/unique-tournament/11/season/{seasonId}/rounds,拿到所有轮次编号。 - 遍历轮次获取比赛ID:对每个轮次,请求
https://api.sofascore.com/api/v1/unique-tournament/11/season/{seasonId}/round/{roundNumber}/events,提取所有比赛的eventId。 - 批量爬取比赛数据:用上面的
get_single_match_data函数处理每个eventId。
代码示例(获取赛季所有比赛ID):
def get_serie_a_season_events(season_id): # 获取所有轮次信息 rounds_url = f'https://api.sofascore.com/api/v1/unique-tournament/11/season/{season_id}/rounds' rounds_resp = requests.get(rounds_url, headers=headers) rounds_data = rounds_resp.json() all_event_ids = [] for round_info in rounds_data: round_num = round_info['round'] # 获取该轮所有比赛 round_events_url = f'https://api.sofascore.com/api/v1/unique-tournament/11/season/{season_id}/round/{round_num}/events' round_events_resp = requests.get(round_events_url, headers=headers) round_events_data = round_events_resp.json() for event in round_events_data['events']: all_event_ids.append(event['id']) time.sleep(1) # 加延迟避免被封 return all_event_ids # 测试获取2024/2025赛季所有比赛ID season_id = 51505 all_event_ids = get_serie_a_season_events(season_id) print("意甲2024/2025赛季所有比赛ID:", all_event_ids)
注意事项
- 每次请求后加1-2秒延迟,避免触发反爬机制导致IP被封。
- API返回的JSON结构清晰,直接通过键值对提取数据即可,无需解析复杂HTML。
- 若请求失败,检查请求头是否正确,必要时更换IP或使用代理。
内容的提问来源于stack exchange,提问作者criug95
相关产品推荐
相关产品推荐

