You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

技术问询:Python爬取Sofascore丹麦超2022/2023赛季进球时间至Excel的问题

问题说明

我使用Excel 2019和Python 3.12.4,想要获取丹麦足球超级联赛2022/2023赛季所有赛事的日期、主客场球队、比赛结果及进球时间,并将数据导入Excel。目前我的Python代码仅能获取单场赛事的进球时间,且获取的进球时间存在错误。现有代码如下:

import requests
from bs4 import BeautifulSoup
import pandas as pd
import os

# Define the headers
headers = {'User-Agent': 'Mozilla/5.0'}

# Scraping the main page (although this part is not used later in the code)
response = requests.get(
    'https://www.sofascore.com/fc-midtjylland-brondby-if/GAsOA#id:12174728',
    headers=headers
)
soup = BeautifulSoup(response.text, 'html.parser')

# Get pregame form data
response_form = requests.get('https://api.sofascore.com/api/v1/event/12174728/pregame-form', headers=headers)
form = response_form.json()

# Get incidents data
response_incidents = requests.get('https://api.sofascore.com/api/v1/event/12174728/incidents', headers=headers)
incidents = response_incidents.json()
incidents = incidents['incidents']
incidents_df = pd.json_normalize(incidents)

# Filter goal incidents
goals_df = incidents_df[incidents_df['incidentType'] == 'goal'].loc[:, ['time']]


# Create a Pandas Excel writer using XlsxWriter as the engine
output_file = os.path.join(os.path.dirname(__file__), 'sofascore_data.xlsx')
with pd.ExcelWriter(output_file, engine='xlsxwriter') as writer:
    goals_df.to_excel(writer, sheet_name='Goals', index=False)


print(f'Data has been written to {output_file}')
修正方案

以下是能批量获取赛季所有赛事数据、修复进球时间问题的完整代码:

import requests
import pandas as pd
import os
from datetime import datetime

# 请求头,模拟浏览器访问
headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'}

# 丹麦超2022/2023赛季的API标识
SEASON_ID = 61326
BASE_URL = 'https://api.sofascore.com/api/v1'

def get_season_events():
    """获取赛季所有赛事的基础信息"""
    url = f"{BASE_URL}/season/{SEASON_ID}/events"
    response = requests.get(url, headers=headers)
    response.raise_for_status()
    events_data = response.json()['events']
    
    # 提取赛事核心信息,只处理已结束赛事
    events_list = []
    for event in events_data:
        if event['status']['type'] != 'finished':
            continue
        events_list.append({
            'event_id': event['id'],
            'date': datetime.fromtimestamp(event['startTimestamp']).strftime('%Y-%m-%d %H:%M'),
            'home_team': event['homeTeam']['name'],
            'away_team': event['awayTeam']['name'],
            'home_score': event['homeScore']['current'],
            'away_score': event['awayScore']['current'],
        })
    return pd.DataFrame(events_list)

def get_event_goals(event_id):
    """获取单场赛事的所有进球时间,并修正格式"""
    url = f"{BASE_URL}/event/{event_id}/incidents"
    response = requests.get(url, headers=headers)
    response.raise_for_status()
    incidents = response.json()['incidents']
    
    goals = []
    for incident in incidents:
        if incident['incidentType'] != 'goal':
            continue
        # 合并常规时间与补时,生成准确的进球时间格式
        time = incident['time']
        added_time = incident.get('addedTime', 0)
        display_time = f"{time}+{added_time}" if added_time > 0 else str(time)
        goals.append({
            'event_id': event_id,
            'goal_time': display_time,
            'scoring_team': incident['team']['name'],
            'player': incident['player']['name']
        })
    return pd.DataFrame(goals)

if __name__ == '__main__':
    # 获取所有赛事基础数据
    events_df = get_season_events()
    print(f"获取到{len(events_df)}场已结束赛事")
    
    # 遍历所有赛事,获取进球数据
    all_goals = []
    for idx, row in events_df.iterrows():
        event_id = row['event_id']
        print(f"正在处理赛事: {row['home_team']} vs {row['away_team']}")
        goals_df = get_event_goals(event_id)
        all_goals.append(goals_df)
    
    # 合并进球数据与赛事基础数据
    goals_combined = pd.concat(all_goals, ignore_index=True)
    final_df = pd.merge(events_df, goals_combined, on='event_id', how='left')
    
    # 写入Excel
    output_file = os.path.join(os.path.dirname(__file__), '丹麦超2022-2023赛季数据.xlsx')
    with pd.ExcelWriter(output_file, engine='xlsxwriter') as writer:
        final_df.to_excel(writer, sheet_name='赛事数据', index=False)
    
    print(f"数据已成功写入文件: {output_file}")

关键说明

  1. 批量获取赛事:通过赛季ID调用API直接获取该赛季所有已结束赛事的基础信息,无需手动指定单场赛事ID
  2. 进球时间修正:结合API返回的time(常规时间)和addedTime(补时)字段,生成符合观赛习惯的进球时间格式(如90+3)
  3. 数据关联:将赛事基础信息(日期、球队、比分)与进球数据绑定,每条进球记录都能对应到具体赛事
  4. 可靠性优化:添加response.raise_for_status()捕获API请求失败的情况,避免无效数据混入

内容的提问来源于stack exchange,提问作者Fred

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 15:24:54