You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于最小时间差过滤JSON赛事分组数据的技术实现求助

赛事数据处理代码完善方案

需求概述

  • 按Day和Championship对file.json中的赛事数据分组排序,计算组内相邻赛事的时间差;
  • 若组内存在至少一组相邻赛事不满足设定的最小时间差,则丢弃该分组,不输出;
  • 输出每场赛事的Home和Away字段;
  • 将时间差格式化为可读文本。

原始数据(file.json)

[{
    "Day": "Giornata 29",
    "Matches": {
        "Home": "Egnatia",
        "Away": "Kukesi",
        "Times": "19.03. 15:00",
        "Championship": "CALCIO\nALBANIA Super League\n2023/2024"
    }
},{
    "Day": "Giornata 29",
    "Matches": {
        "Home": "Egnatia",
        "Away": "Kukesi",
        "Times": "20.03. 16:09",
        "Championship": "CALCIO\nALBANIA Super League\n2023/2024"
    }
},
{
    "Day": "Giornata 29",
    "Matches": {
        "Home": "Egnatia",
        "Away": "Kukesi",
        "Times": "19.03. 16:15",
        "Championship": "CALCIO\nALBANIA Super League\n2023/2024"
    }
},{
    "Day": "Giornata 41",
    "Matches": {
        "Home": "Lincoln",
        "Away": "Leyton Orient",
        "Times": "19.03. 16:00",
        "Championship": "CALCIO\nINGHILTERRA League One\n2023/2024"
    }
},
{
    "Day": "Giornata 41",
    "Matches": {
        "Home": "Lincoln",
        "Away": "Leyton Orient",
        "Times": "19.03. 18:00",
        "Championship": "CALCIO\nINGHILTERRA League One\n2023/2024"
    }
},
{
    "Day": "Giornata 30",
    "Matches": {
        "Home": "Napoli",
        "Away": "Atalanta",
        "Times": "30.03. 12:30",
        "Championship": "CALCIO\nITALIA Serie A\n2023/2024"
    }
}]

基础代码(code.py)

import collections
import datetime
import json

with open('file.json', 'r+') as f:
    fileData = json.load(f)
    # no f.close() is needed as the context manager handles this

# create a default dictionary with nested default dictionary that has a list in it
similar = collections.defaultdict(lambda : collections.defaultdict(list))

for file in fileData:
    # for every file put it into the dictionary based on the day and the championship
    similar[file["Day"]][file["Matches"]["Championship"]].append(file)

# make a dictionary from the defaultdict
championships = {k:dict(v) for k,v in similar.items()}

# loop over the daysuell
for day in championships:
    # loop over the championships
    for championship in championships.get(day):
        # get the values
        values = championships.get(day).get(championship)
        # extract the times and parse the string to a timestamp and sort the list
        times = sorted([datetime.datetime.strptime(val.get("Matches").get("Times"), "%d.%m. %H:%M") for val in values])

        # calculate the time difference for every element and its adjacent element
        # if the list has one or no elements return None
        time_difference = None if len(times) <= 1 else list(map(lambda t: t[-1]-t[0], zip(times, times[1:])))
        
        print(day, championship, time_difference, '\n\n')

完善后的代码

import collections
import datetime
import json

# 自定义最小时间差(可根据实际需求调整,示例为24小时)
MIN_TIME_DELTA = datetime.timedelta(hours=24)

def format_timedelta(td):
    """将datetime.timedelta转换为可读的中文文本"""
    days = td.days
    hours, remainder = divmod(td.seconds, 3600)
    minutes, _ = divmod(remainder, 60)
    parts = []
    if days > 0:
        parts.append(f"{days}天")
    if hours > 0:
        parts.append(f"{hours}小时")
    if minutes > 0:
        parts.append(f"{minutes}分钟")
    return " ".join(parts) if parts else "0分钟"

# 读取JSON数据
with open('file.json', 'r', encoding='utf-8') as f:
    match_data = json.load(f)

# 按Day和Championship嵌套分组
grouped_matches = collections.defaultdict(lambda: collections.defaultdict(list))
for entry in match_data:
    day = entry["Day"]
    champ = entry["Matches"]["Championship"]
    grouped_matches[day][champ].append(entry)

# 遍历处理每个分组
for day, champs in grouped_matches.items():
    for champ, matches in champs.items():
        # 解析时间并按时间顺序排序赛事
        sorted_matches = sorted(
            matches,
            key=lambda x: datetime.datetime.strptime(x["Matches"]["Times"], "%d.%m. %H:%M")
        )
        parsed_times = [
            datetime.datetime.strptime(m["Matches"]["Times"], "%d.%m. %H:%M")
            for m in sorted_matches
        ]
        
        # 计算相邻时间差并验证是否符合要求
        valid_group = True
        time_diffs = []
        if len(parsed_times) > 1:
            for i in range(len(parsed_times)-1):
                diff = parsed_times[i+1] - parsed_times[i]
                time_diffs.append(diff)
                if diff < MIN_TIME_DELTA:
                    valid_group = False
                    break
        
        # 不满足条件则跳过当前分组
        if not valid_group:
            continue
        
        # 输出分组头部信息
        total_matches = len(sorted_matches)
        print(f"{day} {champ} **{total_matches}/{total_matches}**")
        # 输出每场赛事
        for idx, match in enumerate(sorted_matches):
            home = match["Matches"]["Home"]
            away = match["Matches"]["Away"]
            # 单场赛事无相邻差则无后缀,否则标记满足
            suffix = " v" if idx < len(time_diffs) else ""
            print(f"{home} - {away}{suffix}")
        # 输出分隔线
        print("---------------------------")

关键实现说明

  1. 分组逻辑:使用collections.defaultdict实现Day→Championship→赛事列表的嵌套分组,确保数据按规则归类;
  2. 时间处理:解析赛事时间字符串为datetime对象,按时间排序后计算相邻赛事的时间差;
  3. 有效性验证:遍历时间差列表,若存在任意一个小于设定的MIN_TIME_DELTA,则标记分组无效并跳过输出;
  4. 格式化输出:将timedelta转换为“X天Y小时Z分钟”的可读格式,同时按要求输出赛事对阵信息。

输出结果

当MIN_TIME_DELTA设为24小时时,Giornata 29分组中存在相邻赛事时间差为23小时54分(小于24小时),因此该分组被丢弃,最终输出如下:

Giornata 41 CALCIO
INGHILTERRA League One
2023/2024 **2/2**
Lincoln - Leyton Orient v
Lincoln - Leyton Orient v
---------------------------
Giornata 30 CALCIO
ITALIA Serie A
2023/2024 **1/1**
Napoli - Atalanta
---------------------------

内容的提问来源于stack exchange,提问作者x__SHARINGAN____x

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 16:15:54