基于最小时间差过滤JSON赛事分组数据的技术实现求助
赛事数据处理代码完善方案
需求概述
- 按
Day和Championship对file.json中的赛事数据分组排序,计算组内相邻赛事的时间差; - 若组内存在至少一组相邻赛事不满足设定的最小时间差,则丢弃该分组,不输出;
- 输出每场赛事的
Home和Away字段; - 将时间差格式化为可读文本。
原始数据(file.json)
[{ "Day": "Giornata 29", "Matches": { "Home": "Egnatia", "Away": "Kukesi", "Times": "19.03. 15:00", "Championship": "CALCIO\nALBANIA Super League\n2023/2024" } },{ "Day": "Giornata 29", "Matches": { "Home": "Egnatia", "Away": "Kukesi", "Times": "20.03. 16:09", "Championship": "CALCIO\nALBANIA Super League\n2023/2024" } }, { "Day": "Giornata 29", "Matches": { "Home": "Egnatia", "Away": "Kukesi", "Times": "19.03. 16:15", "Championship": "CALCIO\nALBANIA Super League\n2023/2024" } },{ "Day": "Giornata 41", "Matches": { "Home": "Lincoln", "Away": "Leyton Orient", "Times": "19.03. 16:00", "Championship": "CALCIO\nINGHILTERRA League One\n2023/2024" } }, { "Day": "Giornata 41", "Matches": { "Home": "Lincoln", "Away": "Leyton Orient", "Times": "19.03. 18:00", "Championship": "CALCIO\nINGHILTERRA League One\n2023/2024" } }, { "Day": "Giornata 30", "Matches": { "Home": "Napoli", "Away": "Atalanta", "Times": "30.03. 12:30", "Championship": "CALCIO\nITALIA Serie A\n2023/2024" } }]
基础代码(code.py)
import collections import datetime import json with open('file.json', 'r+') as f: fileData = json.load(f) # no f.close() is needed as the context manager handles this # create a default dictionary with nested default dictionary that has a list in it similar = collections.defaultdict(lambda : collections.defaultdict(list)) for file in fileData: # for every file put it into the dictionary based on the day and the championship similar[file["Day"]][file["Matches"]["Championship"]].append(file) # make a dictionary from the defaultdict championships = {k:dict(v) for k,v in similar.items()} # loop over the daysuell for day in championships: # loop over the championships for championship in championships.get(day): # get the values values = championships.get(day).get(championship) # extract the times and parse the string to a timestamp and sort the list times = sorted([datetime.datetime.strptime(val.get("Matches").get("Times"), "%d.%m. %H:%M") for val in values]) # calculate the time difference for every element and its adjacent element # if the list has one or no elements return None time_difference = None if len(times) <= 1 else list(map(lambda t: t[-1]-t[0], zip(times, times[1:]))) print(day, championship, time_difference, '\n\n')
完善后的代码
import collections import datetime import json # 自定义最小时间差(可根据实际需求调整,示例为24小时) MIN_TIME_DELTA = datetime.timedelta(hours=24) def format_timedelta(td): """将datetime.timedelta转换为可读的中文文本""" days = td.days hours, remainder = divmod(td.seconds, 3600) minutes, _ = divmod(remainder, 60) parts = [] if days > 0: parts.append(f"{days}天") if hours > 0: parts.append(f"{hours}小时") if minutes > 0: parts.append(f"{minutes}分钟") return " ".join(parts) if parts else "0分钟" # 读取JSON数据 with open('file.json', 'r', encoding='utf-8') as f: match_data = json.load(f) # 按Day和Championship嵌套分组 grouped_matches = collections.defaultdict(lambda: collections.defaultdict(list)) for entry in match_data: day = entry["Day"] champ = entry["Matches"]["Championship"] grouped_matches[day][champ].append(entry) # 遍历处理每个分组 for day, champs in grouped_matches.items(): for champ, matches in champs.items(): # 解析时间并按时间顺序排序赛事 sorted_matches = sorted( matches, key=lambda x: datetime.datetime.strptime(x["Matches"]["Times"], "%d.%m. %H:%M") ) parsed_times = [ datetime.datetime.strptime(m["Matches"]["Times"], "%d.%m. %H:%M") for m in sorted_matches ] # 计算相邻时间差并验证是否符合要求 valid_group = True time_diffs = [] if len(parsed_times) > 1: for i in range(len(parsed_times)-1): diff = parsed_times[i+1] - parsed_times[i] time_diffs.append(diff) if diff < MIN_TIME_DELTA: valid_group = False break # 不满足条件则跳过当前分组 if not valid_group: continue # 输出分组头部信息 total_matches = len(sorted_matches) print(f"{day} {champ} **{total_matches}/{total_matches}**") # 输出每场赛事 for idx, match in enumerate(sorted_matches): home = match["Matches"]["Home"] away = match["Matches"]["Away"] # 单场赛事无相邻差则无后缀,否则标记满足 suffix = " v" if idx < len(time_diffs) else "" print(f"{home} - {away}{suffix}") # 输出分隔线 print("---------------------------")
关键实现说明
- 分组逻辑:使用
collections.defaultdict实现Day→Championship→赛事列表的嵌套分组,确保数据按规则归类; - 时间处理:解析赛事时间字符串为
datetime对象,按时间排序后计算相邻赛事的时间差; - 有效性验证:遍历时间差列表,若存在任意一个小于设定的
MIN_TIME_DELTA,则标记分组无效并跳过输出; - 格式化输出:将
timedelta转换为“X天Y小时Z分钟”的可读格式,同时按要求输出赛事对阵信息。
输出结果
当MIN_TIME_DELTA设为24小时时,Giornata 29分组中存在相邻赛事时间差为23小时54分(小于24小时),因此该分组被丢弃,最终输出如下:
Giornata 41 CALCIO INGHILTERRA League One 2023/2024 **2/2** Lincoln - Leyton Orient v Lincoln - Leyton Orient v --------------------------- Giornata 30 CALCIO ITALIA Serie A 2023/2024 **1/1** Napoli - Atalanta ---------------------------
内容的提问来源于stack exchange,提问作者x__SHARINGAN____x
相关产品推荐
相关产品推荐

