如何用Python合并JSON中连续相同speaker的对应字段值?
解决方案
首先,先处理JSON文件的读写,再实现连续相同speaker的合并逻辑,具体步骤如下:
1. 修复原始JSON的语法错误
你提供的原始JSON存在语法问题:每个字典内的键值对末尾缺少逗号(比如"says": "Hello."后面需要加,),否则Python无法正常解析。修正后的JSON示例:
{"turns": [{ "speaker": "A", "says": "Hello.", "other": "aaaa" }, { "speaker": "B", "says": "Hi.", "other": "bbbb" }, { "speaker": "B", "says": "I'm busy now.", "other": "ccccc" }, { "speaker": "A", "says": "See you later?", "other": "dddd" }, { "speaker": "B", "says": "Sure.", "other": "eeee" }, { "speaker": "B", "says": "Bye", "other": "ffff" }, { "speaker": "A", "says": "Bye bye.", "other": "gggg" }] }
2. Python实现代码
下面是完整的代码,包含JSON读写和合并逻辑,注释已写清每一步作用:
import json def merge_consecutive_speakers(data): merged_turns = [] for turn in data["turns"]: # 结果列表为空时,直接添加当前条目 if not merged_turns: merged_turns.append(turn.copy()) continue last_turn = merged_turns[-1] # 当前说话者与最后一条一致时,合并内容 if turn["speaker"] == last_turn["speaker"]: # 用空格拼接says字段 last_turn["says"] = f"{last_turn['says']} {turn['says']}" # 用空格拼接other字段 last_turn["other"] = f"{last_turn['other']} {turn['other']}" else: # 说话者不同,添加新条目 merged_turns.append(turn.copy()) return {"turns": merged_turns} # 读取原始JSON文件 with open("input.json", "r", encoding="utf-8") as f: original_data = json.load(f) # 执行合并操作 merged_data = merge_consecutive_speakers(original_data) # 将合并结果写入新JSON文件 with open("output.json", "w", encoding="utf-8") as f: json.dump(merged_data, f, indent=2, ensure_ascii=False)
代码说明
json.load()和json.dump():分别负责读取JSON文件和写入JSON文件,indent=2让输出格式更易读,ensure_ascii=False保证非ASCII字符正常显示。turn.copy():字典是可变对象,直接添加会导致后续修改影响原数据,因此用copy()创建副本避免问题。- 合并逻辑:遍历每个对话条目,判断当前说话者是否与结果列表最后一条的说话者一致,一致则拼接
says和other字段,不一致则新增条目。
运行代码后,output.json中的内容即为你需要的合并结果。
内容的提问来源于stack exchange,提问作者user19023586
相关产品推荐
相关产品推荐

