Python:如何在现有字典中添加嵌套统计字典且避免嵌套循环?
问题描述
给定如下Python字典:
Dialogues = { "dialogue_id": "000001", "dialogue_turns": [ { "turn_number": 0, "interlocutor_id": "0001", "turn_text": "Hi, how are you?" }, { "turn_number": 1, "interlocutor_id": "0002", "turn_text": "Hi, I'm fine thanks. And you?" }, { "turn_number": 2, "interlocutor_id": "0001", "turn_text": "I am good too, are you coming to the class today" }, { "turn_number": 3, "interlocutor_id": "0002", "turn_text": "Yes, see you soon.bye" }, { "turn_number": 4, "interlocutor_id": "0001", "turn_text": "bye" } ] }
需要生成如下语法正确的Dialogues_analyzed字典,添加嵌套统计信息:
Dialogues_analyzed = { "dialogue_id": "000001", "dialogue_analysis": [ {"interlocutor_id": "0001", "total_turns": "3", "total_words": "该用户所有对话的总词数"}, {"interlocutor_id": "0002", "total_turns": "2", "total_words": "该用户所有对话的总词数"} ] }
要求不使用嵌套for循环实现,此前尝试新建字典合并时出现键丢失问题。
解决方案
利用collections.defaultdict聚合对话者的统计数据,全程仅用单层循环,避免嵌套结构,同时确保键不丢失:
from collections import defaultdict # 原对话字典 Dialogues = { "dialogue_id": "000001", "dialogue_turns": [ { "turn_number": 0, "interlocutor_id": "0001", "turn_text": "Hi, how are you?" }, { "turn_number": 1, "interlocutor_id": "0002", "turn_text": "Hi, I'm fine thanks. And you?" }, { "turn_number": 2, "interlocutor_id": "0001", "turn_text": "I am good too, are you coming to the class today" }, { "turn_number": 3, "interlocutor_id": "0002", "turn_text": "Yes, see you soon.bye" }, { "turn_number": 4, "interlocutor_id": "0001", "turn_text": "bye" } ] } # 初始化统计容器:key为对话者ID,value存储轮次和单词数 stats = defaultdict(lambda: {"total_turns": 0, "total_words": 0}) # 单层循环遍历所有对话轮次 for turn in Dialogues["dialogue_turns"]: user_id = turn["interlocutor_id"] # 统计轮次 stats[user_id]["total_turns"] += 1 # 统计单词数(去除标点后拆分,可根据需求调整逻辑) cleaned_text = turn["turn_text"].translate(str.maketrans('', '', ',.?')) # 移除常见标点 stats[user_id]["total_words"] += len([word for word in cleaned_text.split() if word]) # 生成目标结构字典 Dialogues_analyzed = { "dialogue_id": Dialogues["dialogue_id"], "dialogue_analysis": [ { "interlocutor_id": uid, "total_turns": str(data["total_turns"]), "total_words": str(data["total_words"]) } for uid, data in stats.items() ] } # 输出结果 print(Dialogues_analyzed)
关键说明
- 无嵌套循环:仅使用一层遍历对话轮次的for循环,列表推导式也为单层结构,完全避免嵌套循环。
- 解决键丢失问题:直接继承原字典的
dialogue_id,用defaultdict自动为每个新对话者初始化统计项,确保不会遗漏键;最后通过列表推导式将聚合结果转为目标格式,结构清晰无丢失。 - 单词统计优化:用
str.maketrans高效移除标点,再拆分统计有效单词数,逻辑可根据实际需求(如支持多语言、特殊符号)调整。
内容的提问来源于stack exchange,提问作者Donya
相关产品推荐
相关产品推荐

