You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python:如何在现有字典中添加嵌套统计字典且避免嵌套循环?

问题描述

给定如下Python字典:

Dialogues = {
   "dialogue_id": "000001", 
   "dialogue_turns": [
      { "turn_number": 0,
        "interlocutor_id": "0001",
        "turn_text": "Hi, how are you?" },
      { "turn_number": 1,
        "interlocutor_id": "0002",
        "turn_text": "Hi, I'm fine thanks. And you?" },
      { "turn_number": 2,
        "interlocutor_id": "0001",
        "turn_text": "I am good too, are you coming to the class today" },
      { "turn_number": 3,
        "interlocutor_id": "0002",
        "turn_text": "Yes, see you soon.bye" },
      { "turn_number": 4,
        "interlocutor_id": "0001",
        "turn_text": "bye" }
   ]
}

需要生成如下语法正确的Dialogues_analyzed字典,添加嵌套统计信息:

Dialogues_analyzed = {
   "dialogue_id": "000001", 
   "dialogue_analysis": [
      {"interlocutor_id": "0001",
       "total_turns": "3",
       "total_words": "该用户所有对话的总词数"},
      {"interlocutor_id": "0002",
       "total_turns": "2",
       "total_words": "该用户所有对话的总词数"}
   ]
}

要求不使用嵌套for循环实现,此前尝试新建字典合并时出现键丢失问题。

解决方案

利用collections.defaultdict聚合对话者的统计数据,全程仅用单层循环,避免嵌套结构,同时确保键不丢失:

from collections import defaultdict

# 原对话字典
Dialogues = {
   "dialogue_id": "000001", 
   "dialogue_turns": [
      { "turn_number": 0, "interlocutor_id": "0001", "turn_text": "Hi, how are you?" },
      { "turn_number": 1, "interlocutor_id": "0002", "turn_text": "Hi, I'm fine thanks. And you?" },
      { "turn_number": 2, "interlocutor_id": "0001", "turn_text": "I am good too, are you coming to the class today" },
      { "turn_number": 3, "interlocutor_id": "0002", "turn_text": "Yes, see you soon.bye" },
      { "turn_number": 4, "interlocutor_id": "0001", "turn_text": "bye" }
   ]
}

# 初始化统计容器:key为对话者ID,value存储轮次和单词数
stats = defaultdict(lambda: {"total_turns": 0, "total_words": 0})

# 单层循环遍历所有对话轮次
for turn in Dialogues["dialogue_turns"]:
    user_id = turn["interlocutor_id"]
    # 统计轮次
    stats[user_id]["total_turns"] += 1
    # 统计单词数(去除标点后拆分,可根据需求调整逻辑)
    cleaned_text = turn["turn_text"].translate(str.maketrans('', '', ',.?'))  # 移除常见标点
    stats[user_id]["total_words"] += len([word for word in cleaned_text.split() if word])

# 生成目标结构字典
Dialogues_analyzed = {
    "dialogue_id": Dialogues["dialogue_id"],
    "dialogue_analysis": [
        {
            "interlocutor_id": uid,
            "total_turns": str(data["total_turns"]),
            "total_words": str(data["total_words"])
        }
        for uid, data in stats.items()
    ]
}

# 输出结果
print(Dialogues_analyzed)

关键说明

  1. 无嵌套循环:仅使用一层遍历对话轮次的for循环,列表推导式也为单层结构,完全避免嵌套循环。
  2. 解决键丢失问题:直接继承原字典的dialogue_id,用defaultdict自动为每个新对话者初始化统计项,确保不会遗漏键;最后通过列表推导式将聚合结果转为目标格式,结构清晰无丢失。
  3. 单词统计优化:用str.maketrans高效移除标点,再拆分统计有效单词数,逻辑可根据实际需求(如支持多语言、特殊符号)调整。

内容的提问来源于stack exchange,提问作者Donya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 09:15:34