You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何为字典内列表添加值?代码问题修正求助

解决按参与者分组单词的错误问题

嘿,我瞅见你遇到的问题了——不同参与者的单词被混进同一个列表里了对吧?咱们来拆解下问题出在哪,再给你调整代码。

你的原始代码与问题

先看你写的代码:

words = [ {'word': 'we', 'start_time': 90, 'participant': 'str_MIC_Y6E6_con_VnGhveZbaS'}, {'word': "haven't", 'start_time': 91, 'participant': 'str_MIC_Y6E6_con_VnGhveZbaS'}, {'word': 'even', 'start_time': 91, 'participant': 'str_MIC_Y6E6_con_VnGhveZbaS'}, {'word': 'spoken', 'start_time': 91, 'participant': 'str_MIC_Y6E6_con_VnGhveZbaS'}, {'word': 'about', 'start_time': 92, 'participant': 'str_MIC_Y6E6_con_VnGhveZbaS'}, {'word': 'your', 'start_time': 92, 'participant': 'str_MIC_Y6E6_con_VnGhveZbaS'}, {'word': 'newest', 'start_time': 92, 'participant': 'str_MIC_Y6E6_con_VnGhveZbaS'}, {'word': 'some word here', 'start_time': 45, 'participant': 'other user'} ]
words.sort(key=lambda x: x['start_time'])
clean_transcript = []
wordChunk = {'participant': '', 'words': []}
for w in words:
    if wordChunk['participant'] == w['participant']:
        wordChunk['words'].append(w['word'])
    else:
        wordChunk['participant'] = w['participant']
        print(wordChunk['participant'])
        wordChunk['words'].append(w['word'])
        clean_transcript.append(wordChunk)

运行后得到的错误结果是:

[{'participant': 'str_MIC_Y6E6_con_VnGhveZbaS', 'words': ['some word here', 'we', "haven't", 'even', 'spoken', 'about', 'your', 'newest']}]

问题根源

问题出在你一直在复用同一个wordChunk字典对象。当你把wordChunk添加到clean_transcript列表后,后续对wordChunk的任何修改(比如切换参与者、追加单词)都会影响列表里已经存在的那个字典——因为它们指向的是内存里的同一个对象,不是独立的副本。

举个简单的例子:你把一个苹果放进篮子,然后把这个苹果涂成红色,篮子里的苹果自然也变成红色了,因为是同一个苹果。

解决方案一:每次切换参与者时创建新的字典

我们调整循环逻辑,每次遇到新参与者时,先把之前的参与者数据存入列表,再创建一个全新的wordChunk字典。另外,循环结束后别忘了把最后一个参与者的数据也加进去:

words = [ {'word': 'we', 'start_time': 90, 'participant': 'str_MIC_Y6E6_con_VnGhveZbaS'}, {'word': "haven't", 'start_time': 91, 'participant': 'str_MIC_Y6E6_con_VnGhveZbaS'}, {'word': 'even', 'start_time': 91, 'participant': 'str_MIC_Y6E6_con_VnGhveZbaS'}, {'word': 'spoken', 'start_time': 91, 'participant': 'str_MIC_Y6E6_con_VnGhveZbaS'}, {'word': 'about', 'start_time': 92, 'participant': 'str_MIC_Y6E6_con_VnGhveZbaS'}, {'word': 'your', 'start_time': 92, 'participant': 'str_MIC_Y6E6_con_VnGhveZbaS'}, {'word': 'newest', 'start_time': 92, 'participant': 'str_MIC_Y6E6_con_VnGhveZbaS'}, {'word': 'some word here', 'start_time': 45, 'participant': 'other user'} ]
words.sort(key=lambda x: x['start_time'])
clean_transcript = []
wordChunk = None  # 初始化为None,表示还未处理任何参与者

for w in words:
    if wordChunk is None:
        # 第一次处理,创建第一个参与者的chunk
        wordChunk = {'participant': w['participant'], 'words': [w['word']]}
    elif wordChunk['participant'] == w['participant']:
        # 同一参与者,直接追加单词
        wordChunk['words'].append(w['word'])
    else:
        # 不同参与者,先把之前的chunk存入列表,再创建新的
        clean_transcript.append(wordChunk)
        wordChunk = {'participant': w['participant'], 'words': [w['word']]}

# 循环结束后,把最后一个参与者的chunk加入列表
if wordChunk is not None:
    clean_transcript.append(wordChunk)

print(clean_transcript)

运行后会得到正确的结果:

[
    {'participant': 'other user', 'words': ['some word here']},
    {'participant': 'str_MIC_Y6E6_con_VnGhveZbaS', 'words': ['we', "haven't", 'even', 'spoken', 'about', 'your', 'newest']}
]

解决方案二:用字典先分组(更简洁直观)

如果你觉得上面的逻辑有点绕,还可以用一个字典先按参与者分组单词,最后再转换成你需要的列表格式,代码更清晰:

words = [ {'word': 'we', 'start_time': 90, 'participant': 'str_MIC_Y6E6_con_VnGhveZbaS'}, {'word': "haven't", 'start_time': 91, 'participant': 'str_MIC_Y6E6_con_VnGhveZbaS'}, {'word': 'even', 'start_time': 91, 'participant': 'str_MIC_Y6E6_con_VnGhveZbaS'}, {'word': 'spoken', 'start_time': 91, 'participant': 'str_MIC_Y6E6_con_VnGhveZbaS'}, {'word': 'about', 'start_time': 92, 'participant': 'str_MIC_Y6E6_con_VnGhveZbaS'}, {'word': 'your', 'start_time': 92, 'participant': 'str_MIC_Y6E6_con_VnGhveZbaS'}, {'word': 'newest', 'start_time': 92, 'participant': 'str_MIC_Y6E6_con_VnGhveZbaS'}, {'word': 'some word here', 'start_time': 45, 'participant': 'other user'} ]
words.sort(key=lambda x: x['start_time'])

# 用字典按参与者分组单词
participant_groups = {}
for w in words:
    participant = w['participant']
    if participant not in participant_groups:
        participant_groups[participant] = []
    participant_groups[participant].append(w['word'])

# 转换成目标列表格式
clean_transcript = [{'participant': p, 'words': ws} for p, ws in participant_groups.items()]

print(clean_transcript)

这个方法的逻辑更直接:先把所有单词按参与者归类到字典里,再把字典转换成你需要的clean_transcript结构,结果同样正确。

内容的提问来源于stack exchange,提问作者lr_optim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 21:57:40