如何修正字典Bigram生成结果:消除嵌套列表,得到标准二元组列表
修正字典生成Bigram的代码,实现扁平二元组列表
问题描述
需要从指定字典生成Bigram,仅对包含多个单词的短语生成,要求最终结果是直接包含二元组的扁平列表,但现有代码生成了嵌套列表结构,不符合预期。
示例输入字典:
dict_example = {'keywords1': ['africa', 'basic service', 'class', 'develop country', 'disadvantage', 'economic resource', 'social protection system']}
现有代码的错误输出(嵌套列表):
'keywords1': [[('basic', 'service')], [('develop', 'country')], [('economic', 'resource')], [('social', 'protection'), ('protection', 'system')]]
期望输出(扁平二元组列表):
'keywords1': [('basic', 'service'), ('develop', 'country'), ('economic', 'resource'), ('social', 'protection'), ('protection', 'system')]
错误原因分析
原代码存在两个核心问题:
- 使用嵌套列表推导式
[list(nltk.bigrams(v3)) for v3 in v2 if len(v3) > 1],会为每个多词短语生成独立子列表,导致最终结果嵌套 - 冗余的
keywords_bigram_temp字典和内层循环,增加了不必要的逻辑复杂度
修正后的代码
import nltk from nltk.tokenize import word_tokenize dict_example = {'keywords1': ['africa', 'basic service', 'class', 'develop country', 'disadvantage', 'economic resource', 'social protection system']} keywords_bigram = {} for key, phrases in dict_example.items(): bigrams_flat = [] for phrase in phrases: tokens = word_tokenize(phrase) # 仅处理包含多个单词的短语 if len(tokens) > 1: # 用extend将bigrams迭代器的元素直接追加到总列表,避免嵌套 bigrams_flat.extend(nltk.bigrams(tokens)) keywords_bigram[key] = bigrams_flat
关键修改点
- 移除冗余临时字典,直接在主循环中处理每个短语
- 使用
extend()替代嵌套列表推导式:nltk.bigrams()返回迭代器,extend()会将迭代器中的每个二元组逐个添加到总列表,实现扁平化结构 - 简化循环逻辑,直接遍历原字典的键值对,避免不必要的内层遍历
运行结果
执行修正后的代码,输出完全符合预期:
print(keywords_bigram) # {'keywords1': [('basic', 'service'), ('develop', 'country'), ('economic', 'resource'), ('social', 'protection'), ('protection', 'system')]}
内容的提问来源于stack exchange,提问作者mhutagalung
相关产品推荐
相关产品推荐

