You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修正字典Bigram生成结果:消除嵌套列表,得到标准二元组列表

修正字典生成Bigram的代码,实现扁平二元组列表

问题描述

需要从指定字典生成Bigram,仅对包含多个单词的短语生成,要求最终结果是直接包含二元组的扁平列表,但现有代码生成了嵌套列表结构,不符合预期。

示例输入字典:

dict_example = {'keywords1': ['africa',
  'basic service',
  'class',
  'develop country',
  'disadvantage',
  'economic resource',
  'social protection system']}

现有代码的错误输出(嵌套列表):

'keywords1': [[('basic', 'service')],
  [('develop', 'country')],
  [('economic', 'resource')],
  [('social', 'protection'), ('protection', 'system')]]

期望输出(扁平二元组列表):

'keywords1': [('basic', 'service'),
  ('develop', 'country'),
  ('economic', 'resource'),
  ('social', 'protection'),
  ('protection', 'system')]

错误原因分析

原代码存在两个核心问题:

  1. 使用嵌套列表推导式[list(nltk.bigrams(v3)) for v3 in v2 if len(v3) > 1],会为每个多词短语生成独立子列表,导致最终结果嵌套
  2. 冗余的keywords_bigram_temp字典和内层循环,增加了不必要的逻辑复杂度

修正后的代码

import nltk
from nltk.tokenize import word_tokenize

dict_example = {'keywords1': ['africa',
  'basic service',
  'class',
  'develop country',
  'disadvantage',
  'economic resource',
  'social protection system']}

keywords_bigram = {}
for key, phrases in dict_example.items():
    bigrams_flat = []
    for phrase in phrases:
        tokens = word_tokenize(phrase)
        # 仅处理包含多个单词的短语
        if len(tokens) > 1:
            # 用extend将bigrams迭代器的元素直接追加到总列表,避免嵌套
            bigrams_flat.extend(nltk.bigrams(tokens))
    keywords_bigram[key] = bigrams_flat

关键修改点

  • 移除冗余临时字典,直接在主循环中处理每个短语
  • 使用extend()替代嵌套列表推导式:nltk.bigrams()返回迭代器,extend()会将迭代器中的每个二元组逐个添加到总列表,实现扁平化结构
  • 简化循环逻辑,直接遍历原字典的键值对,避免不必要的内层遍历

运行结果

执行修正后的代码,输出完全符合预期:

print(keywords_bigram)
# {'keywords1': [('basic', 'service'), ('develop', 'country'), ('economic', 'resource'), ('social', 'protection'), ('protection', 'system')]}

内容的提问来源于stack exchange,提问作者mhutagalung

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 20:36:27