You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从指定的文本列表生成bi-grams(二元语法)?

如何从给定文本列表生成二元语法(Bi-grams)

二元语法指的是文本序列中连续两个相邻元素组成的成对序列。针对你提供的文本列表,生成二元语法的核心逻辑就是依次取列表中第i个元素与第i+1个元素配对,直到倒数第二个元素为止。

手动生成结果

给定列表 text = ['to','be',',','or','not','to','be',',','that','is','the','question',':'],生成的二元语法如下:

  • ('to', 'be')
  • ('be', ',')
  • (',', 'or')
  • ('or', 'not')
  • ('not', 'to')
  • ('to', 'be')
  • ('be', ',')
  • (',', 'that')
  • ('that', 'is')
  • ('is', 'the')
  • ('the', 'question')
  • ('question', ':')

Python代码实现

如果需要用代码批量生成,有两种简洁的方式:

方法1:循环遍历

text = ['to','be',',','or','not','to','be',',','that','is','the','question',':']
bigrams = []
# 遍历到倒数第二个元素即可
for i in range(len(text) - 1):
    bigrams.append((text[i], text[i+1]))
print(bigrams)

方法2:使用zip函数(更简洁)

text = ['to','be',',','or','not','to','be',',','that','is','the','question',':']
# 将原列表与原列表从第二个元素开始的切片配对
bigrams = list(zip(text, text[1:]))
print(bigrams)

两种方法最终输出的结果与手动生成的一致。

内容的提问来源于stack exchange,提问作者zari

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 23:35:42