如何从指定的文本列表生成bi-grams(二元语法)?
如何从给定文本列表生成二元语法(Bi-grams)
二元语法指的是文本序列中连续两个相邻元素组成的成对序列。针对你提供的文本列表,生成二元语法的核心逻辑就是依次取列表中第i个元素与第i+1个元素配对,直到倒数第二个元素为止。
手动生成结果
给定列表 text = ['to','be',',','or','not','to','be',',','that','is','the','question',':'],生成的二元语法如下:
- ('to', 'be')
- ('be', ',')
- (',', 'or')
- ('or', 'not')
- ('not', 'to')
- ('to', 'be')
- ('be', ',')
- (',', 'that')
- ('that', 'is')
- ('is', 'the')
- ('the', 'question')
- ('question', ':')
Python代码实现
如果需要用代码批量生成,有两种简洁的方式:
方法1:循环遍历
text = ['to','be',',','or','not','to','be',',','that','is','the','question',':'] bigrams = [] # 遍历到倒数第二个元素即可 for i in range(len(text) - 1): bigrams.append((text[i], text[i+1])) print(bigrams)
方法2:使用zip函数(更简洁)
text = ['to','be',',','or','not','to','be',',','that','is','the','question',':'] # 将原列表与原列表从第二个元素开始的切片配对 bigrams = list(zip(text, text[1:])) print(bigrams)
两种方法最终输出的结果与手动生成的一致。
内容的提问来源于stack exchange,提问作者zari
相关产品推荐
相关产品推荐

