如何用Python查找间隔N个token的非连续二元组(bi-grams)?
嘿,我懂你要的是什么——不是那种紧挨着的连续二元组,而是允许两个token之间隔N个其他词的非连续二元组对吧?CountVectorizer确实只支持连续的ngram,所以得换个思路来实现,我给你几个实用的方案:
解决方案:手动生成非连续二元组
CountVectorizer的设计就是针对连续的n元组,所以要实现非连续的,最直接的方式是手动遍历分词后的token列表,生成所有符合间隔要求的二元组。
1. 基础实现(自定义间隔)
先把文本拆成token,然后通过双重循环生成所有满足“最多间隔N个token”的二元组:
sentence = 'i love coding with python' tokens = sentence.split() # 简单分词,复杂场景可以用专业分词工具 # 定义最大允许的间隔token数:比如这里设为3,即两个token之间最多隔3个词 max_allowed_gap = 3 non_contiguous_bigrams = [] for idx, token in enumerate(tokens): # 遍历当前token之后的所有token,只要间隔不超过max_allowed_gap # 索引差最多为 max_allowed_gap + 1(比如间隔0个是差1,间隔3个是差4) for next_idx in range(idx + 1, min(idx + max_allowed_gap + 1, len(tokens))): non_contiguous_bigrams.append((token, tokens[next_idx])) print(non_contiguous_bigrams)
运行后输出就是你预期的结果:
[('i', 'love'), ('i', 'coding'), ('i', 'with'), ('i', 'python'), ('love', 'coding'), ('love', 'with'), ('love', 'python'), ('coding', 'with'), ('coding', 'python'), ('with', 'python')]
2. 统计二元组频率
如果需要统计每个非连续二元组的出现次数,可以用collections.Counter:
from collections import Counter bigram_frequency = Counter(non_contiguous_bigrams) print(bigram_frequency) # 输出:Counter({('i', 'love'): 1, ('i', 'coding'): 1, ('i', 'with'): 1, ('i', 'python'): 1, ('love', 'coding'): 1, ('love', 'with'): 1, ('love', 'python'): 1, ('coding', 'with'): 1, ('coding', 'python'): 1, ('with', 'python'): 1})
3. 适配复杂分词场景
如果你的文本包含标点、大小写或者需要更精准的分词,可以用NLTK或SpaCy的分词工具替代str.split():
用NLTK分词:
from nltk.tokenize import word_tokenize sentence = 'I love coding with Python!' tokens = word_tokenize(sentence) # 之后重复上面的二元组生成逻辑
用SpaCy分词:
import spacy nlp = spacy.load("en_core_web_sm") doc = nlp('I love coding with Python!') tokens = [token.text for token in doc] # 之后重复二元组生成逻辑
4. 封装成可复用函数
把逻辑封装成函数,方便在多个地方调用:
from collections import Counter def generate_non_contiguous_bigrams(text, max_gap=3, tokenizer=str.split): """ 生成文本中的非连续二元组,允许两个token之间最多间隔max_gap个token :param text: 输入文本 :param max_gap: 最大允许间隔的token数,默认3 :param tokenizer: 分词函数,默认用str.split() :return: 二元组列表 """ tokens = tokenizer(text) bigrams = [] for idx in range(len(tokens)): end_idx = min(idx + max_gap + 1, len(tokens)) for next_idx in range(idx + 1, end_idx): bigrams.append((tokens[idx], tokens[next_idx])) return bigrams def count_non_contiguous_bigrams(text, max_gap=3, tokenizer=str.split): """统计非连续二元组的频率""" bigrams = generate_non_contiguous_bigrams(text, max_gap, tokenizer) return Counter(bigrams) # 使用示例 sentence = 'i love coding with python' print(generate_non_contiguous_bigrams(sentence, max_gap=3)) print(count_non_contiguous_bigrams(sentence, max_gap=3))
内容的提问来源于stack exchange,提问作者Steve
相关产品推荐
相关产品推荐

