You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python查找间隔N个token的非连续二元组(bi-grams)?

嘿,我懂你要的是什么——不是那种紧挨着的连续二元组,而是允许两个token之间隔N个其他词的非连续二元组对吧?CountVectorizer确实只支持连续的ngram,所以得换个思路来实现,我给你几个实用的方案:

解决方案:手动生成非连续二元组

CountVectorizer的设计就是针对连续的n元组,所以要实现非连续的,最直接的方式是手动遍历分词后的token列表,生成所有符合间隔要求的二元组。

1. 基础实现(自定义间隔)

先把文本拆成token,然后通过双重循环生成所有满足“最多间隔N个token”的二元组:

sentence = 'i love coding with python'
tokens = sentence.split()  # 简单分词,复杂场景可以用专业分词工具

# 定义最大允许的间隔token数:比如这里设为3,即两个token之间最多隔3个词
max_allowed_gap = 3

non_contiguous_bigrams = []
for idx, token in enumerate(tokens):
    # 遍历当前token之后的所有token,只要间隔不超过max_allowed_gap
    # 索引差最多为 max_allowed_gap + 1(比如间隔0个是差1,间隔3个是差4)
    for next_idx in range(idx + 1, min(idx + max_allowed_gap + 1, len(tokens))):
        non_contiguous_bigrams.append((token, tokens[next_idx]))

print(non_contiguous_bigrams)

运行后输出就是你预期的结果:

[('i', 'love'), ('i', 'coding'), ('i', 'with'), ('i', 'python'), ('love', 'coding'), ('love', 'with'), ('love', 'python'), ('coding', 'with'), ('coding', 'python'), ('with', 'python')]

2. 统计二元组频率

如果需要统计每个非连续二元组的出现次数,可以用collections.Counter:

from collections import Counter

bigram_frequency = Counter(non_contiguous_bigrams)
print(bigram_frequency)
# 输出:Counter({('i', 'love'): 1, ('i', 'coding'): 1, ('i', 'with'): 1, ('i', 'python'): 1, ('love', 'coding'): 1, ('love', 'with'): 1, ('love', 'python'): 1, ('coding', 'with'): 1, ('coding', 'python'): 1, ('with', 'python'): 1})

3. 适配复杂分词场景

如果你的文本包含标点、大小写或者需要更精准的分词,可以用NLTK或SpaCy的分词工具替代str.split():

用NLTK分词:

from nltk.tokenize import word_tokenize
sentence = 'I love coding with Python!'
tokens = word_tokenize(sentence)
# 之后重复上面的二元组生成逻辑

用SpaCy分词:

import spacy
nlp = spacy.load("en_core_web_sm")
doc = nlp('I love coding with Python!')
tokens = [token.text for token in doc]
# 之后重复二元组生成逻辑

4. 封装成可复用函数

把逻辑封装成函数,方便在多个地方调用:

from collections import Counter

def generate_non_contiguous_bigrams(text, max_gap=3, tokenizer=str.split):
    """
    生成文本中的非连续二元组,允许两个token之间最多间隔max_gap个token
    :param text: 输入文本
    :param max_gap: 最大允许间隔的token数,默认3
    :param tokenizer: 分词函数,默认用str.split()
    :return: 二元组列表
    """
    tokens = tokenizer(text)
    bigrams = []
    for idx in range(len(tokens)):
        end_idx = min(idx + max_gap + 1, len(tokens))
        for next_idx in range(idx + 1, end_idx):
            bigrams.append((tokens[idx], tokens[next_idx]))
    return bigrams

def count_non_contiguous_bigrams(text, max_gap=3, tokenizer=str.split):
    """统计非连续二元组的频率"""
    bigrams = generate_non_contiguous_bigrams(text, max_gap, tokenizer)
    return Counter(bigrams)

# 使用示例
sentence = 'i love coding with python'
print(generate_non_contiguous_bigrams(sentence, max_gap=3))
print(count_non_contiguous_bigrams(sentence, max_gap=3))

内容的提问来源于stack exchange,提问作者Steve

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 03:56:30