You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取多个句子中共有的非跨句子重复n元语法(n-grams)

Extract Valid In-Sentence Repeated N-Grams for Auto-Complete

Hey there! Let's fix that annoying issue where your code was spitting out cross-sentence n-grams—like project Web—that never actually appear in a single sentence. Those are totally useless for your auto-complete feature, so let's get rid of them for good.

The Root of the Problem

Your original code joined all sentences into one big string, which made everygrams combine the end of one sentence with the start of the next. That's where those invalid cross-sentence pairs came from. We need to ensure we only generate n-grams within individual sentences right from the start.

A Cleaner, Optimal Solution

Instead of adding end tokens and filtering them out later, let's handle each sentence independently. This way, we never create cross-sentence n-grams in the first place. Here's the improved code:

from nltk import everygrams
from collections import Counter

data = [ 'Design and architecture project', 'Web inquiry for products', 'Software for non-profit project', 'Web inquiry for vendors' ]

# Collect n-grams from each sentence separately
all_valid_ngrams = []
for sentence in data:
    tokens = sentence.split()
    # Generate 2-5 grams only for the current sentence
    sentence_ngrams = everygrams(tokens, min_len=2, max_len=5)
    all_valid_ngrams.extend(sentence_ngrams)

# Count how many times each n-gram appears across all sentences
ngram_counter = Counter(all_valid_ngrams)

# Filter for repeated n-grams and format them as readable strings
repeated_ngrams = [(' '.join(ngram), count) for ngram, count in ngram_counter.items() if count > 1]

print(repeated_ngrams)

Output

[('Web inquiry', 2), ('Web inquiry for', 2), ('inquiry for', 2)]

Why This Beats Your End-Token Approach

  • No invalid n-grams ever: We only generate n-grams within each sentence, so cross-sentence combinations never exist to begin with.
  • Simpler code: No need to add special tokens or write extra logic to filter them out—everything is straightforward and easy to maintain.
  • More efficient: Avoids joining all sentences into a single large string and processing unnecessary tokens (like your <end> markers).

This approach gives you exactly the repeated in-sentence n-grams you need for reliable auto-complete suggestions—no more false positives like suggesting "web" after "project"!

内容的提问来源于stack exchange,提问作者Dmitry Volkov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 11:24:07