如何提取多个句子中共有的非跨句子重复n元语法(n-grams)
Hey there! Let's fix that annoying issue where your code was spitting out cross-sentence n-grams—like project Web—that never actually appear in a single sentence. Those are totally useless for your auto-complete feature, so let's get rid of them for good.
The Root of the Problem
Your original code joined all sentences into one big string, which made everygrams combine the end of one sentence with the start of the next. That's where those invalid cross-sentence pairs came from. We need to ensure we only generate n-grams within individual sentences right from the start.
A Cleaner, Optimal Solution
Instead of adding end tokens and filtering them out later, let's handle each sentence independently. This way, we never create cross-sentence n-grams in the first place. Here's the improved code:
from nltk import everygrams from collections import Counter data = [ 'Design and architecture project', 'Web inquiry for products', 'Software for non-profit project', 'Web inquiry for vendors' ] # Collect n-grams from each sentence separately all_valid_ngrams = [] for sentence in data: tokens = sentence.split() # Generate 2-5 grams only for the current sentence sentence_ngrams = everygrams(tokens, min_len=2, max_len=5) all_valid_ngrams.extend(sentence_ngrams) # Count how many times each n-gram appears across all sentences ngram_counter = Counter(all_valid_ngrams) # Filter for repeated n-grams and format them as readable strings repeated_ngrams = [(' '.join(ngram), count) for ngram, count in ngram_counter.items() if count > 1] print(repeated_ngrams)
Output
[('Web inquiry', 2), ('Web inquiry for', 2), ('inquiry for', 2)]
Why This Beats Your End-Token Approach
- No invalid n-grams ever: We only generate n-grams within each sentence, so cross-sentence combinations never exist to begin with.
- Simpler code: No need to add special tokens or write extra logic to filter them out—everything is straightforward and easy to maintain.
- More efficient: Avoids joining all sentences into a single large string and processing unnecessary tokens (like your
<end>markers).
This approach gives you exactly the repeated in-sentence n-grams you need for reliable auto-complete suggestions—no more false positives like suggesting "web" after "project"!
内容的提问来源于stack exchange,提问作者Dmitry Volkov

