You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用NLTK实现句子语义相似性匹配、替换及频率统计?

Absolutely! While NLTK alone isn’t a one-stop solution for deep semantic similarity, it excels at core text preprocessing, and when paired with complementary NLP tools, you can easily group semantically identical sentences, replace them with a representative, and count their frequencies. Let’s break down how to approach this:

Step 1: Preprocess Sentences with NLTK

First, normalize your text to reduce noise (like case differences, stopwords, or irrelevant tokens) so similarity models focus on meaningful content. NLTK has all the tools you need for this:

import nltk
from nltk.tokenize import word_tokenize
from nltk.corpus import stopwords
from nltk.stem import WordNetLemmatizer

# Download required NLTK resources
nltk.download(['punkt', 'stopwords', 'wordnet'])

def preprocess_sentence(sentence):
    # Lowercase and tokenize
    tokens = word_tokenize(sentence.lower())
    # Remove stopwords and non-alphanumeric tokens
    stop_words = set(stopwords.words('english'))
    filtered_tokens = [token for token in tokens if token.isalnum() and token not in stop_words]
    # Lemmatize (reduce words to their base form)
    lemmatizer = WordNetLemmatizer()
    lemmatized_tokens = [lemmatizer.lemmatize(token) for token in filtered_tokens]
    return ' '.join(lemmatized_tokens)

# Example input sentences
raw_sentences = [
    "I love eating pepperoni pizza",
    "Eating pepperoni pizza is my favorite",
    "Pepperoni pizza is what I enjoy most",
    "The ocean waves are calming",
    "Calming is how the ocean waves feel"
]

# Preprocess all sentences
preprocessed_sentences = [preprocess_sentence(s) for s in raw_sentences]
Step 2: Choose a Semantic Similarity Metric

Next, pick a method to measure how similar sentences are. Here are the most practical options, with NLTK integration:

Option 1: Surface-Level Duplicates (Exact/Near-Exact)

For sentences that are almost identical (typos, word order swaps), use NLTK’s edit_distance to catch near-matches:

from nltk.metrics.distance import edit_distance

def is_similar(s1, s2, threshold=3):
    return edit_distance(s1, s2) <= threshold

# Group exact/near-exact duplicates
groups = []
visited = [False] * len(raw_sentences)

for i in range(len(raw_sentences)):
    if not visited[i]:
        group = [j for j in range(len(raw_sentences)) if is_similar(preprocessed_sentences[i], preprocessed_sentences[j])]
        for idx in group:
            visited[idx] = True
        groups.append(group)

Option 2: Bag-of-Words + Cosine Similarity

For basic semantic overlap (good for short sentences), combine NLTK preprocessing with scikit-learn’s cosine similarity:

from sklearn.feature_extraction.text import CountVectorizer
from sklearn.metrics.pairwise import cosine_similarity

# Convert preprocessed sentences to numerical vectors
vectorizer = CountVectorizer()
bow_vectors = vectorizer.fit_transform(preprocessed_sentences)

# Calculate similarity between every pair of sentences
similarity_matrix = cosine_similarity(bow_vectors)

# Group sentences with similarity above a threshold (adjust based on your data)
threshold = 0.6
groups = []
visited = [False] * len(raw_sentences)

for i in range(len(raw_sentences)):
    if not visited[i]:
        group = [j for j in range(len(raw_sentences)) if similarity_matrix[i][j] >= threshold]
        for idx in group:
            visited[idx] = True
        groups.append(group)

Option 3: Sentence Embeddings (Best for True Semantics)

For deep semantic matching (e.g., "I need a coffee" vs. "Can I get a cup of coffee"), use pre-trained sentence embeddings. NLTK handles preprocessing, and libraries like sentence-transformers provide state-of-the-art embeddings:

from sentence_transformers import SentenceTransformer
from sklearn.metrics.pairwise import cosine_similarity

# Load pre-trained embedding model
model = SentenceTransformer('all-MiniLM-L6-v2')

# Generate embeddings for raw sentences (preprocessing is optional here)
sentence_embeddings = model.encode(raw_sentences)

# Compute similarity and group
similarity_matrix = cosine_similarity(sentence_embeddings)
threshold = 0.75
groups = []
visited = [False] * len(raw_sentences)

for i in range(len(raw_sentences)):
    if not visited[i]:
        group = [j for j in range(len(raw_sentences)) if similarity_matrix[i][j] >= threshold]
        for idx in group:
            visited[idx] = True
        groups.append(group)
Step 3: Count Frequencies & Replace Sentences

Once you have your groups, assign a representative sentence (e.g., the first occurrence or most common one) and count how many sentences fall into each group:

# Create a result dictionary: representative sentence -> frequency
result = {}

for group in groups:
    # Use the first sentence in the group as the representative
    representative = raw_sentences[group[0]]
    result[representative] = len(group)

# Optional: Replace all sentences in each group with the representative
normalized_sentences = []
for i in range(len(raw_sentences)):
    for group in groups:
        if i in group:
            normalized_sentences.append(raw_sentences[group[0]])
            break

print("Frequency Counts:")
for sent, count in result.items():
    print(f"- '{sent}': {count} times")

print("\nNormalized Sentences:")
print(normalized_sentences)
Key Notes
  • Threshold Tuning: Adjust the similarity threshold based on your data—higher values mean stricter matching.
  • NLTK Limitations: NLTK doesn’t have built-in deep semantic models, but it’s a great foundation for preprocessing before using more advanced tools like sentence-transformers or spaCy.
  • Representative Selection: Instead of the first sentence, you could pick the most frequent one (if duplicates exist) or the shortest sentence for readability.

内容的提问来源于stack exchange,提问作者pyception

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:13:01