如何用NLTK实现句子语义相似性匹配、替换及频率统计?
Absolutely! While NLTK alone isn’t a one-stop solution for deep semantic similarity, it excels at core text preprocessing, and when paired with complementary NLP tools, you can easily group semantically identical sentences, replace them with a representative, and count their frequencies. Let’s break down how to approach this:
First, normalize your text to reduce noise (like case differences, stopwords, or irrelevant tokens) so similarity models focus on meaningful content. NLTK has all the tools you need for this:
import nltk from nltk.tokenize import word_tokenize from nltk.corpus import stopwords from nltk.stem import WordNetLemmatizer # Download required NLTK resources nltk.download(['punkt', 'stopwords', 'wordnet']) def preprocess_sentence(sentence): # Lowercase and tokenize tokens = word_tokenize(sentence.lower()) # Remove stopwords and non-alphanumeric tokens stop_words = set(stopwords.words('english')) filtered_tokens = [token for token in tokens if token.isalnum() and token not in stop_words] # Lemmatize (reduce words to their base form) lemmatizer = WordNetLemmatizer() lemmatized_tokens = [lemmatizer.lemmatize(token) for token in filtered_tokens] return ' '.join(lemmatized_tokens) # Example input sentences raw_sentences = [ "I love eating pepperoni pizza", "Eating pepperoni pizza is my favorite", "Pepperoni pizza is what I enjoy most", "The ocean waves are calming", "Calming is how the ocean waves feel" ] # Preprocess all sentences preprocessed_sentences = [preprocess_sentence(s) for s in raw_sentences]
Next, pick a method to measure how similar sentences are. Here are the most practical options, with NLTK integration:
Option 1: Surface-Level Duplicates (Exact/Near-Exact)
For sentences that are almost identical (typos, word order swaps), use NLTK’s edit_distance to catch near-matches:
from nltk.metrics.distance import edit_distance def is_similar(s1, s2, threshold=3): return edit_distance(s1, s2) <= threshold # Group exact/near-exact duplicates groups = [] visited = [False] * len(raw_sentences) for i in range(len(raw_sentences)): if not visited[i]: group = [j for j in range(len(raw_sentences)) if is_similar(preprocessed_sentences[i], preprocessed_sentences[j])] for idx in group: visited[idx] = True groups.append(group)
Option 2: Bag-of-Words + Cosine Similarity
For basic semantic overlap (good for short sentences), combine NLTK preprocessing with scikit-learn’s cosine similarity:
from sklearn.feature_extraction.text import CountVectorizer from sklearn.metrics.pairwise import cosine_similarity # Convert preprocessed sentences to numerical vectors vectorizer = CountVectorizer() bow_vectors = vectorizer.fit_transform(preprocessed_sentences) # Calculate similarity between every pair of sentences similarity_matrix = cosine_similarity(bow_vectors) # Group sentences with similarity above a threshold (adjust based on your data) threshold = 0.6 groups = [] visited = [False] * len(raw_sentences) for i in range(len(raw_sentences)): if not visited[i]: group = [j for j in range(len(raw_sentences)) if similarity_matrix[i][j] >= threshold] for idx in group: visited[idx] = True groups.append(group)
Option 3: Sentence Embeddings (Best for True Semantics)
For deep semantic matching (e.g., "I need a coffee" vs. "Can I get a cup of coffee"), use pre-trained sentence embeddings. NLTK handles preprocessing, and libraries like sentence-transformers provide state-of-the-art embeddings:
from sentence_transformers import SentenceTransformer from sklearn.metrics.pairwise import cosine_similarity # Load pre-trained embedding model model = SentenceTransformer('all-MiniLM-L6-v2') # Generate embeddings for raw sentences (preprocessing is optional here) sentence_embeddings = model.encode(raw_sentences) # Compute similarity and group similarity_matrix = cosine_similarity(sentence_embeddings) threshold = 0.75 groups = [] visited = [False] * len(raw_sentences) for i in range(len(raw_sentences)): if not visited[i]: group = [j for j in range(len(raw_sentences)) if similarity_matrix[i][j] >= threshold] for idx in group: visited[idx] = True groups.append(group)
Once you have your groups, assign a representative sentence (e.g., the first occurrence or most common one) and count how many sentences fall into each group:
# Create a result dictionary: representative sentence -> frequency result = {} for group in groups: # Use the first sentence in the group as the representative representative = raw_sentences[group[0]] result[representative] = len(group) # Optional: Replace all sentences in each group with the representative normalized_sentences = [] for i in range(len(raw_sentences)): for group in groups: if i in group: normalized_sentences.append(raw_sentences[group[0]]) break print("Frequency Counts:") for sent, count in result.items(): print(f"- '{sent}': {count} times") print("\nNormalized Sentences:") print(normalized_sentences)
- Threshold Tuning: Adjust the similarity threshold based on your data—higher values mean stricter matching.
- NLTK Limitations: NLTK doesn’t have built-in deep semantic models, but it’s a great foundation for preprocessing before using more advanced tools like
sentence-transformersor spaCy. - Representative Selection: Instead of the first sentence, you could pick the most frequent one (if duplicates exist) or the shortest sentence for readability.
内容的提问来源于stack exchange,提问作者pyception

