技术咨询:Python环境下如何从语料库中提取Top 100 N-gram、Bigram及Trigram?
Hey there! Let’s work through this problem together—you need to extract the top 100 multi-granularity terms (unigrams, bigrams, trigrams) from your 1500 product reviews, and the tools you’ve tried so far either generate every possible ngram or only handle one length at a time. Here’s a practical, Python-based solution that fixes both issues:
Before generating ngrams, clean up your text to eliminate noise (like stopwords, punctuation, or irrelevant tokens). This ensures we only work with meaningful terms.
import nltk from nltk.corpus import stopwords from nltk.tokenize import word_tokenize from collections import Counter from sklearn.feature_extraction.text import TfidfVectorizer # Download required NLTK resources (run once) nltk.download('punkt') nltk.download('stopwords') # Initialize stopwords (adjust for your language if needed) stop_words = set(stopwords.words('english')) def preprocess_text(text): # Tokenize and convert to lowercase tokens = word_tokenize(text.lower()) # Keep only alphabetic tokens that aren't stopwords filtered_tokens = [token for token in tokens if token.isalpha() and token not in stop_words] return filtered_tokens # Replace this with your actual list of 1500 reviews your_reviews = [ "This wireless headphone has amazing sound quality and long battery life", "The screen resolution of this tablet is crisp, but it gets hot during gaming", # ... add all your reviews here ] # Preprocess all reviews processed_reviews = [preprocess_text(review) for review in your_reviews]
We’ll create unigrams, bigrams, and trigrams from the cleaned tokens, then combine them into a single list. This avoids the "single granularity" limitation you faced.
def generate_ngrams(tokens, n): # Create ngrams by sliding a window of size n over the tokens return [' '.join(tokens[i:i+n]) for i in range(len(tokens) - n + 1)] # Collect all ngrams across all reviews all_ngrams = [] for tokens in processed_reviews: # Add unigrams (single words) all_ngrams.extend(tokens) # Add bigrams (two-word combinations) all_ngrams.extend(generate_ngrams(tokens, 2)) # Add trigrams (three-word combinations) all_ngrams.extend(generate_ngrams(tokens, 3))
Now you have two solid options to rank terms—use raw frequency (for most common terms) or TF-IDF (for terms that are important across the corpus, not just frequent in a few reviews).
Option 1: Rank by Raw Frequency
Great if you want the most commonly mentioned terms:
# Count occurrences of each ngram ngram_frequencies = Counter(all_ngrams) # Get top 100 terms top_100_frequency = ngram_frequencies.most_common(100) # Print results print("Top 100 Terms by Frequency:") for term, count in top_100_frequency: print(f"{term}: {count}")
Option 2: Rank by TF-IDF
Better if you want terms that are distinctive (e.g., a term that’s frequent in reviews of a specific product feature, not just "good" or "great"):
# Convert processed tokens back to full text strings for TF-IDF corpus = [' '.join(tokens) for tokens in processed_reviews] # Configure TF-IDF to handle unigrams, bigrams, and trigrams in one go tfidf_vectorizer = TfidfVectorizer(ngram_range=(1, 3)) tfidf_matrix = tfidf_vectorizer.fit_transform(corpus) # Map each term to its total TF-IDF score across all reviews term_tfidf_scores = dict( zip(tfidf_vectorizer.get_feature_names_out(), tfidf_matrix.sum(axis=0).A1) ) # Sort terms by score (highest first) and take top 100 top_100_tfidf = sorted(term_tfidf_scores.items(), key=lambda x: x[1], reverse=True)[:100] # Print results print("\nTop 100 Terms by TF-IDF:") for term, score in top_100_tfidf: print(f"{term}: {score:.4f}")
If you want to focus on product-specific terms (like "battery life" or "camera resolution"), use POS tagging to keep only nouns, adjectives, and relevant verbs. Here’s how to add that with spaCy:
import spacy # Load spaCy's English model (run `pip install spacy` and `python -m spacy download en_core_web_sm` first) nlp = spacy.load('en_core_web_sm') def preprocess_with_pos(text): doc = nlp(text.lower()) # Keep only nouns, adjectives, and verbs (adjust POS tags as needed) filtered_tokens = [ token.text for token in doc if token.pos_ in ['NOUN', 'ADJ', 'VERB'] and token.text not in stop_words ] return filtered_tokens # Re-process reviews with POS filtering processed_reviews_pos = [preprocess_with_pos(review) for review in your_reviews] # Repeat steps 2 and 3 with processed_reviews_pos instead of processed_reviews
This will filter out meaningless combinations and give you more actionable product terms.
内容的提问来源于stack exchange,提问作者chefnewman

