You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

技术咨询:Python环境下如何从语料库中提取Top 100 N-gram、Bigram及Trigram?

Hey there! Let’s work through this problem together—you need to extract the top 100 multi-granularity terms (unigrams, bigrams, trigrams) from your 1500 product reviews, and the tools you’ve tried so far either generate every possible ngram or only handle one length at a time. Here’s a practical, Python-based solution that fixes both issues:

Step 1: Preprocess Your Reviews First

Before generating ngrams, clean up your text to eliminate noise (like stopwords, punctuation, or irrelevant tokens). This ensures we only work with meaningful terms.

import nltk
from nltk.corpus import stopwords
from nltk.tokenize import word_tokenize
from collections import Counter
from sklearn.feature_extraction.text import TfidfVectorizer

# Download required NLTK resources (run once)
nltk.download('punkt')
nltk.download('stopwords')

# Initialize stopwords (adjust for your language if needed)
stop_words = set(stopwords.words('english'))

def preprocess_text(text):
    # Tokenize and convert to lowercase
    tokens = word_tokenize(text.lower())
    # Keep only alphabetic tokens that aren't stopwords
    filtered_tokens = [token for token in tokens if token.isalpha() and token not in stop_words]
    return filtered_tokens

# Replace this with your actual list of 1500 reviews
your_reviews = [
    "This wireless headphone has amazing sound quality and long battery life",
    "The screen resolution of this tablet is crisp, but it gets hot during gaming",
    # ... add all your reviews here
]

# Preprocess all reviews
processed_reviews = [preprocess_text(review) for review in your_reviews]
Step 2: Generate Multi-Granularity Ngrams

We’ll create unigrams, bigrams, and trigrams from the cleaned tokens, then combine them into a single list. This avoids the "single granularity" limitation you faced.

def generate_ngrams(tokens, n):
    # Create ngrams by sliding a window of size n over the tokens
    return [' '.join(tokens[i:i+n]) for i in range(len(tokens) - n + 1)]

# Collect all ngrams across all reviews
all_ngrams = []
for tokens in processed_reviews:
    # Add unigrams (single words)
    all_ngrams.extend(tokens)
    # Add bigrams (two-word combinations)
    all_ngrams.extend(generate_ngrams(tokens, 2))
    # Add trigrams (three-word combinations)
    all_ngrams.extend(generate_ngrams(tokens, 3))
Step 3: Rank Terms & Get Top 100

Now you have two solid options to rank terms—use raw frequency (for most common terms) or TF-IDF (for terms that are important across the corpus, not just frequent in a few reviews).

Option 1: Rank by Raw Frequency

Great if you want the most commonly mentioned terms:

# Count occurrences of each ngram
ngram_frequencies = Counter(all_ngrams)
# Get top 100 terms
top_100_frequency = ngram_frequencies.most_common(100)

# Print results
print("Top 100 Terms by Frequency:")
for term, count in top_100_frequency:
    print(f"{term}: {count}")

Option 2: Rank by TF-IDF

Better if you want terms that are distinctive (e.g., a term that’s frequent in reviews of a specific product feature, not just "good" or "great"):

# Convert processed tokens back to full text strings for TF-IDF
corpus = [' '.join(tokens) for tokens in processed_reviews]

# Configure TF-IDF to handle unigrams, bigrams, and trigrams in one go
tfidf_vectorizer = TfidfVectorizer(ngram_range=(1, 3))
tfidf_matrix = tfidf_vectorizer.fit_transform(corpus)

# Map each term to its total TF-IDF score across all reviews
term_tfidf_scores = dict(
    zip(tfidf_vectorizer.get_feature_names_out(), 
        tfidf_matrix.sum(axis=0).A1)
)

# Sort terms by score (highest first) and take top 100
top_100_tfidf = sorted(term_tfidf_scores.items(), key=lambda x: x[1], reverse=True)[:100]

# Print results
print("\nTop 100 Terms by TF-IDF:")
for term, score in top_100_tfidf:
    print(f"{term}: {score:.4f}")
Bonus: Filter by Part-of-Speech (POS) for Better Quality

If you want to focus on product-specific terms (like "battery life" or "camera resolution"), use POS tagging to keep only nouns, adjectives, and relevant verbs. Here’s how to add that with spaCy:

import spacy

# Load spaCy's English model (run `pip install spacy` and `python -m spacy download en_core_web_sm` first)
nlp = spacy.load('en_core_web_sm')

def preprocess_with_pos(text):
    doc = nlp(text.lower())
    # Keep only nouns, adjectives, and verbs (adjust POS tags as needed)
    filtered_tokens = [
        token.text for token in doc 
        if token.pos_ in ['NOUN', 'ADJ', 'VERB'] and token.text not in stop_words
    ]
    return filtered_tokens

# Re-process reviews with POS filtering
processed_reviews_pos = [preprocess_with_pos(review) for review in your_reviews]
# Repeat steps 2 and 3 with processed_reviews_pos instead of processed_reviews

This will filter out meaningless combinations and give you more actionable product terms.

内容的提问来源于stack exchange,提问作者chefnewman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 22:39:09