如何使用NLTK的SentiWordNet获取n元语法(含二元组)的同义词
Hey there! The issue you're hitting is super common—WordNet (and SentiWordNet, which builds on it) doesn't work with arbitrary multi-word phrases like "horrible food" or "i tried". These tools are designed around individual words (lemmas) and their specific parts of speech, not full n-grams. So passing a two-word phrase directly to swn.senti_synset() is never going to work. Let's break down how to fix this with practical code examples.
SentiWordNet expects inputs in the format lemma.pos.offset (e.g., horrible.a.01), where each entry corresponds to a single word synset. There’s no database entry for multi-word expressions like "horrible food", so your original code throws an error because it can’t find a matching synset.
Most n-grams can be split into their component words, each with a specific part of speech (POS). For each word, you can retrieve synonyms (or sentiment synsets) and then combine them contextually. First, you’ll need to map NLTK’s POS tags to WordNet’s required labels.
Here’s a corrected, actionable code example:
from nltk.corpus import sentiwordnet as swn, wordnet from nltk.tag import pos_tag from nltk.tokenize import word_tokenize # Map NLTK POS tags to WordNet's POS format def nltk_to_wordnet_pos(nltk_tag): if nltk_tag.startswith('J'): return wordnet.ADJ elif nltk_tag.startswith('V'): return wordnet.VERB elif nltk_tag.startswith('N'): return wordnet.NOUN elif nltk_tag.startswith('R'): return wordnet.ADV else: return None # Skip unsupported POS tags def get_phrase_synonyms(phrase): # Tokenize and tag the phrase with POS labels tokens = word_tokenize(phrase) tagged_words = pos_tag(tokens) phrase_synonyms = [] for word, tag in tagged_words: wn_pos = nltk_to_wordnet_pos(tag) if not wn_pos: continue # Skip words we can't map to WordNet # Get WordNet synsets for the word (using correct POS) synsets = wordnet.synsets(word, pos=wn_pos) if not synsets: continue # Extract synonyms from the first synset (adjust index for more options) synonyms = [lemma.name().replace('_', ' ') for lemma in synsets[0].lemmas()] phrase_synonyms.append({ 'original_word': word, 'part_of_speech': wn_pos, 'synonyms': synonyms }) return phrase_synonyms # Test with your example phrase print(get_phrase_synonyms("horrible food"))
This will return synonyms for each word in the phrase (e.g., "terrible" for "horrible", "cuisine" for "food") which you can combine to form meaningful n-gram synonyms like "terrible cuisine".
Some common collocations (like "ice cream" or "coffee shop") are stored as single synsets in WordNet. You can add a check for these before splitting the phrase into individual words:
def get_multi_word_synset(phrase): # Replace spaces with underscores to match WordNet's formatting phrase_underscored = phrase.replace(' ', '_') # Check all possible POS categories for the phrase for pos in [wordnet.NOUN, wordnet.ADJ, wordnet.VERB, wordnet.ADV]: synsets = wordnet.synsets(phrase_underscored, pos=pos) if synsets: # Extract synonyms for the multi-word synset synonyms = [lemma.name().replace('_', ' ') for lemma in synsets[0].lemmas()] return { 'original_phrase': phrase, 'synonyms': synonyms } return None # No multi-word synset found # Example usage print(get_multi_word_synset("ice cream")) # Returns synonyms like "frozen custard" print(get_multi_word_synset("horrible food")) # Returns None, fall back to Approach 1
If you want synonyms that carry similar emotional tone (since you referenced SentiWordNet), you can filter synonyms based on their positive/negative sentiment scores:
def get_sentiment_matched_synonyms(word, pos): # Get the top SentiWordNet synset for the word senti_synsets = swn.senti_synsets(word, pos=pos) if not senti_synsets: return [] top_synset = senti_synsets[0] original_sentiment = top_synset.pos_score() - top_synset.neg_score() # Get WordNet synonyms for the word wn_synset = wordnet.synset(top_synset.synset.name()) synonyms = [lemma.name().replace('_', ' ') for lemma in wn_synset.lemmas()] # Filter synonyms with similar sentiment (adjust threshold as needed) matched_synonyms = [] for syn in synonyms: syn_senti_synsets = swn.senti_synsets(syn, pos=pos) if syn_senti_synsets: syn_sentiment = syn_senti_synsets[0].pos_score() - syn_senti_synsets[0].neg_score() if abs(syn_sentiment - original_sentiment) < 0.2: matched_synonyms.append(syn) return matched_synonyms
内容的提问来源于stack exchange,提问作者T3J45

