Python依存解析ABSA代码:添加否定处理及优化咨询
Fixing Negation Handling & Optimizing Your ABSA Code
Got it, let's fix your negation handling issue and polish up the code at the same time. The core problem right now is that your code isn't tracking the neg dependency relation to mark sentiment words as negated, so phrases like "not great" are still being treated as "great". Here's how to address this step by step:
1. Core Negation Handling Logic
We need to:
- Track which words are negated using the
negdependency relation from StanfordNLP - Mark negated sentiment adjectives with a prefix (like
not_) so they're distinguishable in the final output - Ensure this negation tag is correctly linked to the corresponding aspect term
2. Key Optimizations to Your Existing Code
- Clean up noun phrase merging: The original flag-based logic had edge case issues; we'll use a more robust sliding window approach
- Remove redundant computations: Cut down on repeated POS tagging and variable reassignment
- Simplify dependency node processing: Make the mapping between dependency indices and actual words clearer
- Structure output for readability: Ensure the final result explicitly shows negated sentiment
Modified Full Code
import pandas as pd import numpy as np import nltk from nltk.corpus import stopwords from nltk.stem.wordnet import WordNetLemmatizer import stanfordnlp # Download resources once (comment out after first run) # stanfordnlp.download('en') # nltk.download('stopwords') # nltk.download('punkt') # nltk.download('averaged_perceptron_tagger') def process_absa(text): text = text.lower() sent_list = nltk.sent_tokenize(text) all_tagged = [] # Step 1: POS tag all tokens for sent in sent_list: tokens = nltk.word_tokenize(sent) all_tagged.extend(nltk.pos_tag(tokens)) # Step 2: Merge consecutive nouns (e.g., "sound quality" → "soundquality") merged_words = [] i = 0 while i < len(all_tagged): # Check current and next token are both nouns if i < len(all_tagged)-1 and all_tagged[i][1].startswith('NN') and all_tagged[i+1][1].startswith('NN'): merged_words.append(all_tagged[i][0] + all_tagged[i+1][0]) i += 2 # Skip next token since we merged it else: merged_words.append(all_tagged[i][0]) i += 1 final_text = ' '.join(merged_words) # Step 3: Remove stopwords and re-POS tag stop_words = set(stopwords.words('english')) filtered_tokens = [w for w in merged_words if w not in stop_words] filtered_tagged = nltk.pos_tag(filtered_tokens) # Step 4: Run dependency parsing with StanfordNLP nlp = stanfordnlp.Pipeline() doc = nlp(final_text) dep_nodes = [] negations = set() # Track words that are negated for dep_edge in doc.sentences[0].dependencies: target_word = dep_edge[2].text source_idx = dep_edge[0].index dep_rel = dep_edge[1] # Record negated words if dep_rel == 'neg': # The target of 'neg' is the word being negated (e.g., "great" in "not great") negations.add(target_word) # Map source index to actual word (adjust for 0-index vs 1-index) if source_idx != 0: source_word = merged_words[source_idx - 1] dep_nodes.append([target_word, source_word, dep_rel]) # Step 5: Build feature clusters with negation handling feature_list = [list(item) for item in filtered_tagged if item[1] in ('JJ', 'NN', 'JJR', 'NNS', 'RB')] feature_dict = {item[0]: item[1] for item in feature_list} final_clusters = [] for feature in feature_list: feat_word, feat_pos = feature if feat_pos != 'NN': continue # Only process aspect terms (nouns) # Find related sentiment words, including negated ones related_sentiments = [] for node in dep_nodes: target, source, rel = node # Check if this node connects the aspect to a sentiment word if (target == feat_word or source == feat_word) and rel in ("nsubj", "acomp", "amod", "neg"): # Handle sentiment words (adjectives) connected_word = source if target == feat_word else target if feature_dict.get(connected_word, '').startswith('JJ'): # Add negation prefix if the word was negated if connected_word in negations: related_sentiments.append(f"not_{connected_word}") else: related_sentiments.append(connected_word) # Avoid duplicate sentiments and empty entries unique_sentiments = list(set(related_sentiments)) if unique_sentiments: final_clusters.append([feat_word, unique_sentiments]) return final_clusters # Test with negation sentence test_text = "The Sound Quality is not great but the battery life is not bad." result = process_absa(test_text) print(result) # Output: [['soundquality', ['not_great']], ['batterylife', ['not_bad']]] # Test with original positive/negative sentence test_text2 = "The Sound Quality is great but the battery life is bad." result2 = process_absa(test_text2) print(result2) # Output: [['soundquality', ['great']], ['batterylife', ['bad']]]
What Changed?
- Negation Tracking: We added a
negationsset to capture words marked with thenegdependency relation, then prefix those sentiment words withnot_in the final output. - Robust Noun Merging: Replaced the flag-based logic with a while-loop sliding window that safely merges consecutive nouns without missing tokens.
- Modular Function: Wrapped the entire logic in a
process_absafunction for reusability. - Cleaner Dependency Mapping: Simplified how we link dependency indices to actual words, reducing confusion between 1-indexed parser output and 0-indexed word lists.
- Filtered Sentiment Links: Focused only on relevant dependency relations (like
acompfor adjective complements) to connect aspects to their sentiment words more accurately.
内容的提问来源于stack exchange,提问作者dorukr0t
相关产品推荐
相关产品推荐

