从零构建文本分类工具:情感分析集成难点求助
Hey bubbaspaarx, great project idea—building this classification tool from scratch without libraries is a solid learning exercise, and tackling sentiment-aware filtering is a smart next step. Let's walk through how you can pull this off, step by step.
Core Approach
Your main challenge is linking sentiment (positive/negative) to the category keywords you're extracting, then filtering out categories tied to negative sentiment. Here's an actionable breakdown:
1. Define Your Reference Libraries
First, formalize two key datasets (you can expand these as needed):
- Category Keyword Map: Map each target category to a list of related terms (including multi-word phrases like "my son"):
category_keywords = { 'father': ['son', 'my son'], 'photography': ['photograph', 'portrait', 'pic', 'photographs'], 'travel': ['holiday', 'trip'], 'spain': ['spain'], 'cooking': ['baked', 'cooking', 'bake', 'cake'], 'chocolate': ['chocolate'] } - Negative Sentiment Word List: List terms that signal dislike for a category:
negative_words = ['hate', 'dislike', 'detest', 'loathe', 'despise']
2. Text Preprocessing
Clean the input text to make analysis easier—lowercase everything, strip punctuation, and split into words:
def preprocess_text(text): text = text.lower() # Remove common punctuation by replacing with spaces for p in '.,!?;': text = text.replace(p, ' ') return text.split()
3. Extract Candidate Categories
First, pull all categories that have matching keywords in the text (this refines your existing word-matching logic to handle multi-word terms):
def extract_candidate_categories(text_words): candidates = set() for category, keywords in category_keywords.items(): for keyword in keywords: keyword_parts = keyword.split() # Check if the keyword phrase exists in the text for i in range(len(text_words) - len(keyword_parts) + 1): if text_words[i:i+len(keyword_parts)] == keyword_parts: candidates.add(category) break # Stop checking other keywords for this category once found return candidates
4. Filter Out Negatively Linked Categories
This is the critical part: check if a category's keywords are near negative sentiment words. We'll use a "context window" (e.g., 3 words before/after the keyword) to determine if the sentiment targets that category:
def filter_negative_categories(text_words, candidates): final_categories = [] window_size = 3 # Adjust based on how far sentiment words can be from the target # First, map each category to the positions of its keywords in the text category_positions = {} for category in candidates: positions = [] for keyword in category_keywords[category]: keyword_parts = keyword.split() for i in range(len(text_words) - len(keyword_parts) + 1): if text_words[i:i+len(keyword_parts)] == keyword_parts: # Record start and end indices of the keyword phrase positions.append((i, i + len(keyword_parts) - 1)) category_positions[category] = positions # Check each category's keyword positions for nearby negative words for category in candidates: is_negative = False for (start_idx, end_idx) in category_positions[category]: # Define the range of words to check around the keyword check_start = max(0, start_idx - window_size) check_end = min(len(text_words)-1, end_idx + window_size) # Look for negative words in this window for word in text_words[check_start:check_end+1]: if word in negative_words: is_negative = True break if is_negative: break if not is_negative: final_categories.append(category) return sorted(final_categories)
5. Test It Out
Run the functions with your example inputs to see it work:
# Test 1: Original positive text text1 = "Whenever I am out walking with my son, I like to take portrait photographs of him to see how he changes over time. My favourite is a pic of him when we were on holiday in Spain and when his face was covered in chocolate from a cake we had baked" words1 = preprocess_text(text1) candidates1 = extract_candidate_categories(words1) print(filter_negative_categories(words1, candidates1)) # Output: ['chocolate', 'cooking', 'father', 'photography', 'spain', 'travel'] # Test 2: Text with negative cooking sentiment text2 = "Whenever I am out walking with my son, I like to take portrait photographs of him. I hate cooking but love the chocolate cake we had in Spain during our holiday" words2 = preprocess_text(text2) candidates2 = extract_candidate_categories(words2) print(filter_negative_categories(words2, candidates2)) # Output: ['chocolate', 'father', 'photography', 'spain', 'travel'] (cooking is excluded)
Next Steps to Improve
- Handle Word Variants: Add simple stemming (e.g., strip "ed"/"ing" from words like "baked" → "bake") to catch more matches without expanding your keyword list.
- Negation Handling: Extend sentiment logic to catch phrases like "don't like" by checking for "not" or "don't" followed by positive words.
- Adjust Context Window: Tweak the
window_sizebased on sentence length—shorter sentences might need a smaller window to avoid false links. - Multi-Sentence Awareness: If the text has multiple sentences, only check sentiment in the same sentence as the category keyword (since sentiment in one sentence rarely targets a category in another).
内容的提问来源于stack exchange,提问作者bubbaspaarx

