Python中TextBlob的sentiment.polarity无法正常使用问题求助
问题:TextBlob情感极性分数始终为负,调用
blob.sentiment.polarity异常 我尝试用TextBlob为商品评论生成情感与极性分数,但发现所有结果的情感分数都是负数,blob.sentiment.polarity似乎无法正常工作;另外我曾尝试在函数里定义polarity_score = blob.polarity(之前写错成blog.polarity),同样失败了。相关代码如下:
import pandas as pd from textblob import TextBlob import spacy # Load spaCy model - using _md nlp = spacy.load('en_core_web_md') # Function to perform sentiment analysis using spaCy def predict_sentiment_spacy(text): try: # Tokenize the text doc = nlp(text) # Calculate sentiment score sentiment_score = doc.sentiment # Attach a label to the sentiment sentiment = 'Positive' if sentiment_score >= 0.5 else 'Negative' return sentiment, sentiment_score except Exception as e: # Handle exception, such as SSL errors print(f"Error in spaCy sentiment analysis: {e}") return 'Error', None # Function to perform sentiment analysis using TextBlob def predict_sentiment_textblob(text): try: # Create a TextBlob object blob = TextBlob(text) # Get sentiment polarity score sentiment_score = blob.sentiment.polarity # Determine sentiment label sentiment = 'Positive' if sentiment_score > 0 else 'Negative' if sentiment_score < 0 else 'Neutral' return sentiment, sentiment_score except Exception as e: # Handle exception print(f"Error in TextBlob sentiment analysis: {e}") return 'Error', None # Function to calculate similarity score using spaCy def calculate_similarity(review1, review2): try: # Tokenize and preprocess the texts doc1 = nlp(review1) doc2 = nlp(review2) # Calculate similarity score similarity_score = doc1.similarity(doc2) return similarity_score except Exception as e: # Handle exception print(f"Error in calculating similarity: {e}") return None # Load the data df = pd.read_csv('amazon_product_reviews.csv') # Select the 'reviews.text' column reviews_data = df['reviews.text'] # Remove missing values clean_data = df.dropna(subset=['reviews.text']) # Perform sentiment analysis and similarity calculation compare_more = True while compare_more: try: # Get user input for indices index1 = int(input("Enter the index of the first review: ")) index2 = int(input("Enter the index of the second review: ")) # Get the reviews review1 = clean_data.iloc[index1]['reviews.text'] review2 = clean_data.iloc[index2]['reviews.text'] # SpaCy Sentiment Analysis sentiment_spacy_1, score_spacy_1 = predict_sentiment_spacy(review1) print(f"Sentiment for Review {index1}: {sentiment_spacy_1}") sentiment_spacy_2, score_spacy_2 = predict_sentiment_spacy(review2) print(f"Sentiment for Review {index2}: {sentiment_spacy_2}") # TextBlob Sentiment Analysis sentiment_textblob_1, score_textblob_1 = predict_sentiment_textblob(review1) print(f"TextBlob Sentiment for Review {index1}: {sentiment_textblob_1}") print(f"TextBlob Sentiment Score for Review {index1}: {score_textblob_1}") sentiment_textblob_2, score_textblob_2 = predict_sentiment_textblob(review2) print(f"TextBlob Sentiment for Review {index2}: {sentiment_textblob_2}") print(f"TextBlob Sentiment Score for Review {index2}: {score_textblob_2}") # Calculate Similarity similarity_score = calculate_similarity(review1, review2) print(f"Similarity Score between Review {index1} and Review {index2}: {similarity_score}") except ValueError: print("Invalid index. Please enter valid integer indices.") # Ask the user if they want to compare more reviews response = input("Do you want to compare more reviews? (yes/no): ").lower() compare_more = response == 'yes' # Print sentiment analysis for all reviews print("\nSentiment Analysis for All Reviews:") for index, review in enumerate(clean_data['reviews.text']): sentiment_spacy, _ = predict_sentiment_spacy(review) sentiment_textblob, score_textblob = predict_sentiment_textblob(review) print(f"Review {index + 1}:") print(f"- SpaCy Predicted Sentiment: {sentiment_spacy}") print(f"- TextBlob Predicted Sentiment: {sentiment_textblob}") print(f"- TextBlob Sentiment Score: {score_textblob}") print('-' * 50)
问题分析与修复方案
1. SpaCy情感分析失效原因
你使用的en_core_web_md模型不包含情感分析功能,doc.sentiment属性不存在,代码中只是被try-except捕获了异常,导致返回错误结果。要让SpaCy支持情感分析,需要安装spacytextblob扩展。
2. TextBlob分数异常的可能原因
- 文本未预处理:原始评论可能包含特殊符号、数字或杂乱格式,干扰TextBlob的语义判断。
- 数据集本身问题:如果数据集负面评论占比极高,也会出现分数普遍为负,但更可能是预处理缺失导致的偏差。
3. 具体修复步骤
修复SpaCy情感分析
先安装依赖:
pip install spacytextblob python -m textblob.download_corpora
修改SpaCy情感分析函数:
def predict_sentiment_spacy(text): try: doc = nlp(text) # 调用spacytextblob的情感极性属性 sentiment_score = doc._.blob.sentiment.polarity # 统一标签判断逻辑 sentiment = 'Positive' if sentiment_score > 0 else 'Negative' if sentiment_score < 0 else 'Neutral' return sentiment, sentiment_score except Exception as e: print(f"Error in spaCy sentiment analysis: {e}") return 'Error', None
优化TextBlob文本预处理
添加文本清洗函数,提升分析准确性:
import re def clean_text(text): # 移除特殊字符、数字,转为小写 text = re.sub(r'[^a-zA-Z\s]', '', text).lower() return text def predict_sentiment_textblob(text): try: cleaned_text = clean_text(text) blob = TextBlob(cleaned_text) sentiment_score = blob.sentiment.polarity sentiment = 'Positive' if sentiment_score > 0 else 'Negative' if sentiment_score < 0 else 'Neutral' return sentiment, sentiment_score except Exception as e: print(f"Error in TextBlob sentiment analysis: {e}") return 'Error', None
修正变量名笔误
你提到的polarity_score = blog.polarity是拼写错误,正确写法是polarity_score = blob.sentiment.polarity,这在你当前代码中已经是正确的,无需修改。
其他注意事项
- 确保
amazon_product_reviews.csv文件路径正确,代码可访问。 - 确认依赖版本兼容:
pandas>=1.0、textblob>=0.17、spacy>=3.0、spacytextblob>=4.0
内容的提问来源于stack exchange,提问作者Nikhil Kalanjee
相关产品推荐
相关产品推荐

