You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中TextBlob的sentiment.polarity无法正常使用问题求助

问题:TextBlob情感极性分数始终为负,调用blob.sentiment.polarity异常

我尝试用TextBlob为商品评论生成情感与极性分数,但发现所有结果的情感分数都是负数,blob.sentiment.polarity似乎无法正常工作;另外我曾尝试在函数里定义polarity_score = blob.polarity(之前写错成blog.polarity),同样失败了。相关代码如下:

import pandas as pd
from textblob import TextBlob
import spacy

# Load spaCy model - using _md
nlp = spacy.load('en_core_web_md')

# Function to perform sentiment analysis using spaCy
def predict_sentiment_spacy(text):
    try:
        # Tokenize the text
        doc = nlp(text)

        # Calculate sentiment score
        sentiment_score = doc.sentiment

        # Attach a label to the sentiment 
        sentiment = 'Positive' if sentiment_score >= 0.5 else 'Negative'

        return sentiment, sentiment_score
    except Exception as e:
        # Handle exception, such as SSL errors
        print(f"Error in spaCy sentiment analysis: {e}")
        return 'Error', None

# Function to perform sentiment analysis using TextBlob
def predict_sentiment_textblob(text):
    try:
        # Create a TextBlob object
        blob = TextBlob(text)

        # Get sentiment polarity score
        sentiment_score = blob.sentiment.polarity

        # Determine sentiment label
        sentiment = 'Positive' if sentiment_score > 0 else 'Negative' if sentiment_score < 0 else 'Neutral'

        return sentiment, sentiment_score
    except Exception as e:
        # Handle exception
        print(f"Error in TextBlob sentiment analysis: {e}")
        return 'Error', None

# Function to calculate similarity score using spaCy
def calculate_similarity(review1, review2):
    try:
        # Tokenize and preprocess the texts
        doc1 = nlp(review1)
        doc2 = nlp(review2)

        # Calculate similarity score
        similarity_score = doc1.similarity(doc2)

        return similarity_score
    except Exception as e:
        # Handle exception
        print(f"Error in calculating similarity: {e}")
        return None

# Load the data
df = pd.read_csv('amazon_product_reviews.csv')

# Select the 'reviews.text' column
reviews_data = df['reviews.text']

# Remove missing values
clean_data = df.dropna(subset=['reviews.text'])

# Perform sentiment analysis and similarity calculation
compare_more = True

while compare_more:
    try:
        # Get user input for indices
        index1 = int(input("Enter the index of the first review: "))
        index2 = int(input("Enter the index of the second review: "))

        # Get the reviews
        review1 = clean_data.iloc[index1]['reviews.text']
        review2 = clean_data.iloc[index2]['reviews.text']

        # SpaCy Sentiment Analysis
        sentiment_spacy_1, score_spacy_1 = predict_sentiment_spacy(review1)
        print(f"Sentiment for Review {index1}: {sentiment_spacy_1}")

        sentiment_spacy_2, score_spacy_2 = predict_sentiment_spacy(review2)
        print(f"Sentiment for Review {index2}: {sentiment_spacy_2}")

        # TextBlob Sentiment Analysis
        sentiment_textblob_1, score_textblob_1 = predict_sentiment_textblob(review1)
        print(f"TextBlob Sentiment for Review {index1}: {sentiment_textblob_1}")
        print(f"TextBlob Sentiment Score for Review {index1}: {score_textblob_1}")

        sentiment_textblob_2, score_textblob_2 = predict_sentiment_textblob(review2)
        print(f"TextBlob Sentiment for Review {index2}: {sentiment_textblob_2}")
        print(f"TextBlob Sentiment Score for Review {index2}: {score_textblob_2}")

        # Calculate Similarity
        similarity_score = calculate_similarity(review1, review2)
        print(f"Similarity Score between Review {index1} and Review {index2}: {similarity_score}")

    except ValueError:
        print("Invalid index. Please enter valid integer indices.")

    # Ask the user if they want to compare more reviews
    response = input("Do you want to compare more reviews? (yes/no): ").lower()
    compare_more = response == 'yes'

# Print sentiment analysis for all reviews
print("\nSentiment Analysis for All Reviews:")
for index, review in enumerate(clean_data['reviews.text']):
    sentiment_spacy, _ = predict_sentiment_spacy(review)
    sentiment_textblob, score_textblob = predict_sentiment_textblob(review)
    print(f"Review {index + 1}:")
    print(f"- SpaCy Predicted Sentiment: {sentiment_spacy}")
    print(f"- TextBlob Predicted Sentiment: {sentiment_textblob}")
    print(f"- TextBlob Sentiment Score: {score_textblob}")
    print('-' * 50)

问题分析与修复方案

1. SpaCy情感分析失效原因

你使用的en_core_web_md模型不包含情感分析功能,doc.sentiment属性不存在,代码中只是被try-except捕获了异常,导致返回错误结果。要让SpaCy支持情感分析,需要安装spacytextblob扩展。

2. TextBlob分数异常的可能原因

  • 文本未预处理:原始评论可能包含特殊符号、数字或杂乱格式,干扰TextBlob的语义判断。
  • 数据集本身问题:如果数据集负面评论占比极高,也会出现分数普遍为负,但更可能是预处理缺失导致的偏差。

3. 具体修复步骤

修复SpaCy情感分析

先安装依赖:

pip install spacytextblob
python -m textblob.download_corpora

修改SpaCy情感分析函数:

def predict_sentiment_spacy(text):
    try:
        doc = nlp(text)
        # 调用spacytextblob的情感极性属性
        sentiment_score = doc._.blob.sentiment.polarity
        # 统一标签判断逻辑
        sentiment = 'Positive' if sentiment_score > 0 else 'Negative' if sentiment_score < 0 else 'Neutral'
        return sentiment, sentiment_score
    except Exception as e:
        print(f"Error in spaCy sentiment analysis: {e}")
        return 'Error', None

优化TextBlob文本预处理

添加文本清洗函数,提升分析准确性:

import re

def clean_text(text):
    # 移除特殊字符、数字,转为小写
    text = re.sub(r'[^a-zA-Z\s]', '', text).lower()
    return text

def predict_sentiment_textblob(text):
    try:
        cleaned_text = clean_text(text)
        blob = TextBlob(cleaned_text)
        sentiment_score = blob.sentiment.polarity
        sentiment = 'Positive' if sentiment_score > 0 else 'Negative' if sentiment_score < 0 else 'Neutral'
        return sentiment, sentiment_score
    except Exception as e:
        print(f"Error in TextBlob sentiment analysis: {e}")
        return 'Error', None

修正变量名笔误

你提到的polarity_score = blog.polarity是拼写错误,正确写法是polarity_score = blob.sentiment.polarity,这在你当前代码中已经是正确的,无需修改。

其他注意事项

  • 确保amazon_product_reviews.csv文件路径正确,代码可访问。
  • 确认依赖版本兼容:pandas>=1.0、textblob>=0.17、spacy>=3.0、spacytextblob>=4.0

内容的提问来源于stack exchange,提问作者Nikhil Kalanjee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 20:00:59