德语Twitter情感分析无法分类正负情感问题求助
问题诊断
你的问题核心是TextBlob默认仅支持英文情感分析,它的内置情感词典和模型都是为英文训练的,处理德语文本时会返回无意义的极性值(大多为0),导致无法正确分类正负情感。另外你的代码里还有一处语法错误:
tweet:textblob.TextBlob(tweet).sentiment.subjectivity)
这行是不完整的语句,应该是用于创建Subjectivity列的代码,需要修正。
解决方案
方案1:使用TextBlob德语扩展(快速适配)
TextBlob有第三方扩展textblob-de专门支持德语情感分析,步骤如下:
- 安装扩展:
pip install textblob-de
- 修改代码:
- 导入
TextBlobDE替换原有的TextBlob - 修正语法错误,完整构建
Subjectivity列 - 调整情感分析的调用逻辑
- 导入
修改后的核心代码:
import tweepy import textblob_de # 导入德语扩展 import pandas as pd import numpy as np import matplotlib.pyplot as plt import re # 认证部分保持不变 authenticator = tweepy.OAuthHandler(api_key, api_key_secret) authenticator.set_access_token(bearer, bearer_secret) api = tweepy.API(authenticator, wait_on_rate_limit=True) # 抓取德语推文部分保持不变 search = f'#afd -filter:retweets' tweet_cursor = tweepy.Cursor(api.search_tweets, q=search, lang='de', tweet_mode='extended').items(1000) tweets = [tweet.full_text for tweet in tweet_cursor] tweets_df = pd.DataFrame(tweets, columns=['Tweets']) # 文本清洗部分保持不变 for _, row in tweets_df.iterrows(): row['Tweets'] = re.sub(r'http\S+', '', row['Tweets']) row['Tweets'] = re.sub(r'#\S+', '', row['Tweets']) row['Tweets'] = re.sub(r'@\S+', '', row['Tweets']) row['Tweets'] = re.sub(r'\\n+', '', row['Tweets']) # 替换为德语情感分析 tweets_df['Polarity'] = tweets_df['Tweets'].map(lambda tweet: textblob_de.TextBlobDE(tweet).sentiment.polarity) tweets_df['Subjectivity'] = tweets_df['Tweets'].map(lambda tweet: textblob_de.TextBlobDE(tweet).sentiment.subjectivity) # 修正语法错误 tweets_df['result'] = tweets_df['Polarity'].map(lambda pol: '+' if pol > 0 else '-') # 统计和可视化部分保持不变 positive = tweets_df[tweets_df.result == '+'].count()['Tweets'] negative = tweets_df[tweets_df.result == '-'].count()['Tweets'] plt.bar([0, 1], [positive, negative], label=['Positive', 'Negative'], color=['green', 'red']) plt.legend() plt.show()
方案2:使用Hugging Face Transformers(更准确)
如果需要更精准的德语情感分析,推荐使用预训练的德语BERT模型,比如oliverguhr/german-sentiment-bert:
- 安装依赖:
pip install transformers torch
- 修改情感分析部分代码:
from transformers import pipeline # 加载德语情感分析管道 sentiment_analyzer = pipeline("sentiment-analysis", model="oliverguhr/german-sentiment-bert") # 替换原有的Polarity和result列生成逻辑 def get_sentiment(tweet): result = sentiment_analyzer(tweet)[0] # 映射模型输出到你的正负标记 return 1 if result['label'] == 'positive' else -1 if result['label'] == 'negative' else 0 tweets_df['Polarity'] = tweets_df['Tweets'].apply(get_sentiment) tweets_df['result'] = tweets_df['Polarity'].map(lambda pol: '+' if pol > 0 else '-')
这个方案的情感分类准确率远高于TextBlob,适合需要高精度的场景。
注意事项
- TextBlob-de的情感分析基于规则和简单词典,对于复杂德语文本的准确率有限;
- 使用Transformers模型时,首次运行会下载预训练权重,需要联网;
- 确保你的Twitter API密钥配置正确,避免请求被限制。
内容的提问来源于stack exchange,提问作者MoBa
相关产品推荐
相关产品推荐

