仇恨言论分类模型异常:所有测试样本均输出同一结果求助
问题:分类模型所有预测结果均为“No Hate and Offensive”,与预期不符
我有一个包含tweets和labels(标签分为:offensive language、hate speech、no hate and offensive)两列的DataFrame。对tweets执行文本清洗后构建了分类模型,但建模完成后,所有测试文本的预测结果全是‘No Hate and Offensive’,和预期不符。以下是我的代码实现:
# Cleaning my text def clean_text(text): text = str(text).lower() text = re.sub('\[.*?\]', '', text) text = re.sub('https?://\S+|www\.\S+', '', text) text = re.sub('<.*?>+', '', text) text = re.sub('[%s]' % re.escape(string.punctuation), '', text) text = re.sub('\n', '', text) text = re.sub('\w*\d\w*', '', text) text = [word for word in text.split(' ') if word not in stopword] text = "".join(text) text = [stemmer.stem(word) for word in text.split(' ')] text = "".join(text) return text data["tweet"] = data["tweet"].apply(clean_text) #Converting my columns to array import numpy as np x = np.array(data["tweet"]) y = np.array(data["labels"]) # Fiting the tweet column for modelling from sklearn.feature_extraction.text import CountVectorizer from sklearn.model_selection import train_test_split count_vec = CountVectorizer() X = count_vec.fit_transform(x) X_train, X_test, y_train, y_test = train_test_split(X, y, test_size = 0.33, random_state = 42) from sklearn.tree import DecisionTreeClassifier dc_tree = DecisionTreeClassifier() dc_tree.fit(X_train, y_train) ##Testing my Model test_text = "Kill all of them and burn it down" test_data = count_vec.transform([test_text]).toarray() print(dc_tree.predict(test_data)) #Output: ['No Hate and Offensive'] #Expectation: ['Hate Speech'] test_text = "practice love and patience to live a good life!" test_data2 = count_vec.transform([test_text]).toarray() print(dc_tree.predict(test_data2)) #Output: ['No Hate and Offensive']
问题原因及解决方案
1. 文本清洗函数存在致命错误
清洗过程中,你将分词后的列表直接用""拼接,导致单词完全连在一起(比如"kill all"变成"killall"),CountVectorizer无法提取有效特征,模型无法学习到类别间的差异。
修改后的clean_text函数:
def clean_text(text): import re import string from nltk.corpus import stopwords from nltk.stem import PorterStemmer stopword = stopwords.words('english') stemmer = PorterStemmer() text = str(text).lower() text = re.sub('\[.*?\]', '', text) text = re.sub('https?://\S+|www\.\S+', '', text) text = re.sub('<.*?>+', '', text) text = re.sub('[%s]' % re.escape(string.punctuation), '', text) text = re.sub('\n', '', text) text = re.sub('\w*\d\w*', '', text) # 用空格连接过滤停用词后的单词,同时过滤空字符串 text = " ".join([word for word in text.split(' ') if word.strip() and word not in stopword]) # 用空格连接词干提取后的单词,同时过滤空字符串 text = " ".join([stemmer.stem(word) for word in text.split(' ') if word.strip()]) return text
2. 测试文本未做同标准清洗
训练数据是经过清洗的,但测试文本直接用原始内容转换,特征分布不一致,导致模型无法正确识别。
修改测试代码:
# 测试文本先清洗再转换 test_text = "Kill all of them and burn it down" cleaned_test = clean_text(test_text) test_data = count_vec.transform([cleaned_test]) print(dc_tree.predict(test_data)) test_text2 = "practice love and patience to live a good life!" cleaned_test2 = clean_text(test_text2) test_data2 = count_vec.transform([cleaned_test2]) print(dc_tree.predict(test_data2))
3. 数据不平衡问题
如果训练集中“No Hate and Offensive”类别的样本占比过高,模型会偏向预测该类别。先检查标签分布:
print(data['labels'].value_counts())
如果存在不平衡,可通过以下方式解决:
- 使用带类别权重的模型:
dc_tree = DecisionTreeClassifier(class_weight='balanced')
- 对少数类进行过采样(如SMOTE)或对多数类进行欠采样。
4. 模型选择局限性
决策树在文本分类任务中表现通常不如线性模型或集成模型,可尝试更换模型:
# 示例:使用逻辑回归,带类别权重 from sklearn.linear_model import LogisticRegression model = LogisticRegression(class_weight='balanced', max_iter=1000) model.fit(X_train, y_train)
内容的提问来源于stack exchange,提问作者Abuchi
相关产品推荐
相关产品推荐

