You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

仇恨言论分类模型异常:所有测试样本均输出同一结果求助

问题:分类模型所有预测结果均为“No Hate and Offensive”,与预期不符

我有一个包含tweets和labels(标签分为:offensive language、hate speech、no hate and offensive)两列的DataFrame。对tweets执行文本清洗后构建了分类模型,但建模完成后,所有测试文本的预测结果全是‘No Hate and Offensive’,和预期不符。以下是我的代码实现:

# Cleaning my text
def clean_text(text):
    text = str(text).lower()
    text = re.sub('\[.*?\]', '', text)
    text = re.sub('https?://\S+|www\.\S+', '', text)
    text = re.sub('<.*?>+', '', text)
    text = re.sub('[%s]' % re.escape(string.punctuation), '', text)
    text = re.sub('\n', '', text)
    text = re.sub('\w*\d\w*', '', text)
    text = [word for word in text.split(' ') if word not in stopword]
    text = "".join(text)
    text = [stemmer.stem(word) for word in text.split(' ')]
    text = "".join(text)
    
    return text

data["tweet"] = data["tweet"].apply(clean_text)


#Converting my columns to array
import numpy as np
x = np.array(data["tweet"])
y = np.array(data["labels"])


# Fiting the tweet column for modelling
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.model_selection import train_test_split
count_vec = CountVectorizer()
X = count_vec.fit_transform(x)

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size = 0.33, random_state = 42)


from sklearn.tree import DecisionTreeClassifier

dc_tree = DecisionTreeClassifier()
dc_tree.fit(X_train, y_train)

##Testing my Model
test_text = "Kill all of them and burn it down"
test_data = count_vec.transform([test_text]).toarray()

print(dc_tree.predict(test_data)) 
#Output: ['No Hate and Offensive'] #Expectation: ['Hate Speech']


test_text = "practice love and patience to live a good life!"
test_data2 = count_vec.transform([test_text]).toarray()
print(dc_tree.predict(test_data2)) 
#Output: ['No Hate and Offensive']

问题原因及解决方案

1. 文本清洗函数存在致命错误

清洗过程中,你将分词后的列表直接用""拼接,导致单词完全连在一起(比如"kill all"变成"killall"),CountVectorizer无法提取有效特征,模型无法学习到类别间的差异。

修改后的clean_text函数:

def clean_text(text):
    import re
    import string
    from nltk.corpus import stopwords
    from nltk.stem import PorterStemmer
    
    stopword = stopwords.words('english')
    stemmer = PorterStemmer()
    
    text = str(text).lower()
    text = re.sub('\[.*?\]', '', text)
    text = re.sub('https?://\S+|www\.\S+', '', text)
    text = re.sub('<.*?>+', '', text)
    text = re.sub('[%s]' % re.escape(string.punctuation), '', text)
    text = re.sub('\n', '', text)
    text = re.sub('\w*\d\w*', '', text)
    # 用空格连接过滤停用词后的单词,同时过滤空字符串
    text = " ".join([word for word in text.split(' ') if word.strip() and word not in stopword])
    # 用空格连接词干提取后的单词,同时过滤空字符串
    text = " ".join([stemmer.stem(word) for word in text.split(' ') if word.strip()])
    
    return text

2. 测试文本未做同标准清洗

训练数据是经过清洗的,但测试文本直接用原始内容转换,特征分布不一致,导致模型无法正确识别。

修改测试代码:

# 测试文本先清洗再转换
test_text = "Kill all of them and burn it down"
cleaned_test = clean_text(test_text)
test_data = count_vec.transform([cleaned_test])
print(dc_tree.predict(test_data))

test_text2 = "practice love and patience to live a good life!"
cleaned_test2 = clean_text(test_text2)
test_data2 = count_vec.transform([cleaned_test2])
print(dc_tree.predict(test_data2))

3. 数据不平衡问题

如果训练集中“No Hate and Offensive”类别的样本占比过高,模型会偏向预测该类别。先检查标签分布:

print(data['labels'].value_counts())

如果存在不平衡,可通过以下方式解决:

  • 使用带类别权重的模型:
dc_tree = DecisionTreeClassifier(class_weight='balanced')
  • 对少数类进行过采样(如SMOTE)或对多数类进行欠采样。

4. 模型选择局限性

决策树在文本分类任务中表现通常不如线性模型或集成模型,可尝试更换模型:

# 示例:使用逻辑回归,带类别权重
from sklearn.linear_model import LogisticRegression
model = LogisticRegression(class_weight='balanced', max_iter=1000)
model.fit(X_train, y_train)

内容的提问来源于stack exchange,提问作者Abuchi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 15:34:58