如何从NLTK分类器获取正负预测的精确率、召回率与F值?
如何用NLTK获取情感分析的精确率、召回率和F值
嘿,我来帮你搞定这个问题!首先得先把你的训练数据补全——最后那条负面样本没写完,我先帮你补全(你可以根据实际情况调整),然后一步步实现特征提取、训练分类器,最后计算你需要的评估指标:
第一步:补全并整理训练数据
import nltk from nltk.tokenize import word_tokenize # 补全后的训练数据 train = [ ('I love this sandwich.', 'pos'), ('This is an amazing place!', 'pos'), ('I feel very good about these beers.', 'pos'), ('This is my best work.', 'pos'), ("What an awesome view", 'pos'), ('I do not like this restaurant', 'neg'), ('I am tired of this stuff.', 'neg'), ("I can't deal with this anymore.", 'neg') # 补全的负面样本 ]
第二步:提取文本特征
NLTK的分类器需要把文本转换成特征字典,这里我们用最简单的“词存在”特征(即文本中出现的每个词作为特征,值为True):
def extract_features(text): # 分词并转小写,避免大小写影响 words = word_tokenize(text.lower()) return {word: True for word in words} # 把训练数据转换成分类器能识别的特征格式 train_features = [(extract_features(text), label) for (text, label) in train]
第三步:拆分数据集并训练分类器
为了评估模型性能,我们需要把数据拆分成训练集和测试集(示例中用前6条训练,后2条测试;实际项目建议用更大的数据集或交叉验证):
from nltk.classify import NaiveBayesClassifier # 拆分训练集和测试集 train_set = train_features[:6] test_set = train_features[6:] # 训练Naive Bayes分类器(NLTK中最常用的文本分类器之一) classifier = NaiveBayesClassifier.train(train_set)
第四步:计算精确率、召回率和F值
这里有两种可靠的计算方式,你可以根据需求选择:
方式1:手动计算(更直观)
先统计每个类的真正例(TP)、假正例(FP)、假负例(FN),再代入公式计算:
# 收集真实标签和预测标签 true_labels = [label for (features, label) in test_set] predicted_labels = [classifier.classify(features) for (features, label) in test_set] # 定义正负类标签 pos_label = 'pos' neg_label = 'neg' # 统计TP/FP/FN # 正类 pos_tp = sum(1 for t, p in zip(true_labels, predicted_labels) if t == pos_label and p == pos_label) pos_fp = sum(1 for t, p in zip(true_labels, predicted_labels) if t != pos_label and p == pos_label) pos_fn = sum(1 for t, p in zip(true_labels, predicted_labels) if t == pos_label and p != pos_label) # 负类 neg_tp = sum(1 for t, p in zip(true_labels, predicted_labels) if t == neg_label and p == neg_label) neg_fp = sum(1 for t, p in zip(true_labels, predicted_labels) if t != neg_label and p == neg_label) neg_fn = sum(1 for t, p in zip(true_labels, predicted_labels) if t == neg_label and p != neg_label) # 计算精确率(Precision = TP/(TP+FP)) pos_precision = pos_tp / (pos_tp + pos_fp) if (pos_tp + pos_fp) > 0 else 0.0 neg_precision = neg_tp / (neg_tp + neg_fp) if (neg_tp + neg_fp) > 0 else 0.0 # 计算召回率(Recall = TP/(TP+FN)) pos_recall = pos_tp / (pos_tp + pos_fn) if (pos_tp + pos_fn) > 0 else 0.0 neg_recall = neg_tp / (neg_tp + neg_fn) if (neg_tp + neg_fn) > 0 else 0.0 # 计算F值(F-measure = 2*(Precision*Recall)/(Precision+Recall)) pos_fmeasure = 2 * (pos_precision * pos_recall) / (pos_precision + pos_recall) if (pos_precision + pos_recall) > 0 else 0.0 neg_fmeasure = 2 * (neg_precision * neg_recall) / (neg_precision + neg_recall) if (neg_precision + neg_recall) > 0 else 0.0 # 打印结果 print(f"正类({pos_label}) - 精确率: {pos_precision:.2f}, 召回率: {pos_recall:.2f}, F值: {pos_fmeasure:.2f}") print(f"负类({neg_label}) - 精确率: {neg_precision:.2f}, 召回率: {neg_recall:.2f}, F值: {neg_fmeasure:.2f}")
方式2:使用NLTK内置的metrics模块
NLTK提供了现成的函数来计算这些指标,但要注意传入正确的参数(避免用集合去重,否则会丢失标签信息):
import nltk.metrics as metrics # 同样先收集真实和预测标签 true_labels = [label for (features, label) in test_set] predicted_labels = [classifier.classify(features) for (features, label) in test_set] # 计算每个类的指标 pos_precision = metrics.precision(set(true_labels), set(predicted_labels), pos_label) pos_recall = metrics.recall(set(true_labels), set(predicted_labels), pos_label) pos_fmeasure = metrics.f_measure(set(true_labels), set(predicted_labels), pos_label) neg_precision = metrics.precision(set(true_labels), set(predicted_labels), neg_label) neg_recall = metrics.recall(set(true_labels), set(predicted_labels), neg_label) neg_fmeasure = metrics.f_measure(set(true_labels), set(predicted_labels), neg_label) # 打印结果 print(f"正类({pos_label}) - 精确率: {pos_precision}, 召回率: {pos_recall}, F值: {pos_fmeasure}")
额外建议:用交叉验证提升评估稳定性
如果你的数据集很小,拆分训练测试集会导致结果不稳定,建议用交叉验证:
import random # 打乱数据保证随机性 random.shuffle(train_features) # 10折交叉验证 fold_count = 10 fold_size = len(train_features) // fold_count # 存储每折的指标 pos_precisions = [] pos_recalls = [] pos_fmeasures = [] neg_precisions = [] neg_recalls = [] neg_fmeasures = [] for i in range(fold_count): # 拆分当前折的训练和测试数据 test_fold = train_features[i*fold_size : (i+1)*fold_size] train_fold = train_features[:i*fold_size] + train_features[(i+1)*fold_size:] # 训练分类器 classifier = NaiveBayesClassifier.train(train_fold) # 收集标签 true_labels = [label for (f, label) in test_fold] predicted_labels = [classifier.classify(f) for (f, label) in test_fold] # 计算当前折的指标(复用方式1的手动计算代码) # ...(此处省略统计和计算代码,可直接复制方式1的逻辑) # 把结果添加到对应列表中 # 最后计算平均指标 avg_pos_precision = sum(pos_precisions) / fold_count # 其他指标同理计算平均值
内容的提问来源于stack exchange,提问作者A.Bassit
相关产品推荐
相关产品推荐

