You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用CamemBERT提取特征后KNN分类准确率为0的问题求助

CamemBERT+KNN分类准确率0.0问题排查

问题背景

尝试用CamemBERT提取文本特征,再用KNN分类器完成分类任务,代码运行70分钟后准确率为0.0。以下是原代码及数据格式:

原代码

import numpy as np
import pandas as pd
from sklearn.model_selection import train_test_split
from transformers import AutoTokenizer, CamembertModel
from sklearn.neighbors import KNeighborsClassifier
from sklearn.metrics import accuracy_score
import torch

# Initialize the Camembert tokenizer and model
tokenizer = AutoTokenizer.from_pretrained("camembert-base")
model = CamembertModel.from_pretrained("camembert-base")

# Select the relevant columns from the DataFrame
cols = ["Intitulé (Ce champ doit respecter la nomenclature suivante : Code action – Libellé)_y", "Domaine sou domaine "]
df = df[cols]

# Convert DataFrame to a dictionary with row indices as keys and cell values as values
data = df.to_dict(orient='split')
data = dict(zip(data['index'], data['data']))

# Collect all the input texts into a list of strings
input_texts = [str(text) for text in data.values()]

# Set the batch size for processing
batch_size = 8

# Initialize lists to store the input features and labels
input_features_list = []
input_labels_list = []

# Process the data in batches
for i in range(0, len(input_texts), batch_size):
    batch_texts = input_texts[i:i + batch_size]

    # Tokenize the batch of texts
    inputs = tokenizer(batch_texts, return_tensors="pt", padding=True, truncation=True)

    # Get the model outputs for the batch
    with torch.no_grad():
        outputs = model(**inputs)

    # Extract the last hidden states and convert them to a numpy array
    last_hidden_states = outputs.last_hidden_state
    input_features = last_hidden_states[:, 0, :].numpy()

    # Add the input features and labels to the corresponding lists
    input_features_list.append(input_features)
    input_labels_list.extend(list(data.keys())[i:i + batch_size])

# Concatenate the input features and labels for all batches
input_features = np.concatenate(input_features_list)
input_labels = input_labels_list

# Split the data into training and testing sets
X_train, X_test, y_train, y_test = train_test_split(input_features, input_labels, test_size=0.2, random_state=42)

# Initialize and train the KNeighborsClassifier
neigh = KNeighborsClassifier(n_neighbors=3)
neigh.fit(X_train, y_train)

# Make predictions on the testing data
y_pred = neigh.predict(X_test)

# Calculate the accuracy
accuracy = accuracy_score(y_test, y_pred)

print("Accuracy:", accuracy)

数据格式

{
    'index': [row_index_1, row_index_2, ...],
    'columns': [column_name_1, column_name_2, ...],
    'data': [
        [cell_value_row_1_col_1, cell_value_row_1_col_2, ...],
        [cell_value_row_2_col_1, cell_value_row_2_col_2, ...],
        ...
    ]
}

核心错误与修复方案

1. 标签完全错误(最致命)

你当前把DataFrame的行索引当成了分类标签(input_labels_list.extend(list(data.keys())[i:i + batch_size])),但实际标签应该是第二列"Domaine sou domaine "的取值。由于train_test_split会打乱行索引,训练集和测试集的行索引完全不重叠,KNN根本无法预测正确的行号,直接导致准确率为0。

2. 文本输入错误

input_texts = [str(text) for text in data.values()]把每行的两个列值(文本+标签)拼成了字符串(比如"[文本内容, 标签值]"),但特征输入应该只使用第一列的文本,不该混入标签信息。

3. 运行效率极低

70分钟的运行时间完全不合理,建议用GPU加速(如果有),大幅减少特征提取时间。

4. KNN的高维适配问题

CamemBERT输出的768维特征会引发维度灾难,KNN的距离计算在高维空间中几乎失效,建议先做PCA降维,或者换用SVM、LightGBM等更适合高维数据的分类器。


修正后的关键代码片段

import numpy as np
import pandas as pd
from sklearn.model_selection import train_test_split
from transformers import AutoTokenizer, CamembertModel
from sklearn.neighbors import KNeighborsClassifier
from sklearn.metrics import accuracy_score
import torch

# 初始化Tokenizer和模型,优先用GPU
tokenizer = AutoTokenizer.from_pretrained("camembert-base")
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = CamembertModel.from_pretrained("camembert-base").to(device)

# 选择列
cols = ["Intitulé (Ce champ doit respecter la nomenclature suivante : Code action – Libellé)_y", "Domaine sou domaine "]
df = df[cols]

# 正确提取文本和标签(直接从df取,不用转字典)
input_texts = df[cols[0]].astype(str).tolist()
input_labels = df[cols[1]].tolist()

batch_size = 16  # 增大batch_size提升效率
input_features_list = []

# 批量提取特征
for i in range(0, len(input_texts), batch_size):
    batch_texts = input_texts[i:i + batch_size]
    # 把tokenizer输入移到对应设备
    inputs = tokenizer(batch_texts, return_tensors="pt", padding=True, truncation=True).to(device)
    
    with torch.no_grad():
        outputs = model(**inputs)
    
    # 提取CLS特征,移回CPU转numpy
    cls_features = outputs.last_hidden_state[:, 0, :].cpu().numpy()
    input_features_list.append(cls_features)

input_features = np.concatenate(input_features_list)

# 拆分数据集
X_train, X_test, y_train, y_test = train_test_split(input_features, input_labels, test_size=0.2, random_state=42)

# 训练KNN(可选:先做PCA降维)
# from sklearn.decomposition import PCA
# pca = PCA(n_components=128)
# X_train = pca.fit_transform(X_train)
# X_test = pca.transform(X_test)

neigh = KNeighborsClassifier(n_neighbors=3)
neigh.fit(X_train, y_train)

y_pred = neigh.predict(X_test)
accuracy = accuracy_score(y_test, y_pred)
print("Accuracy:", accuracy)

内容的提问来源于stack exchange,提问作者Wajih101

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 12:31:05