使用CamemBERT提取特征后KNN分类准确率为0的问题求助
CamemBERT+KNN分类准确率0.0问题排查
问题背景
尝试用CamemBERT提取文本特征,再用KNN分类器完成分类任务,代码运行70分钟后准确率为0.0。以下是原代码及数据格式:
原代码
import numpy as np import pandas as pd from sklearn.model_selection import train_test_split from transformers import AutoTokenizer, CamembertModel from sklearn.neighbors import KNeighborsClassifier from sklearn.metrics import accuracy_score import torch # Initialize the Camembert tokenizer and model tokenizer = AutoTokenizer.from_pretrained("camembert-base") model = CamembertModel.from_pretrained("camembert-base") # Select the relevant columns from the DataFrame cols = ["Intitulé (Ce champ doit respecter la nomenclature suivante : Code action – Libellé)_y", "Domaine sou domaine "] df = df[cols] # Convert DataFrame to a dictionary with row indices as keys and cell values as values data = df.to_dict(orient='split') data = dict(zip(data['index'], data['data'])) # Collect all the input texts into a list of strings input_texts = [str(text) for text in data.values()] # Set the batch size for processing batch_size = 8 # Initialize lists to store the input features and labels input_features_list = [] input_labels_list = [] # Process the data in batches for i in range(0, len(input_texts), batch_size): batch_texts = input_texts[i:i + batch_size] # Tokenize the batch of texts inputs = tokenizer(batch_texts, return_tensors="pt", padding=True, truncation=True) # Get the model outputs for the batch with torch.no_grad(): outputs = model(**inputs) # Extract the last hidden states and convert them to a numpy array last_hidden_states = outputs.last_hidden_state input_features = last_hidden_states[:, 0, :].numpy() # Add the input features and labels to the corresponding lists input_features_list.append(input_features) input_labels_list.extend(list(data.keys())[i:i + batch_size]) # Concatenate the input features and labels for all batches input_features = np.concatenate(input_features_list) input_labels = input_labels_list # Split the data into training and testing sets X_train, X_test, y_train, y_test = train_test_split(input_features, input_labels, test_size=0.2, random_state=42) # Initialize and train the KNeighborsClassifier neigh = KNeighborsClassifier(n_neighbors=3) neigh.fit(X_train, y_train) # Make predictions on the testing data y_pred = neigh.predict(X_test) # Calculate the accuracy accuracy = accuracy_score(y_test, y_pred) print("Accuracy:", accuracy)
数据格式
{ 'index': [row_index_1, row_index_2, ...], 'columns': [column_name_1, column_name_2, ...], 'data': [ [cell_value_row_1_col_1, cell_value_row_1_col_2, ...], [cell_value_row_2_col_1, cell_value_row_2_col_2, ...], ... ] }
核心错误与修复方案
1. 标签完全错误(最致命)
你当前把DataFrame的行索引当成了分类标签(input_labels_list.extend(list(data.keys())[i:i + batch_size])),但实际标签应该是第二列"Domaine sou domaine "的取值。由于train_test_split会打乱行索引,训练集和测试集的行索引完全不重叠,KNN根本无法预测正确的行号,直接导致准确率为0。
2. 文本输入错误
input_texts = [str(text) for text in data.values()]把每行的两个列值(文本+标签)拼成了字符串(比如"[文本内容, 标签值]"),但特征输入应该只使用第一列的文本,不该混入标签信息。
3. 运行效率极低
70分钟的运行时间完全不合理,建议用GPU加速(如果有),大幅减少特征提取时间。
4. KNN的高维适配问题
CamemBERT输出的768维特征会引发维度灾难,KNN的距离计算在高维空间中几乎失效,建议先做PCA降维,或者换用SVM、LightGBM等更适合高维数据的分类器。
修正后的关键代码片段
import numpy as np import pandas as pd from sklearn.model_selection import train_test_split from transformers import AutoTokenizer, CamembertModel from sklearn.neighbors import KNeighborsClassifier from sklearn.metrics import accuracy_score import torch # 初始化Tokenizer和模型,优先用GPU tokenizer = AutoTokenizer.from_pretrained("camembert-base") device = torch.device("cuda" if torch.cuda.is_available() else "cpu") model = CamembertModel.from_pretrained("camembert-base").to(device) # 选择列 cols = ["Intitulé (Ce champ doit respecter la nomenclature suivante : Code action – Libellé)_y", "Domaine sou domaine "] df = df[cols] # 正确提取文本和标签(直接从df取,不用转字典) input_texts = df[cols[0]].astype(str).tolist() input_labels = df[cols[1]].tolist() batch_size = 16 # 增大batch_size提升效率 input_features_list = [] # 批量提取特征 for i in range(0, len(input_texts), batch_size): batch_texts = input_texts[i:i + batch_size] # 把tokenizer输入移到对应设备 inputs = tokenizer(batch_texts, return_tensors="pt", padding=True, truncation=True).to(device) with torch.no_grad(): outputs = model(**inputs) # 提取CLS特征,移回CPU转numpy cls_features = outputs.last_hidden_state[:, 0, :].cpu().numpy() input_features_list.append(cls_features) input_features = np.concatenate(input_features_list) # 拆分数据集 X_train, X_test, y_train, y_test = train_test_split(input_features, input_labels, test_size=0.2, random_state=42) # 训练KNN(可选:先做PCA降维) # from sklearn.decomposition import PCA # pca = PCA(n_components=128) # X_train = pca.fit_transform(X_train) # X_test = pca.transform(X_test) neigh = KNeighborsClassifier(n_neighbors=3) neigh.fit(X_train, y_train) y_pred = neigh.predict(X_test) accuracy = accuracy_score(y_test, y_pred) print("Accuracy:", accuracy)
内容的提问来源于stack exchange,提问作者Wajih101
相关产品推荐
相关产品推荐

