DistilBERT三分类预测报错:TypeError: only size-1 arrays can be converted to Python scalars 解决方案咨询
Hey, let's break down what's causing your error and fix it step by step!
The Core Problem
That TypeError: only size-1 arrays can be converted to Python scalars is popping up because your model is outputting full sequence hidden states (shape (batch_size, sequence_length, hidden_size)) instead of the 3-class logits you need for classification. Look at your tf_output example—it's a 2D array with rows for each token in the sentence, not a 1D array of 3 values for your three classes. On top of that, your softmax calculation was using the wrong axis, making the dimension mismatch even worse.
Fix 1: Make Sure Your Model is Built for 3-Class Classification
First, double-check that your model is set up to output 3 classification labels. If you're using Hugging Face's Transformers library, the easiest way is to use TFDistilBertForSequenceClassification—it automatically handles the <CLS> token (the token DistilBERT uses for classification) and outputs logits for your classes:
from transformers import TFDistilBertForSequenceClassification # Initialize model for 3-class classification model = TFDistilBertForSequenceClassification.from_pretrained( "distilbert-base-uncased", num_labels=3, # Critical: set this to your number of classes output_attentions=False, output_hidden_states=False )
If you built a custom model, you need to explicitly grab the <CLS> token's output (first token in the sequence) and pass it through a dense layer with 3 units:
class CustomDistilBERT(tf.keras.Model): def __init__(self, num_labels=3): super().__init__() self.distilbert = TFDistilBertModel.from_pretrained("distilbert-base-uncased") self.classifier = tf.keras.layers.Dense(num_labels, activation=None) def call(self, inputs): outputs = self.distilbert(inputs) # Grab the <CLS> token output (first position in the sequence) cls_output = outputs.last_hidden_state[:, 0, :] # Generate logits for 3 classes logits = self.classifier(cls_output) return logits
Fix 2: Rewrite Your Prediction Function
Your original code had issues with how you handled tokenizer output and model predictions. Here's the corrected version:
def predict_sentences_from_list(sentence_list, text, tokenizer, model, label_list): # Expand the dict to store predictions and confidences too predictions_dict = {"sentences": [], "predictions": [], "confidences": []} c_counter = 0 p_counter = 0 feedback_list = [] result = FRE(nlp(text)) print("FRE:", result) for sentence in sentence_list: # Use tokenizer() instead of encode() to get proper model inputs (input_ids + attention_mask) predict_input = tokenizer( sentence, truncation=True, padding="max_length", # Standardize sequence length max_length=128, # Adjust based on your sentence lengths return_tensors="tf" ) # Get model predictions—logits are stored in the logits field tf_output = model.predict(predict_input) logits = tf_output.logits # Shape: (1, 3) → 1 sample, 3 class logits print("Logits for sentence:", logits.numpy()) # Calculate softmax to get class probabilities (axis=1 for per-sample calculation) probabilities = tf.nn.softmax(logits, axis=1).numpy()[0] # Shape: (3,) print("Class probabilities:", probabilities) # Get the predicted label index = tf.math.argmax(probabilities).numpy() predicted_label = label_list[index] # Update counts and store results predictions_dict["sentences"].append(sentence) predictions_dict["predictions"].append(predicted_label) predictions_dict["confidences"].append(max(probabilities)) if predicted_label == "Claim": c_counter += 1 elif predicted_label == "Premise": p_counter += 1 return predictions_dict
Why This Works
tokenizer()returns a dictionary withinput_idsandattention_mask—this is the standard input format for Hugging Face TF models, which ensures your model processes the sentence correctly.- We extract
logitsdirectly from the model output—this is the array of 3 values your classification head produces, one for each class. - Using
axis=1intf.nn.softmaxensures we calculate probabilities across the 3 classes for each sentence, instead of the wrong axis which was causing dimension errors.
Extra Tip
For better efficiency, consider batch-processing your sentences instead of looping through them one by one. The tokenizer can handle lists of sentences directly, and the model will process them in a single predict call.
内容的提问来源于stack exchange,提问作者Philipp

