BERT自然语言处理模型训练时输入形状不匹配问题求助
Looks like you're hitting a classic shape mismatch issue—your model expects input sequences of length 50, but your training data (X_train) is coming in with length 52. Let's break down how to fix this:
Root Cause
Your model's input layers are explicitly defined to accept sequences of length 50:
idx = layers.Input((50), dtype="int32", name="input_idx")
But when you feed X_train to model.fit(), one or more of the input tensors (input_ids, masks, segments) has a shape of (None, 52), which doesn't align with the model's expected input shape.
Step-by-Step Solutions
1. Fix Text Preprocessing (Recommended)
The most likely culprit is your tokenization step not properly truncating/padding sequences to 50 tokens. Make sure you're using the BERT tokenizer with explicit max_length, padding, and truncation settings:
tokenizer = transformers.BertTokenizer.from_pretrained("bert-base-uncased") # Process your text data encoded_data = tokenizer( your_text_list, # Replace with your actual text data max_length=50, padding="max_length", # Pad shorter sequences to 50 tokens truncation=True, # Truncate longer sequences to 50 tokens return_tensors="tf" ) # Extract the tensors to use as X_train/X_test X_train = [encoded_data["input_ids"], encoded_data["attention_mask"], encoded_data["token_type_ids"]]
This ensures every input sequence is exactly 50 tokens long, matching your model's input layer definition.
2. Verify Input Shapes Before Training
Double-check the shapes of your input data to confirm the fix:
print("Input IDs shape:", X_train[0].shape) print("Attention Mask shape:", X_train[1].shape) print("Token Type IDs shape:", X_train[2].shape)
All should output (num_samples, 50). If they still show 52, go back to your preprocessing code—you might have forgotten to apply the same tokenizer settings to both training and test data.
3. Adjust Model Input Length (If Necessary)
If you actually need to use sequence length 52 (e.g., your data has meaningful context beyond 50 tokens), update your model's input layers to match:
idx = layers.Input((52), dtype="int32", name="input_idx") masks = layers.Input((52), dtype="int32", name="input_masks") segments = layers.Input((52), dtype="int32", name="input_segments")
Then recompile your model and retrain. Note: This will change the BERT output shape to (None,52,768), but your GlobalAveragePooling1D layer will handle this fine.
Key Note
Always ensure your training and test data are preprocessed with the exact same tokenizer settings. If you padded/truncated training data to 50 but left test data at 52, you'll hit the same error during prediction.
内容的提问来源于stack exchange,提问作者Radhika Singh

