You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BERT自然语言处理模型训练时输入形状不匹配问题求助

Fixing Shape Mismatch Error in BERT NLP Model

Looks like you're hitting a classic shape mismatch issue—your model expects input sequences of length 50, but your training data (X_train) is coming in with length 52. Let's break down how to fix this:

Root Cause

Your model's input layers are explicitly defined to accept sequences of length 50:

idx = layers.Input((50), dtype="int32", name="input_idx")

But when you feed X_train to model.fit(), one or more of the input tensors (input_ids, masks, segments) has a shape of (None, 52), which doesn't align with the model's expected input shape.

Step-by-Step Solutions

The most likely culprit is your tokenization step not properly truncating/padding sequences to 50 tokens. Make sure you're using the BERT tokenizer with explicit max_length, padding, and truncation settings:

tokenizer = transformers.BertTokenizer.from_pretrained("bert-base-uncased")

# Process your text data
encoded_data = tokenizer(
    your_text_list,  # Replace with your actual text data
    max_length=50,
    padding="max_length",  # Pad shorter sequences to 50 tokens
    truncation=True,  # Truncate longer sequences to 50 tokens
    return_tensors="tf"
)

# Extract the tensors to use as X_train/X_test
X_train = [encoded_data["input_ids"], encoded_data["attention_mask"], encoded_data["token_type_ids"]]

This ensures every input sequence is exactly 50 tokens long, matching your model's input layer definition.

2. Verify Input Shapes Before Training

Double-check the shapes of your input data to confirm the fix:

print("Input IDs shape:", X_train[0].shape)
print("Attention Mask shape:", X_train[1].shape)
print("Token Type IDs shape:", X_train[2].shape)

All should output (num_samples, 50). If they still show 52, go back to your preprocessing code—you might have forgotten to apply the same tokenizer settings to both training and test data.

3. Adjust Model Input Length (If Necessary)

If you actually need to use sequence length 52 (e.g., your data has meaningful context beyond 50 tokens), update your model's input layers to match:

idx = layers.Input((52), dtype="int32", name="input_idx")
masks = layers.Input((52), dtype="int32", name="input_masks")
segments = layers.Input((52), dtype="int32", name="input_segments")

Then recompile your model and retrain. Note: This will change the BERT output shape to (None,52,768), but your GlobalAveragePooling1D layer will handle this fine.

Key Note

Always ensure your training and test data are preprocessed with the exact same tokenizer settings. If you padded/truncated training data to 50 but left test data at 52, you'll hit the same error during prediction.

内容的提问来源于stack exchange,提问作者Radhika Singh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 23:17:43