You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于BERT的影评情感分类代码报错求助:token索引超出范围

影评情感分类BERT实现错误排查

问题背景

尝试用BERT、Transformers和TensorFlow实现影评情感分类,运行代码时出现索引越界错误,且模型未正确输出分类结果。

错误信息

Traceback (most recent call last):

  File "C:\Users\home\anaconda3\lib\site-packages\spyder_kernels\py3compat.py", line 356, in compat_exec
    exec(code, globals, locals)

  File "c:\users\home\downloads\mlp.py", line 60, in <module>
    dev_loss, dev_acc = evaluate(mlp, *dev, tf.keras.losses.MeanSquaredError())

  File "c:\users\home\downloads\mlp.py", line 46, in evaluate
    predictions = model(inputs)

  File "C:\Users\home\anaconda3\lib\site-packages\keras\utils\traceback_utils.py", line 67, in error_handler
    raise e.with_traceback(filtered_tb) from None

  File "c:\users\home\downloads\mlp.py", line 39, in call
    outputs = self.model(inputs)

  File "C:\Users\home\anaconda3\lib\site-packages\transformers\modeling_tf_utils.py", line 409, in run_call_with_unpacked_inputs
    return func(self, **unpacked_inputs)

  File "C:\Users\home\anaconda3\lib\site-packages\transformers\models\bert\modeling_tf_bert.py", line 1108, in call
    outputs = self.bert(

  File "C:\Users\home\anaconda3\lib\site-packages\transformers\modeling_tf_utils.py", line 409, in run_call_with_unpacked_inputs
    return func(self, **unpacked_inputs)

  File "C:\Users\home\anaconda3\lib\site-packages\transformers\models\bert\modeling_tf_bert.py", line 781, in call
    embedding_output = self.embeddings(

  File "C:\Users\home\anaconda3\lib\site-packages\transformers\models\bert\modeling_tf_bert.py", line 203, in call
    inputs_embeds = tf.gather(params=self.weight, indices=input_ids)

InvalidArgumentError: Exception encountered when calling layer "embeddings" (type TFBertEmbeddings).

indices[1174,8] = 29550 is not in [0, 28996) [Op:ResourceGather]

Call arguments received:
  • input_ids=tf.Tensor(shape=(1599, 73), dtype=int32)
  • position_ids=None
  • token_type_ids=tf.Tensor(shape=(1599, 73), dtype=int32)
  • inputs_embeds=None
  • past_key_values_length=0
  • training=False

问题根源

  1. Tokenizer与预训练模型不匹配:
    read_dataset函数默认使用bert-base-uncased的Tokenizer,但BertMLP类加载的是bert-base-cased预训练模型。两者词汇表范围不同:bert-base-cased的词汇表大小为28996,而bert-base-uncased的词汇表包含更多小写形式的token,生成的token ID可能超出bert-base-cased模型的索引范围,导致资源收集时的索引越界错误。

  2. 模型Call方法未正确连接分类头:
    当前call方法直接返回BERT模型的原始输出(包含last_hidden_state和pooler_output的元组),没有将BERT的输出传入后续的MLP分类头,既无法得到情感分类的预测值,也会导致损失计算时维度不匹配。

修复方案

1. 统一Tokenizer与预训练模型

将Tokenizer和模型的名称统一,要么都用bert-base-uncased,要么都用bert-base-cased。

2. 修正模型Call方法

从BERT的输出中提取[CLS] token对应的特征(或使用pooler_output),传入分类头得到最终的二分类预测结果。

3. 调整评估与训练逻辑

确保模型输出的预测值维度与标签匹配,同时二分类任务使用BinaryCrossentropy损失函数比MSE更合适。

完整修复代码

import numpy as np
import tensorflow as tf
from transformers import BertTokenizer, TFBertModel

def read_dataset(filename, model_name="bert-base-uncased"):
    """Reads a dataset from the specified path and returns sentences and labels"""

    tokenizer = BertTokenizer.from_pretrained(model_name)
    with open(filename, "r", encoding="utf-8") as f:
        lines = f.readlines()
        # preallocate memory for the data
        sents, labels = list(), np.empty((len(lines), 1), dtype=int)

        for i, line in enumerate(lines):
            text, str_label, _ = line.split("\t")
            labels[i] = int(str_label.split("=")[1] == "POS")
            sents.append(text)
    return dict(tokenizer(sents, padding=True, truncation=True, return_tensors="tf")), labels


class BertMLP(tf.keras.Model):
    def __init__(self, embed_batch_size=100, model_name="bert-base-uncased"):
        super(BertMLP, self).__init__()
        self.bs = embed_batch_size
        self.model = TFBertModel.from_pretrained(model_name)
        self.classification_head = tf.keras.models.Sequential(
            layers = [
                tf.keras.Input(shape=(self.model.config.hidden_size,)),
                tf.keras.layers.Dense(350, activation="tanh"),
                tf.keras.layers.Dense(200, activation="tanh"),
                tf.keras.layers.Dense(50, activation="tanh"),
                tf.keras.layers.Dense(1, activation="sigmoid", use_bias=False)
            ]
        )

    def call(self, inputs):
        # 获取BERT的输出,取[CLS] token的特征(第一个token,对应索引0)
        outputs = self.model(inputs)
        cls_output = outputs.last_hidden_state[:, 0, :]
        # 传入分类头得到预测结果
        return self.classification_head(cls_output)

def evaluate(model, inputs, labels, loss_func):
    mean_loss = tf.keras.metrics.Mean(name="eval_loss")
    accuracy = tf.keras.metrics.BinaryAccuracy(name="eval_accuracy")

    predictions = model(inputs, training=False)
    mean_loss(loss_func(labels, predictions))
    accuracy(labels, predictions)

    return mean_loss.result(), accuracy.result() * 100


if __name__ == "__main__":
    train = read_dataset("datasets/rt-polarity.train.vecs")
    dev = read_dataset("datasets/rt-polarity.dev.vecs")
    test = read_dataset("datasets/rt-polarity.test.vecs")

    mlp = BertMLP()
    mlp.compile(tf.keras.optimizers.SGD(learning_rate=0.01), loss=tf.keras.losses.BinaryCrossentropy())
    dev_loss, dev_acc = evaluate(mlp, *dev, tf.keras.losses.BinaryCrossentropy())
    print("Before training:", f"Dev Loss: {dev_loss}, Dev Acc: {dev_acc}")
    mlp.fit(*train, epochs=10, batch_size=10)
    dev_loss, dev_acc = evaluate(mlp, *dev, tf.keras.losses.BinaryCrossentropy())
    print("After training:", f"Dev Loss: {dev_loss}, Dev Acc: {dev_acc}")

内容的提问来源于stack exchange,提问作者Rewaster

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 11:01:05