You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用非默认Tokenizer初始化TensorFlow自定义Decoder类时遇AttributeError

解决TensorFlow Decoder类实例化时的AttributeError问题

错误详情

---------------------------------------------------------------------------
AttributeError                            Traceback (most recent call last)
<ipython-input-76-debfaa5162f4> in <cell line: 2>()
      1 # That will be sufficient for training. Create an instance of the decoder to test out:
----> 2 decoder = Decoder(tokenizer, UNITS)
      3 logits = decoder(ex_context, tf.expand_dims(ex_tar_in, axis=0))
      4 print(f'encoder output shape: (batch, s, units) {ex_context.shape}')
      5 print(f'input target tokens shape: (batch, t) {ex_tar_in.shape}')

3 frames
/usr/local/lib/python3.10/dist-packages/keras/src/layers/preprocessing/index_lookup.py in set_vocabulary(self, vocabulary, idf_weights)
    517             idf_weights = np.array(idf_weights)
    518 
---> 519         if vocabulary.size == 0:
    520             raise ValueError(
    521                 f"Cannot set an empty vocabulary, you passed {vocabulary}."

AttributeError: 'dict' object has no attribute 'size'

问题原因

原代码是为TensorFlow内置分词器编写的,但你使用了Transformers库的RobertaTokenizer,两者属性规则不匹配:

  • tf.keras.layers.StringLookup的vocabulary参数需要传入可迭代的词汇列表(或带size属性的TensorFlow词汇对象),但RobertaTokenizer的vocab是字典类型,没有size属性
  • 原代码中text_processor.size、自定义START_TOKEN等写法,不符合RobertaTokenizer的属性命名(比如用vocab_size表示词汇量,自带bos_token/eos_token作为起止标记)

解决方案

修改Decoder类的__init__方法,适配RobertaTokenizer的属性和词汇结构:

修改后的Decoder代码

# The decoder
"""
The decoder's job is to generate predictions for the next token at each 
location in the target sequence.

The Decoder Structure:
1. It looks up embeddings for each token in the target sequence.
2. It uses an RNN to process the target sequence, and keep track of what it has generated so far.
3. It uses RNN output as the "query" to the attention layer, when attending to the encoder's output.
4. At each location in the output it predicts the next token.

Note:
When training, the model predicts the next word at each location. So it's 
important that the information only flows in one direction through the model. 
The decoder uses a unidirectional (not bidirectional) 
RNN to process the target sequence.

> When running inference with this model it produces one word at a time, 
and those are fed back into the model.
"""

class Decoder(tf.keras.layers.Layer):
  
  @classmethod
  def add_method(cls, fun):
    setattr(cls, fun.__name__, fun)
    return fun

  def __init__(self, text_processor, units):
    super(Decoder, self).__init__()
    self.text_processor = text_processor  # RobertaTokenizer
    # 替换为RobertaTokenizer的词汇量属性
    self.vocab_size = text_processor.vocab_size  
    # 将Roberta的字典式词汇转为列表,适配StringLookup要求
    vocab_list = list(text_processor.get_vocab().keys())
    
    # 初始化StringLookup层,使用Roberta自带的特殊token
    self.word_to_id = tf.keras.layers.StringLookup(
        vocabulary=vocab_list,
        mask_token=text_processor.mask_token,
        oov_token=text_processor.unk_token)
    
    self.id_to_word = tf.keras.layers.StringLookup(
        vocabulary=vocab_list,
        mask_token=text_processor.mask_token,
        oov_token=text_processor.unk_token,
        invert=True)
    
    # 使用Roberta自带的起止token
    self.start_token = self.word_to_id(text_processor.bos_token)
    self.end_token = self.word_to_id(text_processor.eos_token)

    self.units = units


    # 1. 嵌入层将token ID转为向量
    self.embedding = tf.keras.layers.Embedding(self.vocab_size,
                                               units, mask_zero=True)

    # 2. RNN层记录已生成的序列信息
    self.rnn = tf.keras.layers.GRU(units,
                                   return_sequences=True,
                                   return_state=True,
                                   recurrent_initializer='glorot_uniform')

    # 3. 用RNN输出作为注意力层的查询向量
    self.attention = CrossAttention(units)

    # 4. 全连接层输出每个token的预测logits
    self.output_layer = tf.keras.layers.Dense(self.vocab_size)

关键修改点

  1. 词汇量获取:用text_processor.vocab_size替代text_processor.size,匹配RobertaTokenizer的属性命名
  2. 词汇格式转换:将text_processor.vocab字典转为列表list(text_processor.get_vocab().keys()),满足StringLookup的输入要求
  3. 起止token替换:使用Roberta自带的bos_token(<s>)和eos_token(</s>),替代自定义的START_TOKEN/END_TOKEN

环境版本:Python=3.8,tensorflow=2.14.0,transformers=4.34.0


内容的提问来源于stack exchange,提问作者Abdul Moez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 12:55:33