使用BanglaBERT实现知识蒸馏时遇'Tensor'无numpy属性错误
问题背景
在基于预训练BanglaBERT实现知识蒸馏时,自定义Distiller类的preprocess_data函数中执行text.numpy().decode('utf-8')时抛出错误:'Tensor' object has no attribute 'numpy'。传入的张量类型为tensorflow.python.framework.ops.Tensor,尝试转为EagerTensor也无法解决问题。
根本原因
在Keras的train_step方法中,当模型以Graph模式运行(model.fit默认采用该模式)时,传入的text是Graph Tensor,而非EagerTensor。Graph模式下Tensor的值无法直接通过numpy()获取——该方法仅在Eager模式(即时执行)下有效。你之前的tf.convert_to_tensor(text)操作无效,因为输入本身已是Tensor,转换后仍为Graph Tensor。
此外,当前代码中混用了Python原生操作(如Pandas DataFrame构建、Keras Tokenizer的texts_to_sequences、Python内置max函数),这些操作无法在Graph模式的计算图中正常执行,进一步加剧了问题。
解决方案
方案1:将预处理逻辑移至数据管道(推荐)
将文本解码、Tokenize、归一化等操作放到tf.data.Dataset的map方法中,用tf.py_function包装Python原生逻辑,确保在Eager上下文执行:
# 定义独立的预处理函数 def preprocess_fn(text, label, tokenizer, max_len_pad, onehot_encoder): # 解码Tensor为字符串 text_str = text.numpy().decode('utf-8') # Tokenize seq = tokenizer.texts_to_sequences([text_str])[0] # 避免除以0 seq_max = max(seq) if seq else 1.0 seq_norm = [x / seq_max for x in seq] # 填充序列 seq_norm = pad_sequences([seq_norm], padding='post', maxlen=max_len_pad, dtype='float32')[0] # 处理标签 label_np = label.numpy() onehot_label = onehot_encoder.transform([label_np])[0] return seq_norm, onehot_label # 在数据加载时应用预处理 def tf_preprocess(text, label): seq_norm, onehot_label = tf.py_function( preprocess_fn, inp=[text, label, tokenizer, max_len_pad, onehot_encoder], Tout=[tf.float32, tf.float32] # 根据实际标签类型调整 ) # 显式设置张量形状,避免Keras报错 seq_norm.set_shape((max_len_pad,)) onehot_label.set_shape((onehot_encoder.categories_[0].size,)) return seq_norm, onehot_label # 构建数据管道 train_dataset = train_dataset.map(tf_preprocess)
然后修改Distiller的train_step,直接使用预处理后的数据:
def train_step(self, data): x, y = data # x已预处理完成,y为独热编码标签 # 若必须在train_step中计算教师软标签,需用tf.py_function包装 def get_teacher_preds(text_tensor, label_tensor): text_str = text_tensor.numpy().decode('utf-8') label_np = label_tensor.numpy() teacher_data = pd.DataFrame({'sentences': [text_str], 'label': [label_np]}) _, teacher_output, _ = self.teacher.eval_model(teacher_data) return tf.nn.softmax(teacher_output) teacher_predictions = tf.py_function( get_teacher_preds, inp=[text, label], Tout=[tf.float32] )[0] # 后续损失计算、梯度更新逻辑保持不变...
方案2:在preprocess_data中用tf.py_function包装逻辑
直接修改Distiller类的preprocess_data方法,用tf.py_function包裹所有Python原生操作:
# 预处理函数 def preprocess_data(self, text, label): def tf_preprocess(text_tensor, label_tensor): # 解码为字符串 text_str = text_tensor.numpy().decode('utf-8') # Tokenize与归一化 seq = self.tokenizer.texts_to_sequences([text_str])[0] seq_max = max(seq) if seq else 1.0 seq_norm = [x / seq_max for x in seq] seq_norm = pad_sequences([seq_norm], padding='post', maxlen=self.max_len_pad, dtype='float32')[0] # 处理标签 label_np = label_tensor.numpy() onehot_label = self.onehot_encoder.transform([label_np])[0] return seq_norm, onehot_label # 用tf.py_function包装并指定输出类型 seq_norm, onehot_label = tf.py_function( tf_preprocess, inp=[text, label], Tout=[tf.float32, tf.float32] ) # 设置张量形状 seq_norm.set_shape((self.max_len_pad,)) onehot_label.set_shape((self.onehot_encoder.categories_[0].size,)) return seq_norm, onehot_label
方案3:替换为TensorFlow原生操作(最优)
尽可能用TensorFlow原生API替代Python原生操作,让整个预处理逻辑在Graph模式下运行,避免依赖tf.py_function:
- 使用
tf.keras.layers.TextVectorization替代Keras Tokenizer,实现TensorFlow原生的文本转序列 - 用
tf.reduce_max替代Python内置max函数 - 用
tf.pad替代pad_sequences
示例:
# 初始化TextVectorization层 vectorize_layer = tf.keras.layers.TextVectorization( max_tokens=tokenizer.num_words, output_mode='int', output_sequence_length=self.max_len_pad ) # 适配语料库 vectorize_layer.adapt(train_texts) # 预处理函数(全TensorFlow操作) def preprocess_data(self, text, label): # 文本转序列 seq = vectorize_layer(text) # 归一化(避免除以0) seq_max = tf.reduce_max(seq) seq_max = tf.cond(tf.equal(seq_max, 0), lambda: 1.0, lambda: tf.cast(seq_max, tf.float32)) seq_norm = tf.cast(seq, tf.float32) / seq_max # 处理标签(假设标签已转为数值类型) onehot_label = tf.one_hot(label, depth=self.onehot_encoder.categories_[0].size) return seq_norm, onehot_label
额外注意事项
- 教师模型的
eval_model方法如果依赖Pandas或Python原生操作,同样需要用tf.py_function包装,或者提前离线计算所有样本的软标签并保存,训练时直接加载使用,大幅提升训练效率。 - 若使用
tf.py_function,必须显式设置输出张量的形状,否则Keras无法自动推断输入形状,导致模型编译或训练报错。
内容的提问来源于stack exchange,提问作者Muntasir Ahmed

