如何在Keras模型中添加POS标签处理层?
报错原因
你写的posify_data是面向Python原生字符串实现的逻辑,而Keras的Lambda层在前向传播时传入的输入是tf.Tensor张量对象,原生张量没有split方法,直接调用自然会报错。你需要用tf.py_function把自定义Python逻辑包装成TensorFlow可识别的操作,再封装成层加入模型。
实现方式
方式一:Lambda层包装
import tensorflow as tf import nltk import numpy as np def posify_tensor(input_txt): # 包装适配张量输入的POS转换原生逻辑 def posify_py(txt): # 单样本输入时:张量转numpy字节数组再转字符串 if len(txt.shape) == 0: txt_str = txt.numpy().decode('utf-8') return ' '.join([pair[1] for pair in nltk.pos_tag(txt_str.split())]) # 批量输入时遍历处理每个样本 res = [] for item in txt.numpy(): item_str = item.decode('utf-8') res.append(' '.join([pair[1] for pair in nltk.pos_tag(item_str.split())])) return np.array(res) # 用tf.py_function包装Python函数,指定输入输出类型 result = tf.py_function(func=posify_py, inp=[input_txt], Tout=tf.string) # 固定张量形状,避免后续层报形状不匹配错误 result.set_shape(input_txt.shape) return result # 组装模型,在最前面加入POS转换层 model = tf.keras.Sequential([ tf.keras.layers.Lambda(posify_tensor), encoder, tf.keras.layers.Embedding(input_dim=len(encoder.get_vocabulary()), output_dim=64, mask_zero=True), tf.keras.layers.Bidirectional(tf.keras.layers.LSTM(64)), tf.keras.layers.Dense(64, activation='relu'), tf.keras.layers.Dense(len(encoded_lbls), activation='softmax') ])
方式二:自定义Keras层(推荐)
自定义层的序列化兼容性更好,模型保存加载不会出现Lambda层常见的序列化失败问题:
class POSLayer(tf.keras.layers.Layer): def __init__(self, **kwargs): super().__init__(**kwargs) def call(self, input_txt): def posify_py(txt): if len(txt.shape) == 0: txt_str = txt.numpy().decode('utf-8') return ' '.join([pair[1] for pair in nltk.pos_tag(txt_str.split())]) res = [] for item in txt.numpy(): item_str = item.decode('utf-8') res.append(' '.join([pair[1] for pair in nltk.pos_tag(item_str.split())])) return np.array(res) result = tf.py_function(func=posify_py, inp=[input_txt], Tout=tf.string) result.set_shape(input_txt.shape) return result # 组装模型 model = tf.keras.Sequential([ POSLayer(), encoder, tf.keras.layers.Embedding(input_dim=len(encoder.get_vocabulary()), output_dim=64, mask_zero=True), tf.keras.layers.Bidirectional(tf.keras.layers.LSTM(64)), tf.keras.layers.Dense(64, activation='relu'), tf.keras.layers.Dense(len(encoded_lbls), activation='softmax') ])
注意事项
- 用
tf.py_function包装的Python逻辑无法在GPU上运行,执行时会自动切换到CPU,性能会比纯TensorFlow原生操作低。如果对训练推理性能要求高,建议提前离线把所有文本转换为POS标注序列,再喂给模型训练,不需要把这一步集成到模型结构里。 - 如果你后续需要保存加载模型,使用自定义层的方式仅需额外配置
custom_objects参数即可正常加载,比Lambda层更稳定。
内容的提问来源于stack exchange,提问作者Alaa M.
相关产品推荐
相关产品推荐

