Tensorflow字符级CNN输入维度错误排查与解决
字符级CNN输入维度问题排查与解决
最近我在给一个大型神经网络系统添加两层堆叠的字符级CNN时,卡在了输入维度相关的ValueError上。我的需求是先通过字符替换(根据大小写、是否为数字或字母分类)得到输入单词的正字法表示,再把它喂给CNN。虽然我知道用LSTM/RNN也能实现类似功能,但项目要求必须用CNN,所以只能硬着头皮解决这个问题。
市面上的CNN示例大多是针对MNIST这类图像数据集的,几乎找不到文本数据集的适配案例,我完全摸不清该怎么重塑字符嵌入,才能让它成为CNN的有效输入。下面是我最初尝试运行的代码片段:
# ... # shape = (batch size, max length of sentence, max length of word) self.char_ids = tf.placeholder(tf.int32, shape=[None, None, None], name="char_ids") # ... # Char embedding lookup _char_embeddings = tf.get_variable( name="_char_embeddings", dtype=tf.float32, shape=[self.config.nchars, self.config.dim_char]) char_embeddings = tf.nn.embedding_lookup(_char_embeddings, self.char_ids, name="char_embeddings") # Reshape for CNN? s = tf.shape(char_embeddings) char_embeddings = tf.reshape(char_embeddings, shape=[s[0]*s[1], self.config.dim_char, s[2]]) # Conv #1 conv1 = tf.layers.conv1d( inputs=char_embeddings, filters=64, kernel_size=3, padding="valid", activation=tf.nn.relu) # Conv #2 conv2 = tf.layers.conv1d( inputs=conv1, filters=64, kernel_size=3, padding="valid", activation=tf.nn.relu) pool2 = tf.layers.max_pooling1d(inputs=conv2, pool_size=2, strides=2) # Dense Layer output = tf.layers.dense(inputs=pool2, units=32, activation=tf.nn.relu) # ...
运行后直接报错,错误信息如下:
File "/home/emre/blstm-crf-ner/model/ner_model.py", line 159, in add_word_embeddings_op activation=tf.nn.relu) File "/home/emre/blstm-crf-ner/virtner/lib/python3.4/site-packages/tensorflow/python/layers/convolutional.py", line 411, in conv1d return layer.apply(inputs) File "/home/emre/blstm-crf-ner/virtner/lib/python3.4/site-packages/tensorflow/python/layers/base.py", line 809, in apply return self.__call__(inputs, *args, **kwargs) File "/home/emre/blstm-crf-ner/virtner/lib/python3.4/site-packages/tensorflow/python/layers/base.py", line 680, in __call__ self.build(input_shapes) File "/home/emre/blstm-crf-ner/virtner/lib/python3.4/site-packages/tensorflow/python/layers/convolutional.py", line 132, in build raise ValueError('The channel dimension of the inputs ' ValueError: The channel dimension of the inputs should be defined. Found `None`.
问题解决更新
后来我参考了一些NLP领域介绍CNN应用的博客,再加上vijay m的帮助,终于搞懂了问题所在:CNN必须预先指定输入维度,这和RNN/LSTM可以用sequence_length动态处理的方式不一样。
我调整了代码,提前对单词做了统一处理——短单词补全、长单词截断,固定了单词的最大长度,然后重塑嵌入维度时直接使用这个固定值。最终可以正常运行的代码片段如下:
# Char embedding lookup _char_embeddings = tf.get_variable( name="_char_embeddings", dtype=tf.float32, shape=[self.config.nchars, self.config.dim_char]) char_embeddings = tf.nn.embedding_lookup(_char_embeddings, self.char_ids, name="char_embeddings") # max_len_of_word: 20 # Just pad shorter words and truncate the longer ones. s = tf.shape(char_embeddings) char_embeddings = tf.reshape(char_embeddings, shape=[-1, self.config.dim_char, self.config.max_len_of_word]) # Conv #1 conv1 = tf.layers.conv1d( inputs=char_embeddings, filters=64, kernel_size=3, padding="valid", activation=tf.nn.relu) # Conv #2 conv2 = tf.layers.conv1d( inputs=conv1, filters=64, kernel_size=3, padding="valid", activation=tf.nn.relu) pool2 = tf.layers.max_pooling1d(inputs=conv2, pool_size=2, strides=2) # Dense Layer output = tf.layers.dense(inputs=pool2, units=32, activation=tf.nn.relu)
内容的提问来源于stack exchange,提问作者emrekgn
相关产品推荐
相关产品推荐

