TensorFlow中如何对可变尺寸图像执行conv2d_transpose?
解决TensorFlow中可变尺寸图像的转置卷积问题
你遇到的问题确实是因为输入的动态形状(None)导致的——当你用get_shape().as_list()获取静态形状时,可变维度会返回None,而Python无法对None和整数执行乘法运算,所以触发了TypeError。
要处理可变尺寸的输入,我们需要改用TensorFlow的动态形状操作,也就是在图运行时获取实际的输入尺寸,而不是依赖静态的形状信息。下面是修改后的完整实现:
修改后的代码
import tensorflow as tf def conv2d_transpose(inputs, filters_shape, strides, name, padding="SAME", activation=None): filters = get_conv_filters(filters_shape, name) # 使用tf.shape获取动态形状(运行时的实际尺寸) inputs_shape = tf.shape(inputs) # 计算动态的输出形状 output_shape = calc_output_shape(inputs_shape, filters_shape, strides, padding) strides = [1, *strides, 1] conv_transpose = tf.nn.conv2d_transpose( inputs, filters, output_shape=output_shape, strides=strides, padding=padding, name=name+"transpose" ) if activation is not None: conv_transpose = activation(conv_transpose) return conv_transpose def get_conv_filters(filters_size, name): conv_weights = tf.Variable(tf.truncated_normal(filters_size), name=name + "weights") return conv_weights def calc_output_shape(inputs_shape, filters_shape, strides, padding): # inputs_shape是张量,提取各个维度的动态值 batch_size = inputs_shape[0] inputs_height = inputs_shape[1] inputs_width = inputs_shape[2] _, filters_height, filters_width, _, after_n_channel = filters_shape strides_height, strides_width = strides # 使用tf.cond处理padding分支(动态条件判断) def same_padding_output(): output_height = tf.multiply(inputs_height, strides_height) output_width = tf.multiply(inputs_width, strides_width) return tf.stack([batch_size, output_height, output_width, after_n_channel]) def valid_padding_output(): output_height = tf.add(tf.multiply(tf.subtract(inputs_height, 1), strides_height), filters_height) output_width = tf.add(tf.multiply(tf.subtract(inputs_width, 1), strides_width), filters_width) return tf.stack([batch_size, output_height, output_width, after_n_channel]) return tf.cond(tf.equal(padding, "SAME"), same_padding_output, valid_padding_output)
关键修改点说明
动态形状获取:
把inputs.get_shape().as_list()替换成tf.shape(inputs),这样得到的是一个张量,包含运行时输入的实际尺寸值,而非静态的None。动态计算输出形状:
- 改用TensorFlow的运算函数(
tf.multiply、tf.add、tf.subtract)替代Python的算术运算符,确保所有计算都在TensorFlow图中进行。 - 用
tf.cond处理padding的分支判断,因为静态的if-else无法处理动态张量的条件。
- 改用TensorFlow的运算函数(
输出形状的构建:
直接在calc_output_shape中用tf.stack构建输出形状张量,避免在外部处理None值。
测试代码
你可以用以下代码验证修改后的实现:
import numpy as np input_images = tf.placeholder(tf.float32, [None, None, None, 3]) transpose_layer = conv2d_transpose( input_images, filters_shape=[3,3,3,3], strides=[2,2], name="conv_3_transpose", padding="SAME", activation=tf.nn.relu ) with tf.Session() as sess: sess.run(tf.global_variables_initializer()) # 测试不同尺寸的输入 test_input1 = np.random.rand(1, 64, 64, 3) output1 = sess.run(transpose_layer, feed_dict={input_images: test_input1}) print(f"输入(1,64,64,3)的输出形状:{output1.shape}") # 输出(1, 128, 128, 3) test_input2 = np.random.rand(2, 128, 96, 3) output2 = sess.run(transpose_layer, feed_dict={input_images: test_input2}) print(f"输入(2,128,96,3)的输出形状:{output2.shape}") # 输出(2, 256, 192, 3)
这样就能完美处理可变尺寸的输入图像了,无论你输入的height和width是多少,转置卷积都会根据实际尺寸计算正确的输出形状。
内容的提问来源于stack exchange,提问作者KiHyun Nam
相关产品推荐
相关产品推荐

