You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow中如何对可变尺寸图像执行conv2d_transpose?

解决TensorFlow中可变尺寸图像的转置卷积问题

你遇到的问题确实是因为输入的动态形状(None)导致的——当你用get_shape().as_list()获取静态形状时,可变维度会返回None,而Python无法对None和整数执行乘法运算,所以触发了TypeError。

要处理可变尺寸的输入,我们需要改用TensorFlow的动态形状操作,也就是在图运行时获取实际的输入尺寸,而不是依赖静态的形状信息。下面是修改后的完整实现:

修改后的代码

import tensorflow as tf

def conv2d_transpose(inputs, filters_shape, strides, name, padding="SAME", activation=None):
    filters = get_conv_filters(filters_shape, name)
    # 使用tf.shape获取动态形状(运行时的实际尺寸)
    inputs_shape = tf.shape(inputs)
    # 计算动态的输出形状
    output_shape = calc_output_shape(inputs_shape, filters_shape, strides, padding)
    strides = [1, *strides, 1]
    conv_transpose = tf.nn.conv2d_transpose(
        inputs, 
        filters, 
        output_shape=output_shape, 
        strides=strides, 
        padding=padding, 
        name=name+"transpose"
    )
    if activation is not None:
        conv_transpose = activation(conv_transpose)
    return conv_transpose

def get_conv_filters(filters_size, name):
    conv_weights = tf.Variable(tf.truncated_normal(filters_size), name=name + "weights")
    return conv_weights

def calc_output_shape(inputs_shape, filters_shape, strides, padding):
    # inputs_shape是张量,提取各个维度的动态值
    batch_size = inputs_shape[0]
    inputs_height = inputs_shape[1]
    inputs_width = inputs_shape[2]
    _, filters_height, filters_width, _, after_n_channel = filters_shape
    strides_height, strides_width = strides
    
    # 使用tf.cond处理padding分支(动态条件判断)
    def same_padding_output():
        output_height = tf.multiply(inputs_height, strides_height)
        output_width = tf.multiply(inputs_width, strides_width)
        return tf.stack([batch_size, output_height, output_width, after_n_channel])
    
    def valid_padding_output():
        output_height = tf.add(tf.multiply(tf.subtract(inputs_height, 1), strides_height), filters_height)
        output_width = tf.add(tf.multiply(tf.subtract(inputs_width, 1), strides_width), filters_width)
        return tf.stack([batch_size, output_height, output_width, after_n_channel])
    
    return tf.cond(tf.equal(padding, "SAME"), same_padding_output, valid_padding_output)

关键修改点说明

  1. 动态形状获取:
    把inputs.get_shape().as_list()替换成tf.shape(inputs),这样得到的是一个张量,包含运行时输入的实际尺寸值,而非静态的None。

  2. 动态计算输出形状:

    • 改用TensorFlow的运算函数(tf.multiply、tf.add、tf.subtract)替代Python的算术运算符,确保所有计算都在TensorFlow图中进行。
    • 用tf.cond处理padding的分支判断,因为静态的if-else无法处理动态张量的条件。
  3. 输出形状的构建:
    直接在calc_output_shape中用tf.stack构建输出形状张量,避免在外部处理None值。

测试代码

你可以用以下代码验证修改后的实现:

import numpy as np

input_images = tf.placeholder(tf.float32, [None, None, None, 3])
transpose_layer = conv2d_transpose(
    input_images, 
    filters_shape=[3,3,3,3], 
    strides=[2,2], 
    name="conv_3_transpose", 
    padding="SAME", 
    activation=tf.nn.relu
)

with tf.Session() as sess:
    sess.run(tf.global_variables_initializer())
    # 测试不同尺寸的输入
    test_input1 = np.random.rand(1, 64, 64, 3)
    output1 = sess.run(transpose_layer, feed_dict={input_images: test_input1})
    print(f"输入(1,64,64,3)的输出形状:{output1.shape}")  # 输出(1, 128, 128, 3)
    
    test_input2 = np.random.rand(2, 128, 96, 3)
    output2 = sess.run(transpose_layer, feed_dict={input_images: test_input2})
    print(f"输入(2,128,96,3)的输出形状:{output2.shape}")  # 输出(2, 256, 192, 3)

这样就能完美处理可变尺寸的输入图像了,无论你输入的height和width是多少,转置卷积都会根据实际尺寸计算正确的输出形状。

内容的提问来源于stack exchange,提问作者KiHyun Nam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:52:15