无标注数据下基于强化学习的图像旋转角预测:Keras实现求助
嘿,你的这个思路完全可行!这其实属于自监督学习的范畴——不需要标注的真实旋转角,完全可以通过「让旋转后的图像匹配度最高」这个目标来训练模型。我来一步步教你在Keras里实现,两种方案任你选,都是新手友好的:
核心思路回顾
你的需求是:输入一对图像(img1 和旋转后的 img2),模型预测旋转角度 θ_pred,把 img1 旋转 θ_pred 后和 img2 计算MSE,最小化这个MSE就是训练目标。关键是绕开Keras标准损失需要y_true的限制,直接用输入图像和预测结果计算损失。
方案一:自定义训练循环(新手易理解,灵活性高)
这种方式完全脱离Keras的model.fit框架,手动控制训练流程,适合你这种需要自定义损失逻辑的场景。
步骤1:构建预测模型
先做一个简单的CNN模型,输入是两张图像,输出是预测的旋转角度(建议用弧度,因为TensorFlow的旋转函数默认接受弧度):
import tensorflow as tf from tensorflow import keras from tensorflow.keras import layers import numpy as np def build_angle_predictor(input_shape): # 共享特征提取 backbone,让模型学习旋转相关的特征 shared_backbone = keras.Sequential([ layers.Conv2D(32, (3,3), activation='relu', input_shape=input_shape), layers.MaxPooling2D(), layers.Conv2D(64, (3,3), activation='relu'), layers.MaxPooling2D(), layers.Flatten(), layers.Dense(128, activation='relu'), ]) # 定义两个输入:原始图和旋转后的图 img1_input = layers.Input(shape=input_shape) img2_input = layers.Input(shape=input_shape) # 分别提取两张图的特征,再拼接 feat_img1 = shared_backbone(img1_input) feat_img2 = shared_backbone(img2_input) combined_feats = layers.concatenate([feat_img1, feat_img2]) # 输出预测的旋转角度(弧度,范围0~2π) pred_angle = layers.Dense(1)(combined_feats) return keras.Model(inputs=[img1_input, img2_input], outputs=pred_angle)
步骤2:定义损失计算函数
这个函数负责把预测角度应用到img1,然后和img2算MSE:
def calculate_mse_loss(img1, img2, pred_angle): # 用TensorFlow内置函数旋转图像,默认填充0,你可以改fill_mode='nearest'避免黑边 rotated_img1 = tf.image.rotate(img1, pred_angle) # 计算批量的平均MSE return tf.reduce_mean(tf.square(rotated_img1 - img2))
步骤3:手动训练循环
用tf.GradientTape记录梯度,手动更新模型参数:
# 初始化模型和优化器 input_shape = (64, 64, 3) # 根据你的图像尺寸调整 model = build_angle_predictor(input_shape) optimizer = keras.optimizers.Adam(learning_rate=1e-4) # 模拟你的数据集(替换成你真实的图像对加载逻辑) def data_generator(batch_size=32): while True: # 生成随机测试图(实际中换成你的真实图像对) img1_batch = np.random.rand(batch_size, *input_shape) # 随机生成真实旋转角度(实际中你不需要这个,只是模拟数据用) true_angles = np.random.uniform(0, 2*np.pi, size=(batch_size,1)) img2_batch = tf.image.rotate(img1_batch, true_angles).numpy() yield (img1_batch, img2_batch) # 开始训练 epochs = 10 steps_per_epoch = 100 # 根据你的数据集大小调整 train_gen = data_generator(batch_size=32) for epoch in range(epochs): print(f"=== Epoch {epoch+1}/{epochs} ===") total_loss = 0.0 for step in range(steps_per_epoch): img1_batch, img2_batch = next(train_gen) # 用GradientTape记录梯度 with tf.GradientTape() as tape: pred_angles = model([img1_batch, img2_batch], training=True) loss = calculate_mse_loss(img1_batch, img2_batch, pred_angles) # 计算梯度并更新模型参数 grads = tape.gradient(loss, model.trainable_variables) optimizer.apply_gradients(zip(grads, model.trainable_variables)) total_loss += loss.numpy() avg_loss = total_loss / steps_per_epoch print(f"Average Loss: {avg_loss:.4f}\n")
方案二:用模型内部损失(更贴合Keras原生流程)
如果你想继续用model.fit,可以通过model.add_loss()把损失直接嵌入模型内部,这样就不需要传入y_true了。
构建带内部损失的模型
def build_model_with_internal_loss(input_shape): shared_backbone = keras.Sequential([ layers.Conv2D(32, (3,3), activation='relu', input_shape=input_shape), layers.MaxPooling2D(), layers.Conv2D(64, (3,3), activation='relu'), layers.MaxPooling2D(), layers.Flatten(), layers.Dense(128, activation='relu'), ]) img1_input = layers.Input(shape=input_shape) img2_input = layers.Input(shape=input_shape) feat_img1 = shared_backbone(img1_input) feat_img2 = shared_backbone(img2_input) combined_feats = layers.concatenate([feat_img1, feat_img2]) pred_angle = layers.Dense(1)(combined_feats) # 关键:在模型内部计算损失并添加到模型的损失列表 rotated_img1 = tf.image.rotate(img1_input, pred_angle) mse_loss = tf.reduce_mean(tf.square(rotated_img1 - img2_input)) model = keras.Model(inputs=[img1_input, img2_input], outputs=pred_angle) model.add_loss(mse_loss) # 不需要y_true,模型自己会计算损失 return model
用model.fit训练
model = build_model_with_internal_loss(input_shape) model.compile(optimizer=keras.optimizers.Adam(learning_rate=1e-4)) # 训练时不需要传入y,直接传输入数据即可 model.fit(train_gen, epochs=10, steps_per_epoch=100)
新手注意事项
- 旋转函数细节:
tf.image.rotate默认会在旋转后填充0,如果你不想有黑边,可以设置fill_mode='nearest'或者fill_mode='reflect'。 - 角度单位:一定要注意是弧度还是角度!TensorFlow的旋转函数只认弧度,如果你习惯用角度,记得在预测后转成弧度(
角度 * np.pi / 180)。 - 模型调试:先从小图像(比如64x64)和简单模型开始测试,确保损失能持续下降,再换成你真实的大图像和复杂模型。
- 验证效果:训练一段时间后,拿一组测试图像对,用模型预测角度,手动旋转
img1后和img2对比,看看视觉上是不是匹配,同时计算MSE是否真的变小了。
内容的提问来源于stack exchange,提问作者AnandJ
相关产品推荐
相关产品推荐

