You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无标注数据下基于强化学习的图像旋转角预测:Keras实现求助

嘿,你的这个思路完全可行!这其实属于自监督学习的范畴——不需要标注的真实旋转角,完全可以通过「让旋转后的图像匹配度最高」这个目标来训练模型。我来一步步教你在Keras里实现,两种方案任你选,都是新手友好的:

核心思路回顾

你的需求是:输入一对图像(img1 和旋转后的 img2),模型预测旋转角度 θ_pred,把 img1 旋转 θ_pred 后和 img2 计算MSE,最小化这个MSE就是训练目标。关键是绕开Keras标准损失需要y_true的限制,直接用输入图像和预测结果计算损失。


方案一:自定义训练循环(新手易理解,灵活性高)

这种方式完全脱离Keras的model.fit框架,手动控制训练流程,适合你这种需要自定义损失逻辑的场景。

步骤1:构建预测模型

先做一个简单的CNN模型,输入是两张图像,输出是预测的旋转角度(建议用弧度,因为TensorFlow的旋转函数默认接受弧度):

import tensorflow as tf
from tensorflow import keras
from tensorflow.keras import layers
import numpy as np

def build_angle_predictor(input_shape):
    # 共享特征提取 backbone,让模型学习旋转相关的特征
    shared_backbone = keras.Sequential([
        layers.Conv2D(32, (3,3), activation='relu', input_shape=input_shape),
        layers.MaxPooling2D(),
        layers.Conv2D(64, (3,3), activation='relu'),
        layers.MaxPooling2D(),
        layers.Flatten(),
        layers.Dense(128, activation='relu'),
    ])
    
    # 定义两个输入:原始图和旋转后的图
    img1_input = layers.Input(shape=input_shape)
    img2_input = layers.Input(shape=input_shape)
    
    # 分别提取两张图的特征,再拼接
    feat_img1 = shared_backbone(img1_input)
    feat_img2 = shared_backbone(img2_input)
    combined_feats = layers.concatenate([feat_img1, feat_img2])
    
    # 输出预测的旋转角度(弧度,范围0~2π)
    pred_angle = layers.Dense(1)(combined_feats)
    
    return keras.Model(inputs=[img1_input, img2_input], outputs=pred_angle)

步骤2:定义损失计算函数

这个函数负责把预测角度应用到img1,然后和img2算MSE:

def calculate_mse_loss(img1, img2, pred_angle):
    # 用TensorFlow内置函数旋转图像,默认填充0,你可以改fill_mode='nearest'避免黑边
    rotated_img1 = tf.image.rotate(img1, pred_angle)
    # 计算批量的平均MSE
    return tf.reduce_mean(tf.square(rotated_img1 - img2))

步骤3:手动训练循环

用tf.GradientTape记录梯度,手动更新模型参数:

# 初始化模型和优化器
input_shape = (64, 64, 3)  # 根据你的图像尺寸调整
model = build_angle_predictor(input_shape)
optimizer = keras.optimizers.Adam(learning_rate=1e-4)

# 模拟你的数据集(替换成你真实的图像对加载逻辑)
def data_generator(batch_size=32):
    while True:
        # 生成随机测试图(实际中换成你的真实图像对)
        img1_batch = np.random.rand(batch_size, *input_shape)
        # 随机生成真实旋转角度(实际中你不需要这个,只是模拟数据用)
        true_angles = np.random.uniform(0, 2*np.pi, size=(batch_size,1))
        img2_batch = tf.image.rotate(img1_batch, true_angles).numpy()
        yield (img1_batch, img2_batch)

# 开始训练
epochs = 10
steps_per_epoch = 100  # 根据你的数据集大小调整
train_gen = data_generator(batch_size=32)

for epoch in range(epochs):
    print(f"=== Epoch {epoch+1}/{epochs} ===")
    total_loss = 0.0
    
    for step in range(steps_per_epoch):
        img1_batch, img2_batch = next(train_gen)
        
        # 用GradientTape记录梯度
        with tf.GradientTape() as tape:
            pred_angles = model([img1_batch, img2_batch], training=True)
            loss = calculate_mse_loss(img1_batch, img2_batch, pred_angles)
        
        # 计算梯度并更新模型参数
        grads = tape.gradient(loss, model.trainable_variables)
        optimizer.apply_gradients(zip(grads, model.trainable_variables))
        
        total_loss += loss.numpy()
    
    avg_loss = total_loss / steps_per_epoch
    print(f"Average Loss: {avg_loss:.4f}\n")

方案二:用模型内部损失(更贴合Keras原生流程)

如果你想继续用model.fit,可以通过model.add_loss()把损失直接嵌入模型内部,这样就不需要传入y_true了。

构建带内部损失的模型

def build_model_with_internal_loss(input_shape):
    shared_backbone = keras.Sequential([
        layers.Conv2D(32, (3,3), activation='relu', input_shape=input_shape),
        layers.MaxPooling2D(),
        layers.Conv2D(64, (3,3), activation='relu'),
        layers.MaxPooling2D(),
        layers.Flatten(),
        layers.Dense(128, activation='relu'),
    ])
    
    img1_input = layers.Input(shape=input_shape)
    img2_input = layers.Input(shape=input_shape)
    
    feat_img1 = shared_backbone(img1_input)
    feat_img2 = shared_backbone(img2_input)
    combined_feats = layers.concatenate([feat_img1, feat_img2])
    pred_angle = layers.Dense(1)(combined_feats)
    
    # 关键:在模型内部计算损失并添加到模型的损失列表
    rotated_img1 = tf.image.rotate(img1_input, pred_angle)
    mse_loss = tf.reduce_mean(tf.square(rotated_img1 - img2_input))
    model = keras.Model(inputs=[img1_input, img2_input], outputs=pred_angle)
    model.add_loss(mse_loss)  # 不需要y_true,模型自己会计算损失
    
    return model

用model.fit训练

model = build_model_with_internal_loss(input_shape)
model.compile(optimizer=keras.optimizers.Adam(learning_rate=1e-4))

# 训练时不需要传入y,直接传输入数据即可
model.fit(train_gen, epochs=10, steps_per_epoch=100)

新手注意事项

  1. 旋转函数细节:tf.image.rotate默认会在旋转后填充0,如果你不想有黑边,可以设置fill_mode='nearest'或者fill_mode='reflect'。
  2. 角度单位:一定要注意是弧度还是角度!TensorFlow的旋转函数只认弧度,如果你习惯用角度,记得在预测后转成弧度(角度 * np.pi / 180)。
  3. 模型调试:先从小图像(比如64x64)和简单模型开始测试,确保损失能持续下降,再换成你真实的大图像和复杂模型。
  4. 验证效果:训练一段时间后,拿一组测试图像对,用模型预测角度,手动旋转img1后和img2对比,看看视觉上是不是匹配,同时计算MSE是否真的变小了。

内容的提问来源于stack exchange,提问作者AnandJ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:44:13