如何在Keras中为对比学习生成批次以确保正样本对同组
类感知批次采样(Class-Aware Batch Sampling):对比学习的批次生成策略
你需要的这种批次生成策略叫做类感知批次采样,核心是保证每个批次内的每个类别至少包含2个样本,避免出现单类别样本孤立的情况,确保对比学习能有足够的正样本对来优化嵌入相似度。
实现方法(Keras/TensorFlow)
下面提供两种常用的实现方式,分别适配内存加载数据和大数据集场景:
1. 基于tf.data.Dataset的实现(适合内存可加载数据)
先构建类别到样本索引的映射,再自定义批次生成逻辑:
import numpy as np import tensorflow as tf # 假设已有输入数据和标签(示例标签:[0,0,0,0,1,1,1,2,2]) data = ... # 形状为 (样本数, 特征维度/视频帧维度) labels = np.array([0,0,0,0,1,1,1,2,2]) # 构建类别到样本索引的字典 class_to_indices = {} for idx, label in enumerate(labels): if label not in class_to_indices: class_to_indices[label] = [] class_to_indices[label].append(idx) classes = list(class_to_indices.keys()) def batch_generator(batch_size): while True: batch_indices = [] # 每次从类别中采样,保证每个类别至少取2个样本 while len(batch_indices) < batch_size: remaining_slots = batch_size - len(batch_indices) # 随机选一个类别,且该类别剩余可用样本≥2 available_classes = [c for c in classes if len([i for i in class_to_indices[c] if i not in batch_indices]) >=2] if not available_classes: # 无新类别可选时,从已选类别补充样本 selected_cls = np.random.choice([c for c in classes if len(class_to_indices[c])>0]) available_idxs = [i for i in class_to_indices[selected_cls] if i not in batch_indices] take = min(remaining_slots, len(available_idxs)) batch_indices.extend(np.random.choice(available_idxs, take, replace=False)) else: selected_cls = np.random.choice(available_classes) # 每个类别最少取2个,最多取剩余批次容量 take = np.random.randint(2, min(len(class_to_indices[selected_cls]), remaining_slots)+1) batch_indices.extend(np.random.choice(class_to_indices[selected_cls], take, replace=False)) # 打乱批次内样本顺序 np.random.shuffle(batch_indices) yield data[batch_indices], labels[batch_indices] # 转换为tf.data.Dataset batch_size = 5 dataset = tf.data.Dataset.from_generator( lambda: batch_generator(batch_size), output_signature=( tf.TensorSpec(shape=(batch_size,) + data.shape[1:], dtype=data.dtype), tf.TensorSpec(shape=(batch_size,), dtype=labels.dtype) ) ) # 添加缓存和预取优化 dataset = dataset.cache().prefetch(tf.data.AUTOTUNE)
2. 自定义Keras Sequence(适合大数据集/内存不足场景)
继承Keras的Sequence类,实现按需加载批次:
import numpy as np from tensorflow.keras.utils import Sequence class ClassConstrainedSequence(Sequence): def __init__(self, data, labels, batch_size): self.data = data self.labels = labels self.batch_size = batch_size # 构建类别到样本索引的映射 self.class_to_indices = {} for idx, label in enumerate(labels): if label not in self.class_to_indices: self.class_to_indices[label] = [] self.class_to_indices[label].append(idx) # 过滤掉样本数不足2的类别(避免无法生成正样本对) self.valid_classes = [c for c in self.class_to_indices if len(self.class_to_indices[c]) >=2] assert len(self.valid_classes) > 0, "存在类别样本数不足2个,无法生成有效对比批次" def __len__(self): # 估算总批次数 return len(self.labels) // self.batch_size def __getitem__(self, idx): batch_indices = [] while len(batch_indices) < self.batch_size: remaining = self.batch_size - len(batch_indices) # 随机选可用类别 selected_cls = np.random.choice(self.valid_classes) # 该类别未被选入当前批次的样本索引 available_idxs = [i for i in self.class_to_indices[selected_cls] if i not in batch_indices] if len(available_idxs) < 2: continue # 抽取2到剩余容量之间的样本数 take = np.random.randint(2, min(len(available_idxs), remaining)+1) batch_indices.extend(np.random.choice(available_idxs, take, replace=False)) # 打乱批次内顺序 np.random.shuffle(batch_indices) return self.data[batch_indices], self.labels[batch_indices] def on_epoch_end(self): # 每个epoch后打乱类别内的样本顺序,增加随机性 for cls in self.class_to_indices: np.random.shuffle(self.class_to_indices[cls]) # 使用示例 batch_size = 5 train_seq = ClassConstrainedSequence(data, labels, batch_size) # 传入模型训练 model.fit(train_seq, epochs=10)
为什么这个策略有效?
标准随机批次采样可能出现单个类别仅1个样本的情况,导致对比学习损失计算时正样本对不足,模型无法有效学习同类样本的相似嵌入。类感知批次采样保证每个批次内每个类别至少有2个样本,确保训练过程中始终有足够的正样本对来优化嵌入空间,让同类样本的嵌入更接近,异类样本的嵌入更疏远。
内容的提问来源于stack exchange,提问作者knightofcake
相关产品推荐
相关产品推荐

