You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Keras中为对比学习生成批次以确保正样本对同组

类感知批次采样(Class-Aware Batch Sampling):对比学习的批次生成策略

你需要的这种批次生成策略叫做类感知批次采样,核心是保证每个批次内的每个类别至少包含2个样本,避免出现单类别样本孤立的情况,确保对比学习能有足够的正样本对来优化嵌入相似度。

实现方法(Keras/TensorFlow)

下面提供两种常用的实现方式,分别适配内存加载数据和大数据集场景:

1. 基于tf.data.Dataset的实现(适合内存可加载数据)

先构建类别到样本索引的映射,再自定义批次生成逻辑:

import numpy as np
import tensorflow as tf

# 假设已有输入数据和标签(示例标签:[0,0,0,0,1,1,1,2,2])
data = ...  # 形状为 (样本数, 特征维度/视频帧维度)
labels = np.array([0,0,0,0,1,1,1,2,2])

# 构建类别到样本索引的字典
class_to_indices = {}
for idx, label in enumerate(labels):
    if label not in class_to_indices:
        class_to_indices[label] = []
    class_to_indices[label].append(idx)
classes = list(class_to_indices.keys())

def batch_generator(batch_size):
    while True:
        batch_indices = []
        # 每次从类别中采样,保证每个类别至少取2个样本
        while len(batch_indices) < batch_size:
            remaining_slots = batch_size - len(batch_indices)
            # 随机选一个类别,且该类别剩余可用样本≥2
            available_classes = [c for c in classes if len([i for i in class_to_indices[c] if i not in batch_indices]) >=2]
            if not available_classes:
                # 无新类别可选时,从已选类别补充样本
                selected_cls = np.random.choice([c for c in classes if len(class_to_indices[c])>0])
                available_idxs = [i for i in class_to_indices[selected_cls] if i not in batch_indices]
                take = min(remaining_slots, len(available_idxs))
                batch_indices.extend(np.random.choice(available_idxs, take, replace=False))
            else:
                selected_cls = np.random.choice(available_classes)
                # 每个类别最少取2个,最多取剩余批次容量
                take = np.random.randint(2, min(len(class_to_indices[selected_cls]), remaining_slots)+1)
                batch_indices.extend(np.random.choice(class_to_indices[selected_cls], take, replace=False))
        
        # 打乱批次内样本顺序
        np.random.shuffle(batch_indices)
        yield data[batch_indices], labels[batch_indices]

# 转换为tf.data.Dataset
batch_size = 5
dataset = tf.data.Dataset.from_generator(
    lambda: batch_generator(batch_size),
    output_signature=(
        tf.TensorSpec(shape=(batch_size,) + data.shape[1:], dtype=data.dtype),
        tf.TensorSpec(shape=(batch_size,), dtype=labels.dtype)
    )
)
# 添加缓存和预取优化
dataset = dataset.cache().prefetch(tf.data.AUTOTUNE)

2. 自定义Keras Sequence(适合大数据集/内存不足场景)

继承Keras的Sequence类,实现按需加载批次:

import numpy as np
from tensorflow.keras.utils import Sequence

class ClassConstrainedSequence(Sequence):
    def __init__(self, data, labels, batch_size):
        self.data = data
        self.labels = labels
        self.batch_size = batch_size
        # 构建类别到样本索引的映射
        self.class_to_indices = {}
        for idx, label in enumerate(labels):
            if label not in self.class_to_indices:
                self.class_to_indices[label] = []
            self.class_to_indices[label].append(idx)
        # 过滤掉样本数不足2的类别(避免无法生成正样本对)
        self.valid_classes = [c for c in self.class_to_indices if len(self.class_to_indices[c]) >=2]
        assert len(self.valid_classes) > 0, "存在类别样本数不足2个,无法生成有效对比批次"

    def __len__(self):
        # 估算总批次数
        return len(self.labels) // self.batch_size

    def __getitem__(self, idx):
        batch_indices = []
        while len(batch_indices) < self.batch_size:
            remaining = self.batch_size - len(batch_indices)
            # 随机选可用类别
            selected_cls = np.random.choice(self.valid_classes)
            # 该类别未被选入当前批次的样本索引
            available_idxs = [i for i in self.class_to_indices[selected_cls] if i not in batch_indices]
            if len(available_idxs) < 2:
                continue
            # 抽取2到剩余容量之间的样本数
            take = np.random.randint(2, min(len(available_idxs), remaining)+1)
            batch_indices.extend(np.random.choice(available_idxs, take, replace=False))
        
        # 打乱批次内顺序
        np.random.shuffle(batch_indices)
        return self.data[batch_indices], self.labels[batch_indices]

    def on_epoch_end(self):
        # 每个epoch后打乱类别内的样本顺序,增加随机性
        for cls in self.class_to_indices:
            np.random.shuffle(self.class_to_indices[cls])

# 使用示例
batch_size = 5
train_seq = ClassConstrainedSequence(data, labels, batch_size)
# 传入模型训练
model.fit(train_seq, epochs=10)

为什么这个策略有效?

标准随机批次采样可能出现单个类别仅1个样本的情况,导致对比学习损失计算时正样本对不足,模型无法有效学习同类样本的相似嵌入。类感知批次采样保证每个批次内每个类别至少有2个样本,确保训练过程中始终有足够的正样本对来优化嵌入空间,让同类样本的嵌入更接近,异类样本的嵌入更疏远。

内容的提问来源于stack exchange,提问作者knightofcake

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 15:25:14