CLIP多标签分类问题:替换Softmax为Sigmoid后结果异常如何解决?
问题原因及解决方法
你遇到的问题核心是温度参数(100.0)不适合Sigmoid激活:
- CLIP输出的图像与文本特征点积是余弦相似度,范围在[-1, 1]之间。
- 单标签场景用Softmax时,100.0的高温度会放大差异,让正确类别概率趋近于1,其他趋近于0。但Sigmoid是对每个元素单独计算概率,高温度下大部分点积(即使是负的)乘以100后也会接近0,而
Sigmoid(0)=0.5,导致所有概率都集中在0.5附近。
解决步骤
1. 降低温度参数
把缩放系数从100.0调低到合适值(比如5.0-20.0,可根据你的数据集调整),让点积缩放后的范围落在[-5,5]左右,这样Sigmoid能输出更极端的概率值。
2. 修改代码示例
替换Softmax为Sigmoid,并调整温度:
import torch from PIL import Image import open_clip model, _, preprocess = open_clip.create_model_and_transforms('ViT-B-32-quickgelu', pretrained='laion400m_e32') tokenizer = open_clip.get_tokenizer('ViT-B-32-quickgelu') image = preprocess(Image.open("CLIP.png")).unsqueeze(0) text = tokenizer(["a diagram", "a dog", "a cat"]) with torch.no_grad(), torch.cuda.amp.autocast(): image_features = model.encode_image(image) text_features = model.encode_text(text) image_features /= image_features.norm(dim=-1, keepdim=True) text_features /= text_features.norm(dim=-1, keepdim=True) # 调整温度为5.0(可根据实际效果微调) temperature = 5.0 text_probs = torch.sigmoid(temperature * image_features @ text_features.T) print("Label probs:", text_probs)
3. 可选:微调温度或模型
如果调整温度后效果仍不理想,可以:
- 在你的多标签数据集上微调温度参数(找到能让正负样本概率差异最大的值)。
- 对CLIP模型进行少量微调,适配多标签任务(冻结大部分参数,只微调分类头或最后几层)。
内容的提问来源于stack exchange,提问作者craaaft
相关产品推荐
相关产品推荐

