基于Deeplearning4j提取CNN层激活特征的技术问询
作为刚接触CNN的开发者,你用Deeplearning4j调用预训练VGG16提取fc2层激活值的思路很靠谱,先来看你的实现代码,再逐个解答你的问题:
你的实现代码
public static void main(String[] args) throws IOException { ComputationGraph vgg16transfer = getComputationGraph(); for (File file : new File(ImageClassifier.class.getClassLoader().getResource("mydirectory").getFile()).listFiles()) { Map<String, INDArray> stringINDArrayMap = extractTwo(file, vgg16transfer); //Extract the features from the last fully connected layers saveCompressed(file,stringINDArrayMap.get("fc2")); } } /** * Retrieves the VGG16 computation graph * @return ComputationGraph from the pretrained VGG16 * @throws IOException */ public static ComputationGraph getComputationGraph() throws IOException { ZooModel zooModel = new VGG16(); return (ComputationGraph) zooModel.initPretrained(PretrainedType.IMAGENET); } /** * Compresses the input INDArray and writes it to file * @param imageFile the original image file * @param array INDArray to be saved (features) * @throws IOException */ private static void saveCompressed(File imageFile, INDArray array) throws IOException { INDArray compress = BasicNDArrayCompressor.getInstance().compress(array); Nd4j.write(compress,new DataOutputStream(new FileOutputStream(new File("features/" + imageFile.getName()+ "feat")))); } /** * Given an input image and a ComputationGraph it calls the feedForward method after rescaling the image. * @param imageFile the image whose features need to be extracted * @param vgg16 the ComputationGraph to be used. * @return a map of activations for each layer * @throws IOException */ public static Map<String, INDArray> extractTwo(File imageFile, ComputationGraph vgg16) throws IOException { // Convert file to INDArray NativeImageLoader loader = new NativeImageLoader(224, 224, 3); INDArray image = loader.asMatrix(imageFile); // Mean subtraction pre-processing step for VGG DataNormalization scaler = new VGG16ImagePreProcessor(); scaler.transform(image); //Call the feedForward method to get a map of activations for each layer return vgg16.feedForward(image, false); }
问题1:代码是否确实提取了可用于后续任务的特征?
完全没问题!你的代码逻辑非常规范:
- 正确加载了预训练VGG16模型,用
feedForward(image, false)获取各层激活(false表示不计算反向传播,只做前向推理,效率更高) - 用
VGG16ImagePreProcessor完成了VGG要求的均值减法预处理,这一步是关键,保证输入符合预训练模型的预期 - 提取的fc2层是VGG16的倒数第二层全连接层,输出是4096维的全局特征,非常适合后续的图像检索、分类等任务
你可以简单验证:读取保存的特征文件,用Nd4j.read()加载后打印array.shape(),应该输出[1, 4096](每个图像对应一个4096维向量),说明特征提取正常。
问题2:如何对提取的特征进行PCA/白化处理?
PCA(降维)和白化(去除特征间相关性、统一方差)是常用的特征预处理手段,在Deeplearning4j里可以这样实现:
步骤1:收集所有样本特征
先把所有提取的fc2特征合并成一个大的INDArray,形状为[N, 4096](N是样本总数,每行对应一个图像的特征)。
步骤2:标准化(白化前置步骤)
先对特征做标准化,让每个维度均值为0、方差为1:
DataNormalization scaler = new StandardScaler(); scaler.fit(allFeatures); // 计算所有样本的均值和方差 INDArray normalizedFeatures = scaler.transform(allFeatures); // 应用标准化
步骤3:PCA降维
用Deeplearning4j的PCA类实现降维,比如把4096维降到256维:
PCA pca = new PCA(256); // 设置目标维度 pca.fit(normalizedFeatures); // 学习PCA变换 INDArray pcaFeatures = pca.transform(normalizedFeatures); // 应用降维
步骤4:白化处理
如果需要做白化(让特征协方差矩阵为单位矩阵),可以在PCA后手动实现:
// 计算协方差矩阵 INDArray covMatrix = Nd4j.covariance(normalizedFeatures); // 特征值分解 INDArray[] eigenDecomp = Nd4j.eig(covMatrix); INDArray eigenValues = eigenDecomp[0].real(); INDArray eigenVectors = eigenDecomp[1].real(); // 白化:(标准化特征) × 特征向量 × 特征值平方根的倒数(加小epsilon避免除以0) INDArray whitenedFeatures = normalizedFeatures.mmul(eigenVectors) .mul(Nd4j.diag(eigenValues.add(1e-8).sqrt().rdiv(1)));
问题3:能否将特征编码为VLAD格式?
当然可以!VLAD(Vector of Locally Aggregated Descriptors)是一种把局部特征聚合为全局特征的方法,虽然最初用于卷积层局部特征,但也能适配fc2的全局特征,实现步骤如下:
步骤1:用K-Means构建视觉词典
先对所有fc2特征做K-Means聚类,得到k个聚类中心(比如k=64):
KMeans kMeans = new KMeans.Builder() .nClusters(64) .maxIterations(50) .build(); kMeans.fit(allFeatures); INDArray centers = kMeans.getClusterCenters(); // 形状为[64, 4096]
步骤2:对单个特征计算VLAD
针对每个图像的fc2特征,计算其到每个聚类中心的残差,累加后归一化:
public static INDArray computeVLAD(INDArray singleFeature, INDArray centers) { int k = centers.rows(); int featDim = centers.columns(); INDArray vlad = Nd4j.zeros(k, featDim); // 初始化VLAD矩阵 // 找到距离当前特征最近的聚类中心 INDArray distances = Nd4j.zeros(k); for (int i = 0; i < k; i++) { distances.putScalar(i, singleFeature.distance2(centers.getRow(i))); } int nearestCenterIdx = Nd4j.argMin(distances, 0).getInt(0); // 计算残差并累加到对应中心的槽位 INDArray residual = singleFeature.sub(centers.getRow(nearestCenterIdx)); vlad.putRow(nearestCenterIdx, vlad.getRow(nearestCenterIdx).add(residual)); // 槽内L2归一化(原论文推荐) for (int i = 0; i < k; i++) { INDArray slot = vlad.getRow(i); if (slot.norm2() > 1e-8) { slot.divi(slot.norm2()); } } // 全局L2归一化并展平为一维向量 vlad = vlad.reshape(1, k * featDim); vlad.divi(vlad.norm2()); return vlad; }
注意:如果用卷积层的特征图(比如conv5_3),每个图像会有多个局部特征,VLAD的聚合效果会更明显,但用fc2的全局特征也能生成有效的VLAD表示。
问题4:欧氏距离对比效果不佳,是否需要预处理?
必须做预处理!fc2特征的原始数值范围差异很大,欧氏距离对数值尺度非常敏感,直接对比会导致效果很差。推荐做以下预处理:
1. L2归一化(最关键)
把每个特征向量归一化到单位长度,这样欧氏距离和余弦相似度等价(||a - b||² = 2(1 - cos(a,b))),而余弦相似度更适合衡量特征的相似性。在保存特征前添加归一化:
INDArray fc2Feature = stringINDArrayMap.get("fc2"); fc2Feature.divi(fc2Feature.norm2()); // 原地归一化,把向量缩放到单位长度 saveCompressed(file, fc2Feature);
2. PCA/白化降维
通过PCA降低特征维度,去除冗余信息和噪声,能进一步提升距离对比的效果,尤其是当样本量较大时。
3. 使用VLAD特征
VLAD特征本身经过了多轮归一化,其距离对比的效果通常比原始fc2特征更好,适合图像检索场景。
内容的提问来源于stack exchange,提问作者arocketman

