如何用Python将VGG生成的4096维向量降维至2000以下适配Postgres/pgvector
4096维图像向量降维至2000以下的实现方案
1. 主成分分析(PCA)
这是最常用的线性降维方法,能在保留最大方差的前提下降低维度,计算效率高,适配大规模数据,非常适合图像相似度检索场景。
代码示例:
import numpy as np from sklearn.decomposition import PCA # 假设vgg_features是批量VGG生成的4096维向量,shape为(n_samples, 4096) vgg_features = np.random.rand(1000, 4096) # 示例数据 # 初始化PCA,指定目标维度(如1024,需小于2000) pca = PCA(n_components=1024) # 拟合并完成降维 reduced_features = pca.fit_transform(vgg_features) print(reduced_features.shape) # 输出 (1000, 1024)
若需保留指定比例的方差(如95%),可自动计算维度:
pca = PCA(n_components=0.95) reduced_features = pca.fit_transform(vgg_features) print(f"自动选定的维度: {pca.n_components_}")
2. 截断奇异值分解(Truncated SVD)
原理与PCA类似,但无需对数据做中心化处理,计算速度更快,适配稠密或稀疏特征场景。
代码示例:
from sklearn.decomposition import TruncatedSVD svd = TruncatedSVD(n_components=1536) # 选定1536维,小于2000 reduced_features = svd.fit_transform(vgg_features) print(reduced_features.shape) # 输出 (1000, 1536)
3. 直接提取VGG中间层特征(无需额外降维)
VGG16/VGG19的全连接层输出为4096维,但可直接提取更早的中间层特征,经池化后得到低维向量,同时保留更多空间特征信息。
代码示例(基于Keras):
from tensorflow.keras.applications.vgg16 import VGG16, preprocess_input from tensorflow.keras.preprocessing import image from tensorflow.keras.models import Model import numpy as np # 加载不含顶层分类器的VGG16 base_model = VGG16(weights='imagenet', include_top=False) # 构建模型,输出block5_pool层特征 model = Model(inputs=base_model.input, outputs=base_model.get_layer('block5_pool').output) # 处理单张示例图片 img_path = 'test.jpg' img = image.load_img(img_path, target_size=(224, 224)) x = image.img_to_array(img) x = np.expand_dims(x, axis=0) x = preprocess_input(x) # 获取特征并做全局平均池化,将7x7x512转为512维 features = model.predict(x) reduced_features = np.mean(features, axis=(1, 2)) print(reduced_features.shape) # 输出 (1, 512)
4. 神经网络特征压缩(适配有标注数据场景)
若有图像分类/检索的标注数据,可训练小型神经网络将4096维特征映射至目标维度,能保留更多语义信息,提升检索效果。
代码示例:
import tensorflow as tf from tensorflow.keras.layers import Input, Dense from tensorflow.keras.models import Model # 定义压缩模型结构 input_layer = Input(shape=(4096,)) hidden = Dense(2048, activation='relu')(input_layer) output_layer = Dense(1024, activation=None)(hidden) compression_model = Model(inputs=input_layer, outputs=output_layer) # 采用自监督训练,以原特征为标签 compression_model.compile(optimizer='adam', loss='mse') # 假设X_train为训练用的4096维特征集 compression_model.fit(X_train, X_train, epochs=50, batch_size=32) # 完成降维 reduced_features = compression_model.predict(vgg_features) print(reduced_features.shape) # 输出 (n_samples, 1024)
内容的提问来源于stack exchange,提问作者Pio92
相关产品推荐
相关产品推荐

