如何在Colab笔记本中实现FIt-SNE及优化现有t-SNE代码
问题背景
我正在处理Kaggle狗情绪数据集,先通过PCA降维后运行普通t-SNE,之后尝试了Multicore t-SNE并提升迭代次数,但可视化效果仍不理想。现需了解如何在Colab中实现FIt-SNE,同时寻求现有代码的优化建议。
数据集基础信息:
- labels shape: (4000,)
- images shape: (4000, 110592)
- images size: (192, 192, 3)
在Colab中实现FIt-SNE
FIt-SNE是t-SNE的优化版本,在高维数据降维的效率与效果上均有提升,以下是Colab中的具体实现步骤:
1. 安装依赖与编译FIt-SNE
Colab环境中需先配置依赖,再克隆仓库并编译核心代码:
# 安装系统依赖 !apt-get update && apt-get install -y libopenblas-dev !pip install cython # 克隆FIt-SNE仓库并编译 !git clone https://github.com/KlugerLab/FIt-SNE.git %cd FIt-SNE !g++ -std=c++11 -O3 src/sptree.cpp src/tsne.cpp src/nbodyfft.cpp -o bin/fast_tsne -pthread -lopenblas -ffast-math -march=native %cd ..
2. 调用FIt-SNE进行降维
编译完成后,导入Python接口并对PCA降维后的数据执行转换:
# 添加FIt-SNE路径到Python环境 import sys sys.path.append('/content/FIt-SNE') from fast_tsne import fast_tsne # 执行FIt-SNE,可根据数据调整参数 Z = fast_tsne(imgPCA, perplexity=30, theta=0.5, n_iter=7500) # 用原有可视化函数查看结果 plot_embedding(Z)
参数说明:
perplexity:建议范围5-50,可尝试30-40适配狗情绪数据集theta:取值0-1,越小精度越高但速度越慢,默认0.5平衡速度与效果n_iter:建议至少5000次迭代,可根据效果调整至10000
现有代码优化建议
1. 数据集加载优化
- 预分配内存:替换列表
append的方式,提前预分配numpy数组,减少内存拷贝开销:# 预分配数组存储图像 images = np.empty((4000, img_size[0], img_size[1], img_size[2]), dtype=np.uint8) labels = [] for idx, image_row in tqdm(enumerate(labels_df.iloc), desc="loading images", unit="images", total=4000): img_path = f"{directory}{image_row[2]}/{image_row[1]}" img = cv2.imread(img_path, cv2.IMREAD_COLOR)[:, :, ::-1] # BGR转RGB images[idx] = cv2.resize(img, img_size[0:2]) labels.append(image_row[2]) images = images.reshape(4000, num_px) labels = np.array(labels) - 异常处理:添加图片读取失败判断,避免损坏图片中断程序:
img = cv2.imread(img_path, cv2.IMREAD_COLOR) if img is None: print(f"加载失败:{img_path}") continue
2. PCA降维优化
- 标准化数据:图像像素值范围0-255,PCA前先做标准化(均值0、方差1),提升降维效果:
from sklearn.preprocessing import StandardScaler scaler = StandardScaler() images_scaled = scaler.fit_transform(images) pca = decomposition.PCA(n_components=60) imgPCA = pca.fit_transform(images_scaled) - 动态选择维度:通过累计方差贡献选择最优PCA维度,而非固定60:
pca = decomposition.PCA() pca.fit(images_scaled) cumulative_variance = np.cumsum(pca.explained_variance_ratio_) optimal_pc = np.argmax(cumulative_variance >= 0.95) + 1 # 取累计方差95%对应的维度 print(f"最优PCA维度:{optimal_pc}") pca = decomposition.PCA(n_components=optimal_pc) imgPCA = pca.fit_transform(images_scaled)
3. t-SNE参数优化
无论使用哪种t-SNE变体,以下参数可重点调整:
perplexity:核心参数,建议尝试10、30、50,适配数据集分布learning_rate:默认200,可尝试100-500,过低会导致点聚集,过高会导致点分散n_iter:普通t-SNE建议至少1000次,Multicore/FIt-SNE可尝试5000-10000次
4. 可视化优化
- 添加图例:补充标签与颜色的对应图例,提升可读性:
def plot_embedding(Z, show_axis=False): plt.figure(figsize=(10, 8)) unique_labels = np.unique(labels) label_map = {label: i for i, label in enumerate(unique_labels)} color = np.array([label_map[l] for l in labels]) # 按标签分组绘制散点并添加图例 for idx, label in enumerate(unique_labels): mask = labels == label plt.scatter(Z[mask, 0], Z[mask, 1], label=label, cmap="jet", s=10) plt.colorbar(ticks=range(len(unique_labels)), label="情绪标签") plt.legend(bbox_to_anchor=(1.05, 1), loc='upper left') plt.title('2D t-SNE 可视化结果') if not show_axis: plt.axis("off") plt.axis("equal") plt.tight_layout() plt.show() - 调整点大小:通过
s=10缩小散点,避免点重叠导致的特征被掩盖
原代码整理
数据集加载代码
import pandas as pd import cv2 import numpy as np from tqdm import tqdm import matplotlib.pyplot as plt img_size = (192,192,3) num_px = img_size[0] * img_size[1] * img_size[2] directory = '/content/drive/MyDrive/Colab Notebooks/ML/Dog_Emotion/' images = [] labels = [] labels_df = pd.read_csv(directory + "labels.csv") n_images = 0 for image in tqdm(labels_df.iloc, desc = "loading images", unit = "images", total = 4000): images.append(np.asarray(cv2.resize(cv2.imread(directory + image[2] + '/' + image[1], cv2.IMREAD_COLOR), img_size[0:2])[:, :, ::-1])) labels.append(image[2]) images, labels = np.array(images).reshape(4000, num_px), np.array(labels) print(f'labels shape: {labels.shape}') print(f'images shape: {images.shape}') print(f'images size: {img_size}') def plot_embedding(Z, show_axis="False"): plt.figure(figsize=(10, 8)) map = {label: i for i, label in enumerate(np.unique(labels))} color = np.array([map[l] for l in labels]) plt.scatter(Z[:, 0], Z[:, 1], c = color, cmap = "jet") plt.colorbar() plt.title('2d t-SNE Visualization') if not show_axis: plt.axis("off") plt.axis("equal") plt.show()
PCA+普通t-SNE代码
from sklearn import decomposition from sklearn.manifold import TSNE pc = 60 pca = decomposition.PCA(n_components=pc) _ = pca.fit(images) imgPCA = pca.transform(images) tsne = TSNE(n_components=2) Z = tsne.fit_transform(imgPCA) plot_embedding(Z)
Multicore t-SNE代码
!pip install git+https://github.com/DmitryUlyanov/Multicore-TSNE.git from MulticoreTSNE import MulticoreTSNE Z = MulticoreTSNE(n_jobs=4, n_iter=10000).fit_transform(imgPCA) plot_embedding(Z, show_axis = True)
内容的提问来源于stack exchange,提问作者Michele Magrini
相关产品推荐
相关产品推荐

