You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Colab笔记本中实现FIt-SNE及优化现有t-SNE代码

问题背景

我正在处理Kaggle狗情绪数据集,先通过PCA降维后运行普通t-SNE,之后尝试了Multicore t-SNE并提升迭代次数,但可视化效果仍不理想。现需了解如何在Colab中实现FIt-SNE,同时寻求现有代码的优化建议。

数据集基础信息:

  • labels shape: (4000,)
  • images shape: (4000, 110592)
  • images size: (192, 192, 3)

在Colab中实现FIt-SNE

FIt-SNE是t-SNE的优化版本,在高维数据降维的效率与效果上均有提升,以下是Colab中的具体实现步骤:

1. 安装依赖与编译FIt-SNE

Colab环境中需先配置依赖,再克隆仓库并编译核心代码:

# 安装系统依赖
!apt-get update && apt-get install -y libopenblas-dev
!pip install cython

# 克隆FIt-SNE仓库并编译
!git clone https://github.com/KlugerLab/FIt-SNE.git
%cd FIt-SNE
!g++ -std=c++11 -O3 src/sptree.cpp src/tsne.cpp src/nbodyfft.cpp -o bin/fast_tsne -pthread -lopenblas -ffast-math -march=native
%cd ..

2. 调用FIt-SNE进行降维

编译完成后,导入Python接口并对PCA降维后的数据执行转换:

# 添加FIt-SNE路径到Python环境
import sys
sys.path.append('/content/FIt-SNE')

from fast_tsne import fast_tsne

# 执行FIt-SNE,可根据数据调整参数
Z = fast_tsne(imgPCA, perplexity=30, theta=0.5, n_iter=7500)

# 用原有可视化函数查看结果
plot_embedding(Z)

参数说明:

  • perplexity:建议范围5-50,可尝试30-40适配狗情绪数据集
  • theta:取值0-1,越小精度越高但速度越慢,默认0.5平衡速度与效果
  • n_iter:建议至少5000次迭代,可根据效果调整至10000

现有代码优化建议

1. 数据集加载优化

  • 预分配内存:替换列表append的方式,提前预分配numpy数组,减少内存拷贝开销:
    # 预分配数组存储图像
    images = np.empty((4000, img_size[0], img_size[1], img_size[2]), dtype=np.uint8)
    labels = []
    
    for idx, image_row in tqdm(enumerate(labels_df.iloc), desc="loading images", unit="images", total=4000):
        img_path = f"{directory}{image_row[2]}/{image_row[1]}"
        img = cv2.imread(img_path, cv2.IMREAD_COLOR)[:, :, ::-1]  # BGR转RGB
        images[idx] = cv2.resize(img, img_size[0:2])
        labels.append(image_row[2])
    
    images = images.reshape(4000, num_px)
    labels = np.array(labels)
    
  • 异常处理:添加图片读取失败判断,避免损坏图片中断程序:
    img = cv2.imread(img_path, cv2.IMREAD_COLOR)
    if img is None:
        print(f"加载失败:{img_path}")
        continue
    

2. PCA降维优化

  • 标准化数据:图像像素值范围0-255,PCA前先做标准化(均值0、方差1),提升降维效果:
    from sklearn.preprocessing import StandardScaler
    
    scaler = StandardScaler()
    images_scaled = scaler.fit_transform(images)
    
    pca = decomposition.PCA(n_components=60)
    imgPCA = pca.fit_transform(images_scaled)
    
  • 动态选择维度:通过累计方差贡献选择最优PCA维度,而非固定60:
    pca = decomposition.PCA()
    pca.fit(images_scaled)
    cumulative_variance = np.cumsum(pca.explained_variance_ratio_)
    optimal_pc = np.argmax(cumulative_variance >= 0.95) + 1  # 取累计方差95%对应的维度
    print(f"最优PCA维度:{optimal_pc}")
    
    pca = decomposition.PCA(n_components=optimal_pc)
    imgPCA = pca.fit_transform(images_scaled)
    

3. t-SNE参数优化

无论使用哪种t-SNE变体,以下参数可重点调整:

  • perplexity:核心参数,建议尝试10、30、50,适配数据集分布
  • learning_rate:默认200,可尝试100-500,过低会导致点聚集,过高会导致点分散
  • n_iter:普通t-SNE建议至少1000次,Multicore/FIt-SNE可尝试5000-10000次

4. 可视化优化

  • 添加图例:补充标签与颜色的对应图例,提升可读性:
    def plot_embedding(Z, show_axis=False):
        plt.figure(figsize=(10, 8))
        unique_labels = np.unique(labels)
        label_map = {label: i for i, label in enumerate(unique_labels)}
        color = np.array([label_map[l] for l in labels])
        
        # 按标签分组绘制散点并添加图例
        for idx, label in enumerate(unique_labels):
            mask = labels == label
            plt.scatter(Z[mask, 0], Z[mask, 1], label=label, cmap="jet", s=10)
        
        plt.colorbar(ticks=range(len(unique_labels)), label="情绪标签")
        plt.legend(bbox_to_anchor=(1.05, 1), loc='upper left')
        plt.title('2D t-SNE 可视化结果')
        if not show_axis:
            plt.axis("off")
        plt.axis("equal")
        plt.tight_layout()
        plt.show()
    
  • 调整点大小:通过s=10缩小散点,避免点重叠导致的特征被掩盖

原代码整理

数据集加载代码

import pandas as pd
import cv2
import numpy as np
from tqdm import tqdm
import matplotlib.pyplot as plt

img_size = (192,192,3)
num_px = img_size[0] * img_size[1] * img_size[2]

directory = '/content/drive/MyDrive/Colab Notebooks/ML/Dog_Emotion/' 
images = []
labels = []
labels_df = pd.read_csv(directory + "labels.csv")
n_images = 0

for image in tqdm(labels_df.iloc, desc = "loading images", unit = "images", total = 4000):
  images.append(np.asarray(cv2.resize(cv2.imread(directory + image[2] + '/' + image[1], cv2.IMREAD_COLOR), img_size[0:2])[:, :, ::-1]))
  labels.append(image[2])

images, labels = np.array(images).reshape(4000, num_px), np.array(labels)

print(f'labels shape: {labels.shape}')
print(f'images shape: {images.shape}')
print(f'images size: {img_size}')

def plot_embedding(Z, show_axis="False"):
  plt.figure(figsize=(10, 8))
  map = {label: i for i, label in enumerate(np.unique(labels))}
  color = np.array([map[l] for l in labels])
  plt.scatter(Z[:, 0], Z[:, 1], c = color, cmap = "jet")
  plt.colorbar()
  plt.title('2d t-SNE Visualization')
  if not show_axis:
    plt.axis("off")
  plt.axis("equal")
  plt.show()

PCA+普通t-SNE代码

from sklearn import decomposition
from sklearn.manifold import TSNE

pc = 60
pca = decomposition.PCA(n_components=pc)
_ = pca.fit(images)
imgPCA = pca.transform(images)

tsne = TSNE(n_components=2)
Z = tsne.fit_transform(imgPCA)
plot_embedding(Z)

Multicore t-SNE代码

!pip install git+https://github.com/DmitryUlyanov/Multicore-TSNE.git

from MulticoreTSNE import MulticoreTSNE

Z = MulticoreTSNE(n_jobs=4, n_iter=10000).fit_transform(imgPCA)
plot_embedding(Z, show_axis = True)

内容的提问来源于stack exchange,提问作者Michele Magrini

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 04:45:11