You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用TSNE绘制PyTorch张量向量报错:'list' object has no attribute 'shape'

问题描述

使用ESM-1b模型训练蛋白质序列后,得到PyTorch张量形式的向量列表,尝试用t-SNE可视化时出现错误:

'list' object has no attribute 'shape'

相关代码:
生成向量列表的代码:

sequence_representations = []
for i, (_, seq) in enumerate(new_list):
   sequence_representations.append(token_representations[i, 1 : len(seq) + 1].mean(0))

sequence_representations示例:

[tensor([-0.0054,  0.1090, -0.0046,  ...,  0.0465,  0.0426, -0.0675]),
 tensor([-0.0025,  0.0228, -0.0521,  ..., -0.0611,  0.1010, -0.0103]),
 tensor([ 0.1168, -0.0189, -0.0121,  ..., -0.0388,  0.0586, -0.0285]),......

报错的t-SNE代码:

X_embedded = TSNE(n_components=2, learning_rate='auto', init='random').fit_transform(sequence_representations) #报错位置
解决方法

sklearn的t-SNE要求输入是二维数组结构(形状为(样本数, 特征数)),而你传入的是PyTorch张量的列表,列表没有shape属性,因此报错。只需将张量列表转换为符合要求的numpy数组即可,步骤如下:

  1. 合并张量并转换为numpy数组
import torch
from sklearn.manifold import TSNE

# 将张量列表堆叠成二维张量
sequence_tensor = torch.stack(sequence_representations)
# 如果张量在GPU上,先转到CPU再转numpy;CPU张量可直接转
sequence_np = sequence_tensor.cpu().numpy()
  1. 运行t-SNE并可视化
# 执行t-SNE降维
X_embedded = TSNE(n_components=2, learning_rate='auto', init='random').fit_transform(sequence_np)

# 用matplotlib绘制可视化结果
import matplotlib.pyplot as plt

plt.scatter(X_embedded[:, 0], X_embedded[:, 1])
plt.title("蛋白质序列表示的t-SNE可视化")
plt.show()

关键说明

  • torch.stack()会把列表中所有形状相同的一维张量,堆叠成一个二维张量(每个样本对应一行),这是t-SNE需要的输入格式。
  • 如果你的张量存储在GPU上,必须调用.cpu()先转移到CPU,再调用.numpy()转换为numpy数组,否则会出现设备不匹配的错误。

内容的提问来源于stack exchange,提问作者eneko valero

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 10:45:32