能否导出UMAP嵌入的坐标点并转换为指定的字典结构化格式?
解决方案
你需要的二维坐标直接存储在训练后得到的embedding对象的embedding_属性中,该属性是形状为(样本数量, 2)的numpy数组,每个样本的顺序和你输入的data、df['name']的顺序完全对应。
生成你要求的字典列表结构
运行以下代码即可得到你需要的格式:
# 取出UMAP降维后的坐标数组 coords = embedding.embedding_ # 生成指定结构的列表 result = [ {"name": name, "x-value": float(x), "y-value": float(y)} for name, (x, y) in zip(df['name'], coords) ]
如果后续需要更通用的结构化存储格式,也可以直接导出为CSV文件,方便其他可视化工具读取:
import pandas as pd export_df = pd.DataFrame({ "name": df["name"], "x-value": coords[:, 0], "y-value": coords[:, 1] }) export_df.to_csv("umap_embedding_coords.csv", index=False)
内容的提问来源于stack exchange,提问作者fdgnhgdh
相关产品推荐
相关产品推荐

