You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:从多级索引DataFrame生成含行索引的坐标数组

问题描述

我有一个存储不同身体部位随时间变化的x、y坐标的多级索引DataFrame,结构如下:

segment         0                         1  ...      98        99                
coords          k       x       y         k  ...       y         k       x       y
0        0.008525  312.05  361.65  0.011500  ...  329.97  0.012414  621.83  327.77
1        0.004090  312.32  359.98  0.007290  ...  329.00  0.034572  623.31  327.13
2        0.006645  313.42  359.11  0.011194  ...  330.53  0.003275  621.18  327.55
3        0.008367  314.79  361.47  0.013591  ...  329.58  0.026624  624.32  327.76
4        0.005160  315.91  364.54  0.009056  ...  329.97  0.026840  624.54  327.97
...           ...     ...     ...       ...  ...     ...       ...     ...     ...
40006   -0.081192  323.60  354.73 -0.070411  ...  431.78  0.088513  432.43  433.49
40007   -0.050125  319.29  357.99 -0.074568  ...  431.00  0.470994  436.47  432.65

该DataFrame形状为40008行×300列,其中k值无需保留。为满足绘图需求,需要将数据转换为每行包含[行索引, x坐标, y坐标]的数组,最终维度应为(4000800, 3),示例格式如下:

[[index0, x_i0_s0, y_i0_s0],
[index0, x_i0_s1,y_i0_s1],
[index0, x_i0_s2,y_i0_s2],
...
[index40007, x_i40007_s97, y_i40007_s97],
[index40007, x_i40007_s98,y_i40007_s98],
[index40007, x_i40007_s99,y_i40007_s99]]

实际数据示例:

[[0, 312.05, 361.65],
...
[40007, 436.47, 432.65]]

尝试过以下代码提取x、y列:

x = df.xs(('x',), level=('coords',), axis=1)
y = df.xs(('y',), level=('coords',), axis=1)

之后用np.stack:

x = df.xs(('x',), level=('coords',), axis=1)
y = df.xs(('y',), level=('coords',), axis=1)
result = np.stack((x,y)), axis=2)  # 存在语法错误

得到形状为(40008, 100, 2)的数组,未达到预期效果,求解决方案。

解决方案

方法一:基于NumPy的数组操作

import numpy as np

# 提取x、y的数值数组
x = df.xs('x', level='coords', axis=1).values
y = df.xs('y', level='coords', axis=1).values

# 堆叠x和y,得到(40008, 100, 2)的数组
coords = np.stack([x, y], axis=2)

# 展平坐标数组为(4000800, 2)
flattened_coords = coords.reshape(-1, 2)

# 生成重复的行索引(每个原始行对应100个部位数据)
indices = np.repeat(df.index.values, x.shape[1])

# 拼接索引与坐标,得到目标格式数组
result = np.column_stack([indices, flattened_coords])

方法二:基于Pandas的重塑操作(更简洁)

# 移除所有k列
df_filtered = df.drop('k', level='coords', axis=1)

# 将segment列转为行,同时保留原始行索引
melted = df_filtered.stack(level='segment').reset_index()

# 提取需要的列并转为numpy数组
result = melted[['level_0', 'x', 'y']].values

两种方法最终都会得到维度为(4000800, 3)的数组,完全符合需求。

内容的提问来源于stack exchange,提问作者Ulises Rey

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 14:01:17