You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效按Timestamp分组导出Pandas DataFrame的坐标列表?

高效实现按Timestamp分组导出嵌套坐标列表

给定如下Pandas DataFrame:

timestamp   x_1  y_1   x_2  y_2
0  1649010310     1    4     7    7
1  1649010310     7    8     0    1
2  1649010310     5    1     0    0
3  1649010311     2    0     1    6
4  1649010311     8    4     2    7

需求是按timestamp字段分组,导出包含x_1、y_1、x_2、y_2坐标的嵌套列表,目标格式如下:

[[[1, 4, 7, 7],
  [7, 8, 0, 1],
  [5, 1, 0, 0]],
 [[2, 0, 1, 6],
  [8, 4, 2, 7]]]

已知可以通过遍历DataFrame的for循环实现,但希望找到Pandas或Numpy中更高效的实现方式。


Pandas 实现方案

直接用groupby配合apply和values.tolist()就能搞定,不用手动写循环,代码如下:

import pandas as pd

# 构造示例数据
df = pd.DataFrame({
    'timestamp': [1649010310, 1649010310, 1649010310, 1649010311, 1649010311],
    'x_1': [1,7,5,2,8],
    'y_1': [4,8,1,0,4],
    'x_2': [7,0,0,1,2],
    'y_2': [7,1,0,6,7]
})

# 分组转换为目标嵌套列表
result = df.groupby('timestamp')[['x_1', 'y_1', 'x_2', 'y_2']].apply(lambda g: g.values.tolist()).tolist()
print(result)

先按timestamp分组,对每组的指定列提取数值并转为列表,最后把分组结果整体转成列表,完全符合要求,效率比手动循环高得多。

Numpy 高性能方案

如果数据量特别大,用Numpy的分割操作性能会更好,代码如下:

import numpy as np
import pandas as pd

# 构造示例数据
df = pd.DataFrame({
    'timestamp': [1649010310, 1649010310, 1649010310, 1649010311, 1649010311],
    'x_1': [1,7,5,2,8],
    'y_1': [4,8,1,0,4],
    'x_2': [7,0,0,1,2],
    'y_2': [7,1,0,6,7]
})

# 获取分组的索引分割点
_, split_indices = np.unique(df['timestamp'], return_index=True)
# 分割坐标数组并转为嵌套列表
coords = df[['x_1', 'y_1', 'x_2', 'y_2']].values
result = [group.tolist() for group in np.split(coords, split_indices[1:])]
print(result)

通过np.unique拿到分组的起始索引,再用np.split切割坐标数组,最后转成列表,这种方式避开了Pandas分组的额外开销,大数据场景下优势明显。


内容的提问来源于stack exchange,提问作者teejk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 00:32:51