如何将含numpy一维数组列的DataFrame按数组元素拆分为多行?
高效展开含numpy数组列的大DataFrame(个人级→个人-点数级)
针对你的需求——将含numpy一维数组列的个人级DataFrame,展开为个人-点数级的新DataFrame,以下是两种高效实现方案,分别适配不同数据量级:
方案1:Pandas explode方法(简洁易读,适用于常规数据量)
该方法代码直观,维护成本低,适合大多数非极端数据规模的场景:
示例代码
import pandas as pd import numpy as np # 构造你的PERSON DataFrame(示例数据) person_df = pd.DataFrame({ 'PERSON_ID': [1, 2, 3, 4], 'PERSON_NAME': ['A', 'B', 'C', 'D'], 'PERSON_POINTS': [np.array([1,2,3]), np.array([4,5,6,7]), np.array([8]), np.array([9,10])], 'PERSON_DISTANCES': [np.array([2,4,6]), np.array([2,4,6,8]), np.array([6]), np.array([4,8])] }) # 将numpy数组转为列表(适配部分低版本Pandas对numpy数组的explode支持) person_df['PERSON_POINTS'] = person_df['PERSON_POINTS'].tolist() person_df['PERSON_DISTANCES'] = person_df['PERSON_DISTANCES'].tolist() # 同时展开两个数组列,重置索引 person_points_df = person_df.explode( ['PERSON_POINTS', 'PERSON_DISTANCES'], ignore_index=True ) # 重命名列并修正数据类型 person_points_df = person_points_df.rename(columns={ 'PERSON_POINTS': 'PERSON_POINT', 'PERSON_DISTANCES': 'PERSON_DISTANCE' }).astype({ 'PERSON_POINT': int, 'PERSON_DISTANCE': int }) # 输出结果 print(person_points_df)
方案2:Numpy底层操作(性能最优,适用于超大数据量)
当你的DataFrame行数达到百万级甚至更高时,直接使用Numpy的数组操作能大幅提升效率,避免Pandas上层逻辑的额外开销:
示例代码
import pandas as pd import numpy as np # 构造你的PERSON DataFrame(示例数据) person_df = pd.DataFrame({ 'PERSON_ID': [1, 2, 3, 4], 'PERSON_NAME': ['A', 'B', 'C', 'D'], 'PERSON_POINTS': [np.array([1,2,3]), np.array([4,5,6,7]), np.array([8]), np.array([9,10])], 'PERSON_DISTANCES': [np.array([2,4,6]), np.array([2,4,6,8]), np.array([6]), np.array([4,8])] }) # 获取每个numpy数组的长度,用于重复ID和名称 array_lengths = np.array([len(arr) for arr in person_df['PERSON_POINTS']]) # 重复个人标识列 repeated_ids = np.repeat(person_df['PERSON_ID'].values, array_lengths) repeated_names = np.repeat(person_df['PERSON_NAME'].values, array_lengths) # 扁平化数组列 flattened_points = np.concatenate(person_df['PERSON_POINTS'].values) flattened_distances = np.concatenate(person_df['PERSON_DISTANCES'].values) # 构造目标DataFrame person_points_df = pd.DataFrame({ 'PERSON_ID': repeated_ids, 'PERSON_NAME': repeated_names, 'PERSON_POINT': flattened_points, 'PERSON_DISTANCE': flattened_distances }) # 输出结果 print(person_points_df)
结果验证
两种方案都会生成符合你要求的Person_Points DataFrame,示例输出如下:
PERSON_ID PERSON_NAME PERSON_POINT PERSON_DISTANCE 0 1 A 1 2 1 1 A 2 4 2 1 A 3 6 3 2 B 4 2 4 2 B 5 4 5 2 B 6 6 6 2 B 7 8 7 3 C 8 6 8 4 D 9 4 9 4 D 10 8
内容的提问来源于stack exchange,提问作者user19725009
相关产品推荐
相关产品推荐

