如何基于第二个NumPy数组将第一个数组拆分为列表数组
NumPy按指定长度高效拆分不等长数组
问题背景
现有两个NumPy数组:一个存储任意值,另一个存储大于1的整数(且整数之和等于第一个数组的长度),需要根据第二个数组的长度值拆分第一个数组。示例如下:
import numpy as np values = np.array(["a", "b", "c", "d", "e", "f", "g", "h"]) lengths = np.array([1, 3, 2, 2]) print(len(values) == sum(lengths)) # 输出 True
期望得到的拆分结果:
output = np.array([["a"], ["b", "c", "d"], ["e", "f"], ["g", "h"]], dtype=object)
直接用Python循环处理超大规模数组(如数亿元素)时速度极慢,需要用原生NumPy操作实现高效拆分。
解决方案
NumPy本身没有直接生成不等长数组的原生函数,但可以通过计算拆分索引+批量切片的方式实现,完全规避遍历每个元素的低效操作。以下是两种高效实现方式:
方法1:手动计算起始索引(直观可控)
import numpy as np values = np.array(["a", "b", "c", "d", "e", "f", "g", "h"]) lengths = np.array([1, 3, 2, 2]) # 计算每个分组的起始索引 starts = np.concatenate([[0], np.cumsum(lengths[:-1])]) # 按索引切片提取分组,转为object类型数组 output = np.array([values[s:s+l] for s, l in zip(starts, lengths)], dtype=object) # 若需要转为列表形式(和示例完全一致) output = np.array([values[s:s+l].tolist() for s, l in zip(starts, lengths)], dtype=object) print(output)
方法2:使用np.split(代码更简洁)
np.split内部已经封装了拆分点的计算逻辑,本质和方法1一致,效率相同:
import numpy as np values = np.array(["a", "b", "c", "d", "e", "f", "g", "h"]) lengths = np.array([1, 3, 2, 2]) # 计算拆分点(累加除最后一个元素外的lengths值) split_points = np.cumsum(lengths[:-1]) # 拆分并转为object类型数组 output = np.array(np.split(values, split_points), dtype=object) # 转为列表形式 output = np.array([arr.tolist() for arr in output], dtype=object) print(output)
性能说明
- 核心操作
np.cumsum和np.split都是NumPy底层实现的向量操作,时间复杂度为O(n),远快于Python循环(底层为C实现,无Python解释器开销)。 - 列表推导仅遍历
lengths的长度(远小于数亿级的元素数量),不会产生性能瓶颈。 - 最终的
object类型数组是NumPy处理不等长分组的唯一可行方式,因为NumPy原生仅支持同构数组,不等长数据只能以对象形式存储子数组/列表。
内容的提问来源于stack exchange,提问作者Ted
相关产品推荐
相关产品推荐

