You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何向量化统一NumPy列表数组长度并正确推断dtype?

向量化方式填充DataFrame中的不等长列表

问题背景

现有如下DataFrame:

import pandas as pd
import numpy as np

data = {'col_a': [['a', 'b'], ['a', 'b', 'c'], ['a'], ['a', 'b', 'c', 'd'], ['a', 'b', 'c'], ['a', 'b', 'c', 'd']],
        'col_b':[[1, 3], [1, 0, 0], [4], [1, 1, 2, 0], [0, 0, 5], [3, 1, 2, 5]]}
df= pd.DataFrame(data)

需求:以向量化方式调整col_a和col_b中的子列表,使所有子列表长度等于最长子列表的长度。其中col_a空缺处填充'None',col_b空缺处填充nan,最终输出目标格式如下:

col_a               col_b
0     [a, b, None, None]    [1, 3, nan, nan]
1        [a, b, c, None]      [1, 0, 0, nan]
2  [a, None, None, None]  [4, nan, nan, nan]
3           [a, b, c, d]        [1, 1, 2, 0]
4        [a, b, c, None]      [0, 0, 5, nan]
5           [a, b, c, d]        [3, 1, 2, 5]

尝试代码及错误

尝试了以下向量化代码:

# Convert the column to a NumPy array with object dtype
col_np = df['col_a'].to_numpy()

# Find the maximum length of the lists using NumPy operations
max_length = np.max(np.frompyfunc(len, 1, 1)(col_np))

# Create a mask for padding
mask = np.arange(max_length) < np.frompyfunc(len, 1, 1)(col_np)[:, None]

# Pad the lists with None where necessary
result = np.where(mask, col_np, 'None')

出现错误:

ValueError: operands could not be broadcast together with shapes (6,4) (6,) ()

错误原因

错误出在np.where(mask, col_np, 'None')这一行:col_np是形状为(6,)的一维数组(每个元素是列表),而mask是(6,4)的二维数组,两者无法直接广播匹配,导致维度不兼容。

正确向量化解决方案

要实现向量化填充,需要先将不等长列表转换为二维数组,再进行填充,最后转回列表格式:

处理col_a

# 获取所有子列表长度,转为numpy数组
lengths = np.frompyfunc(len, 1, 1)(df['col_a'].to_numpy()).astype(int)
max_len = lengths.max()

# 创建二维数组,初始填充'None'
col_a_padded = np.full((len(df), max_len), 'None', dtype=object)

# 生成索引矩阵,定位每个子列表的有效元素
row_idx = np.repeat(np.arange(len(df)), lengths)
col_idx = np.concatenate([np.arange(l) for l in lengths])

# 将原列表的有效元素填入对应位置
col_a_padded[row_idx, col_idx] = np.concatenate(df['col_a'].to_numpy())

# 转换回列表格式
df['col_a'] = col_a_padded.tolist()

处理col_b

col_b需要填充nan,处理逻辑类似,只是填充值和 dtype 不同:

lengths_b = np.frompyfunc(len, 1, 1)(df['col_b'].to_numpy()).astype(int)

# 创建二维数组,初始填充nan
col_b_padded = np.full((len(df), max_len), np.nan)

# 填入有效元素
row_idx_b = np.repeat(np.arange(len(df)), lengths_b)
col_idx_b = np.concatenate([np.arange(l) for l in lengths_b])
col_b_padded[row_idx_b, col_idx_b] = np.concatenate(df['col_b'].to_numpy())

# 转换回列表格式
df['col_b'] = col_b_padded.tolist()

完整代码

import pandas as pd
import numpy as np

data = {'col_a': [['a', 'b'], ['a', 'b', 'c'], ['a'], ['a', 'b', 'c', 'd'], ['a', 'b', 'c'], ['a', 'b', 'c', 'd']],
        'col_b':[[1, 3], [1, 0, 0], [4], [1, 1, 2, 0], [0, 0, 5], [3, 1, 2, 5]]}
df= pd.DataFrame(data)

# 计算最长子列表长度
lengths_a = np.frompyfunc(len, 1, 1)(df['col_a'].to_numpy()).astype(int)
max_len = lengths_a.max()

# 处理col_a
col_a_padded = np.full((len(df), max_len), 'None', dtype=object)
row_idx_a = np.repeat(np.arange(len(df)), lengths_a)
col_idx_a = np.concatenate([np.arange(l) for l in lengths_a])
col_a_padded[row_idx_a, col_idx_a] = np.concatenate(df['col_a'].to_numpy())
df['col_a'] = col_a_padded.tolist()

# 处理col_b
lengths_b = np.frompyfunc(len, 1, 1)(df['col_b'].to_numpy()).astype(int)
col_b_padded = np.full((len(df), max_len), np.nan)
row_idx_b = np.repeat(np.arange(len(df)), lengths_b)
col_idx_b = np.concatenate([np.arange(l) for l in lengths_b])
col_b_padded[row_idx_b, col_idx_b] = np.concatenate(df['col_b'].to_numpy())
df['col_b'] = col_b_padded.tolist()

print(df)

运行后即可得到目标输出。

内容的提问来源于stack exchange,提问作者Lihka_nonem

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 16:17:07