You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas按分组补全观测数到指定值且缺失值填充0的高效方法

Pandas分组补全缺失行的高效实现方法

你可以选择以下两种常用且高效的实现方案,均无需手动循环分组:

方案1:groupby + reindex(性能最优,适合大样本多分组场景)

核心逻辑是将观测序号x设为索引后,按分组索引i拆分,每个分组统一重索引到1~n的完整序列,缺失位直接填充0:

import pandas as pd

n = 5
test = pd.DataFrame({'i':[1,1,1,1,2,2],'x':[1,2,3,4,1,2],'v':[1,2,3,4,5,6]})

result = (
    test.set_index('x')
    .groupby('i', group_keys=False)
    .apply(lambda df: df.reindex(range(1, n+1), fill_value=0))
    .reset_index()
    [['i', 'x', 'v']]
)

输出结果完全匹配你给出的desired预期值。

方案2:全量组合左连接(逻辑直观,易理解)

先生成所有分组i和序号x的笛卡尔积全表,再左连接原表,缺失的观测值v统一填0即可:

n = 5
# 生成所有i和x的完整组合
full_index = pd.MultiIndex.from_product(
    [test['i'].unique(), range(1, n+1)], 
    names=['i', 'x']
).to_frame(index=False)

# 左连原表,空值补0
result = full_index.merge(test, on=['i', 'x'], how='left').fillna({'v':0}).astype({'v':int})

两种方案都能实现需求,可根据你的数据规模和习惯选择。

内容的提问来源于stack exchange,提问作者oli5679

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 19:39:04