You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在pandas DataFrame中补全托盘缺失点位的对应数据行

缺失点位补全实现方案

直接用pandas向量化操作实现,比嵌套循环效率更高,代码如下:

第一步:构造示例数据(可跳过,直接用你自己的df即可)

import pandas as pd

data = [
    [16229767, 5, 2, 1, 'T1', 123],
    [16229767, 5, 1, 0, 'T1', 123],
    [16229767, 5, 3, 0, 'T1', 123],
    [16229767, 5, 4, 0, 'T1', 123],
    [16229767, 3, 3, 1, 'T9', 38],
    [16229767, 3, 1, 0, 'T9', 38],
    [16229767, 3, 4, 0, 'T9', 38],
    [29767162, 7, 1, 0, 'T4', 991],
    [29767162, 7, 4, 1, 'T4', 991]
]
df = pd.DataFrame(data, columns=['timestamp', 't_idx', 'position', 'error', 'type', 'SNR'])

第二步:核心补全逻辑

# 1. 提取每个「时间戳+托盘ID」组合的公共固定字段,同组合下type、SNR取值一致
group_fixed = df.groupby(['timestamp', 't_idx'], as_index=False)[['type', 'SNR']].first()

# 2. 为每个组合扩展出1-4的全部点位,生成理论完整表
full_pos = group_fixed.assign(position = [list(range(1,5))]*len(group_fixed)).explode('position', ignore_index=True)

# 3. 关联原表的error字段,缺失点位的error统一填充为1
df_full = full_pos.merge(
    df[['timestamp', 't_idx', 'position', 'error']], 
    on=['timestamp', 't_idx', 'position'], 
    how='left'
)
df_full['error'] = df_full['error'].fillna(1).astype(int)

# 4. 调整列顺序与原表一致,重置索引
df_full = df_full[['timestamp', 't_idx', 'position', 'error', 'type', 'SNR']].reset_index(drop=True)

说明

  • 输出的df_full就是补全后的结果,和你预期的补全行完全一致
  • 如果同个「时间戳+托盘ID」组合下type、SNR存在多个取值,可将first()替换为mode().iloc[0]取众数,或根据你的业务规则调整取值逻辑
  • 相比嵌套循环遍历的写法,向量化操作在数据量较大时性能提升非常明显

内容的提问来源于stack exchange,提问作者Drimer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 17:27:02