Jupyter Notebook中删除0均值行及对应同客户日期行的数据清洗问题
数据清洗解决方案
你原来的代码有几个明显问题:
- 循环遍历的是
peak mean列的数值,不是行索引,所以用i去取Customer列会报错,因为i是0或者其他数值,不是行的位置标识。 drop方法默认不会修改原DataFrame,要么加inplace=True,要么把结果重新赋值给变量,不然删了等于白删。- 嵌套循环的逻辑既不符合需求,效率还极低,pandas处理数据尽量别用循环遍历行。
下面给你两种符合需求的高效实现方式,根据你的实际场景选择:
场景1:删除所有和0值行同Customer+Date的行
如果只要某组(同一Customer+Date)里有一行peak mean为0,就把这组所有行都删掉,用这个方法:
# 第一步:找出所有peak mean为0的行对应的Customer和Date组合 bad_pairs = both_index[both_index['peak mean'] == 0][['Customer', 'Date']] # 第二步:筛选出所有属于这些组合的行,删除它们 both_index = both_index[~both_index.apply( lambda row: (row['Customer'], row['Date']) in bad_pairs.itertuples(index=False), axis=1 )]
或者更简洁的合并写法:
both_index = both_index[~both_index[['Customer', 'Date']].isin( both_index[both_index['peak mean'] == 0][['Customer', 'Date']] ).all(axis=1)]
场景2:仅删除0值行,以及它上下相邻且同Customer+Date的行
如果只删0值行本身,以及它紧挨着的上下行(前提是上下行的Customer和Date和0值行一致),用这个方法:
# 先找出所有peak mean为0的行的索引 zero_rows = both_index[both_index['peak mean'] == 0].index # 收集要删除的索引集合 to_remove = set() for idx in zero_rows: # 先把当前0值行加进去 to_remove.add(idx) # 检查前一行是否存在,且Customer和Date匹配 if idx > both_index.index[0]: prev_idx = idx - 1 if (both_index.loc[prev_idx, 'Customer'] == both_index.loc[idx, 'Customer'] and both_index.loc[prev_idx, 'Date'] == both_index.loc[idx, 'Date']): to_remove.add(prev_idx) # 检查后一行是否存在,且Customer和Date匹配 if idx < both_index.index[-1]: next_idx = idx + 1 if (both_index.loc[next_idx, 'Customer'] == both_index.loc[idx, 'Customer'] and both_index.loc[next_idx, 'Date'] == both_index.loc[idx, 'Date']): to_remove.add(next_idx) # 执行删除 both_index = both_index.drop(to_remove)
补充说明
如果你的DataFrame用的是自定义索引(不是默认整数索引),需要调整前后行的获取方式:
pos = both_index.index.get_loc(idx) if pos > 0: prev_idx = both_index.index[pos-1] # 后续匹配逻辑同上
内容的提问来源于stack exchange,提问作者astronaut19
相关产品推荐
相关产品推荐

