删除重复行后DataFrame.info索引未更新问题咨询
问题
我有一个包含2111行的数据集,删除27条重复行后,DataFrame.info()的输出仍显示行索引范围为0至2110,但实际报告行数为2085。尝试过DataFrame.drop_duplicates()的inplace=True和inplace=False参数,结果一致。是否存在需要调用的DataFrame元数据刷新命令?
预处理前DataFrame.info()输出
!!!!!!!!!!!!!!!!!! Size and shape and info before preprocess 40109 (2111, 19) <bound method DataFrame.info of id Gender Age Height Weight ... TUE CALC MTRANS NObeyesdad BMI 0 1 female 21 1.6200 64 ... 3 to 5 no public_transportation normal_weight 24.3865 1 2 female 21 1.5200 56 ... 0 to 2 sometimes public_transportation normal_weight 24.2382 2 3 male 23 1.8000 77 ... 3 to 5 frequently public_transportation normal_weight 23.7654 3 4 male 27 1.8000 87 ... 0 to 2 frequently walking overweight_level_i 26.8519 4 5 male 22 1.7800 90 ... 0 to 2 sometimes public_transportation overweight_level_ii 28.3424 ... ... ... ... ... ... ... ... ... ... ... ... 2106 2,107 female 21 1.7107 131 ... 3 to 5 sometimes public_transportation obesity_type_iii 44.9015 2107 2,108 female 22 1.7486 134 ... 3 to 5 sometimes public_transportation obesity_type_iii 43.7419 2108 2,109 female 23 1.7522 134 ... 3 to 5 sometimes public_transportation obesity_type_iii 43.5438 2109 2,110 female 24 1.7394 133 ... 3 to 5 sometimes public_transportation obesity_type_iii 44.0715 2110 2,111 female 24 1.7388 133 ... 3 to 5 sometimes public_transportation obesity_type_iii 44.1443 [2111 rows x 19 columns]
删除重复行后DataFrame.info()输出
!!!!!!!!!!!!!!!!!! Size and shape After preprocess 37512 (2084, 18) <bound method DataFrame.info of Gender Age Height Weight FHWO FAVC FCVC NCP CAEC SMOKE CH2O SCC FAF TUE CALC MTRANS NObeyesdad BMI 0 2 21 1.6200 64 2 1 2 3 2 1 2 1 1 2 1 3 2 24.3865 1 2 21 1.5200 56 2 1 3 3 2 2 3 2 4 1 2 3 2 24.2382 2 1 23 1.8000 77 2 1 2 3 2 1 2 1 3 2 3 3 2 23.7654 3 1 27 1.8000 87 1 1 3 3 2 1 2 1 3 1 3 5 3 26.8519 4 1 22 1.7800 90 1 1 2 1 2 1 2 1 1 1 2 3 4 28.3424 ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... 2106 2 21 1.7107 131 2 2 3 3 2 1 2 1 3 2 2 3 7 44.9015 2107 2 22 1.7486 134 2 2 3 3 2 1 2 1 2 2 2 3 7 43.7419 2108 2 23 1.7522 134 2 2 3 3 2 1 2 1 2 2 2 3 7 43.5438 2109 2 24 1.7394 133 2 2 3 3 2 1 3 1 2 2 2 3 7 44.0715 2110 2 24 1.7388 133 2 2 3 3 2 1 3 1 2 2 2 3 7 44.1443 [2084 rows x 18 columns]
注:info输出显示最后一行索引为2110,但实际行数为2084。
尝试过的代码
# Example of inplace = False and inplace=True return_df = return_df.drop_duplicates(inplace=False) return_df.drop_duplicates(inplace=True)
相关删除重复行代码
# removing duplicates count_dup = return_df.duplicated().sum() if (verbose > 0): print (f"Number of Duplicates : {count_dup}") if count_dup > 0: if (verbose > 0): print ("Dropping Duplicates") # return_df.drop_duplicates(inplace=True) return_df = return_df.drop_duplicates(inplace=False) else: if (verbose > 0): print ("No duplicates found.") return return_df
解决方案
这不是元数据未刷新的问题,drop_duplicates()仅删除重复行,但不会重置DataFrame的索引。原索引会被保留,只是被删除行对应的索引会消失,因此最后一行的索引仍为原最大值2110,但实际有效行数已变为2084。
要让索引与实际行数匹配,只需调用reset_index()方法:
方法1:重置索引并保留原索引为新列
return_df = return_df.drop_duplicates().reset_index()
方法2:重置索引并丢弃原索引(推荐)
若无需保留原索引,添加drop=True参数:
return_df = return_df.drop_duplicates().reset_index(drop=True)
整合到现有代码
修改删除重复行的逻辑:
# removing duplicates count_dup = return_df.duplicated().sum() if (verbose > 0): print (f"Number of Duplicates : {count_dup}") if count_dup > 0: if (verbose > 0): print ("Dropping Duplicates") # 删除重复行并重置索引 return_df = return_df.drop_duplicates().reset_index(drop=True) else: if (verbose > 0): print ("No duplicates found.") return return_df
处理后,DataFrame.info()输出的索引范围会变为0到2083,与实际行数完全匹配。
内容的提问来源于stack exchange,提问作者Jess Stuart
相关产品推荐
相关产品推荐

