You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

删除重复行后DataFrame.info索引未更新问题咨询

问题

我有一个包含2111行的数据集,删除27条重复行后,DataFrame.info()的输出仍显示行索引范围为0至2110,但实际报告行数为2085。尝试过DataFrame.drop_duplicates()的inplace=True和inplace=False参数,结果一致。是否存在需要调用的DataFrame元数据刷新命令?

预处理前DataFrame.info()输出

!!!!!!!!!!!!!!!!!!     Size and shape and info before preprocess
40109
(2111, 19)
<bound method DataFrame.info of          id  Gender  Age  Height  Weight  ...     TUE        CALC                 MTRANS           NObeyesdad      BMI
0         1  female   21  1.6200      64  ...  3 to 5          no  public_transportation        normal_weight  24.3865
1         2  female   21  1.5200      56  ...  0 to 2   sometimes  public_transportation        normal_weight  24.2382
2         3    male   23  1.8000      77  ...  3 to 5  frequently  public_transportation        normal_weight  23.7654
3         4    male   27  1.8000      87  ...  0 to 2  frequently                walking   overweight_level_i  26.8519
4         5    male   22  1.7800      90  ...  0 to 2   sometimes  public_transportation  overweight_level_ii  28.3424
...     ...     ...  ...     ...     ...  ...     ...         ...                    ...                  ...      ...
2106  2,107  female   21  1.7107     131  ...  3 to 5   sometimes  public_transportation     obesity_type_iii  44.9015
2107  2,108  female   22  1.7486     134  ...  3 to 5   sometimes  public_transportation     obesity_type_iii  43.7419
2108  2,109  female   23  1.7522     134  ...  3 to 5   sometimes  public_transportation     obesity_type_iii  43.5438
2109  2,110  female   24  1.7394     133  ...  3 to 5   sometimes  public_transportation     obesity_type_iii  44.0715
2110  2,111  female   24  1.7388     133  ...  3 to 5   sometimes  public_transportation     obesity_type_iii  44.1443

[2111 rows x 19 columns]

删除重复行后DataFrame.info()输出

!!!!!!!!!!!!!!!!!!     Size and shape After preprocess
37512     
(2084, 18)
<bound method DataFrame.info of       Gender  Age  Height  Weight  FHWO  FAVC  FCVC  NCP  CAEC  SMOKE  CH2O  SCC  FAF  TUE  CALC  MTRANS  NObeyesdad      BMI
0          2   21  1.6200      64     2     1     2    3     2      1     2    1    1    2     1       3           2  24.3865
1          2   21  1.5200      56     2     1     3    3     2      2     3    2    4    1     2       3           2  24.2382
2          1   23  1.8000      77     2     1     2    3     2      1     2    1    3    2     3       3           2  23.7654
3          1   27  1.8000      87     1     1     3    3     2      1     2    1    3    1     3       5           3  26.8519
4          1   22  1.7800      90     1     1     2    1     2      1     2    1    1    1     2       3           4  28.3424
...      ...  ...     ...     ...   ...   ...   ...  ...   ...    ...   ...  ...  ...  ...   ...     ...         ...      ...
2106       2   21  1.7107     131     2     2     3    3     2      1     2    1    3    2     2       3           7  44.9015
2107       2   22  1.7486     134     2     2     3    3     2      1     2    1    2    2     2       3           7  43.7419
2108       2   23  1.7522     134     2     2     3    3     2      1     2    1    2    2     2       3           7  43.5438
2109       2   24  1.7394     133     2     2     3    3     2      1     3    1    2    2     2       3           7  44.0715
2110       2   24  1.7388     133     2     2     3    3     2      1     3    1    2    2     2       3           7  44.1443

[2084 rows x 18 columns]

注:info输出显示最后一行索引为2110,但实际行数为2084。

尝试过的代码

# Example of inplace = False and inplace=True
return_df = return_df.drop_duplicates(inplace=False)
return_df.drop_duplicates(inplace=True)

相关删除重复行代码

# removing duplicates
count_dup = return_df.duplicated().sum()
if (verbose > 0):
    print (f"Number of Duplicates : {count_dup}")
if count_dup > 0:
    if (verbose > 0):
        print ("Dropping Duplicates")
    # return_df.drop_duplicates(inplace=True)
    return_df = return_df.drop_duplicates(inplace=False)
else:
    if (verbose > 0):
        print ("No duplicates found.")

return return_df

解决方案

这不是元数据未刷新的问题,drop_duplicates()仅删除重复行,但不会重置DataFrame的索引。原索引会被保留,只是被删除行对应的索引会消失,因此最后一行的索引仍为原最大值2110,但实际有效行数已变为2084。

要让索引与实际行数匹配,只需调用reset_index()方法:

方法1:重置索引并保留原索引为新列

return_df = return_df.drop_duplicates().reset_index()

方法2:重置索引并丢弃原索引(推荐)

若无需保留原索引,添加drop=True参数:

return_df = return_df.drop_duplicates().reset_index(drop=True)

整合到现有代码

修改删除重复行的逻辑:

# removing duplicates
count_dup = return_df.duplicated().sum()
if (verbose > 0):
    print (f"Number of Duplicates : {count_dup}")
if count_dup > 0:
    if (verbose > 0):
        print ("Dropping Duplicates")
    # 删除重复行并重置索引
    return_df = return_df.drop_duplicates().reset_index(drop=True)
else:
    if (verbose > 0):
        print ("No duplicates found.")

return return_df

处理后,DataFrame.info()输出的索引范围会变为0到2083,与实际行数完全匹配。

内容的提问来源于stack exchange,提问作者Jess Stuart

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 06:32:31