Pandas .sort_values()排序后DataFrame值错乱、report_id消失问题求助
Pandas排序后数据列错乱、report_id消失问题
我用Pandas加载包含report_id、when、what三列的short_desc.csv,先执行了以下预处理代码:
import pandas as pd # 读取csv文件 shortDesc = pd.read_csv('short_desc.csv') # 筛选report_id为数字且非空的记录 shortDesc = shortDesc[shortDesc['report_id'].str.isdigit().notnull()] # 将when列的UNIX时间戳转换为datetime格式 shortDesc['when'] = pd.to_datetime(shortDesc['when'], unit='s')
预处理完成后,数据显示一切正常。
之后我想通过按when列排序再去重的方式,保留每个report_id对应的最新记录,代码如下:
shortDesc = shortDesc.sort_values(by='when').drop_duplicates(['report_id'], keep='last')
但单独调用shortDesc.sort_values(by=['when'], inplace=False)时,出现了异常:
what列的数值出现跨列错乱的情况report_id列的值直接消失了
奇怪的是,另一个结构相同的DataFrame(仅删除了what列)使用完全相同的代码,却能得到正确结果。
内容的提问来源于stack exchange,提问作者Master Oogway
相关产品推荐
相关产品推荐

