求助:Python Pandas实现CSV文件各列单独降序排序
问题
尝试用Python的Pandas库对CSV文件的每一列单独按从高到低排序,但运行现有代码后,只有第一列完成排序,其余列会跟着对应行的数据移动,没法实现各列独立排序(比如原数据中B列最高值是98,并没有被单独排序到该列的顶部)。
现有代码如下:
import pandas as pd bus1_data = pd.read_csv("p1Bus1Data.csv",header=None) bus1_data.rename(columns={0: 'quantity bids', 1: 'price of bids', 2: 'quantity of offers', 3: 'price of offers'}, inplace=True) bus1_data.to_csv('test_with_col.csv', index=False) bus2_data = pd.read_csv("p1Bus2Data.csv",header=None) bus2_data.rename(columns={0: 'quantity bids', 1: 'price of bids', 2: 'quantity of offers', 3: 'price of offers'}, inplace=True) bus2_data.to_csv('test_with_col.csv', index=False) sorted_bus1_data = bus1_data.sort_values(['quantity bids','price of bids','quantity of offers','price of offers'], ascending=[False,False,False,False]) sorted_bus1_data.to_csv('p1Bus1Data_sorted.csv', index=False) sorted_bus2_data = bus2_data.sort_values(by=["quantity bids", "price of bids", "quantity of offers", "price of offers"],ascending=[False,False,False,False]) sorted_bus2_data.to_csv('p1Bus2Data_sorted.csv', index=False) print(sorted_bus1_data.head()) print(sorted_bus2_data.head())
解决方案
你当前使用的sort_values()方法是按指定列的优先级对整行数据进行排序,排序时行内的所有数据会保持关联,因此无法实现各列独立排序的效果。要让每列单独降序排序,需要用apply()方法对DataFrame的每一列单独处理:
修改后的代码如下:
import pandas as pd # 处理bus1数据 bus1_data = pd.read_csv("p1Bus1Data.csv", header=None) bus1_data.rename(columns={0: 'quantity bids', 1: 'price of bids', 2: 'quantity of offers', 3: 'price of offers'}, inplace=True) # 对每一列单独降序排序并重置索引 sorted_bus1_data = bus1_data.apply(lambda col: col.sort_values(ascending=False).reset_index(drop=True)) sorted_bus1_data.to_csv('p1Bus1Data_sorted.csv', index=False) # 处理bus2数据 bus2_data = pd.read_csv("p1Bus2Data.csv", header=None) bus2_data.rename(columns={0: 'quantity bids', 1: 'price of bids', 2: 'quantity of offers', 3: 'price of offers'}, inplace=True) sorted_bus2_data = bus2_data.apply(lambda col: col.sort_values(ascending=False).reset_index(drop=True)) sorted_bus2_data.to_csv('p1Bus2Data_sorted.csv', index=False) print(sorted_bus1_data.head()) print(sorted_bus2_data.head())
代码说明
apply(lambda col: col.sort_values(ascending=False)):遍历DataFrame的每一列,对每一列单独执行降序排序操作reset_index(drop=True):排序后原行索引会被保留,用该方法重置索引,让每列的排序结果从0开始连续编号,避免索引混乱
内容的提问来源于stack exchange,提问作者Juan Garcia
相关产品推荐
相关产品推荐

