You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对Pandas DataFrame每行的列单独按降序重新排序?

问题:对城镇选举数据按行降序排序得票率

我有一份已清理完成的地方选举数据,每行代表一个城镇,每列代表一位候选人的得票率。列名按候选人在选票上的出现顺序命名,而非得票多少排序。我希望将每行的列按得票率从高到低排序,让每行首列对应得票最高的候选人。

原始DataFrame示例:

town         cand_1 cand_2  cand_3   cand_4   

town_1       0.64    0.24    0.12    NaN   
town_2       0.21    0.47    0.31    0.01    
town_3       0.08    0.12    0.75    0.05    
town_4       0.56    0.30    0.14    NaN   

可以看到,town_3中cand_3得票最高,town_2中cand_2得票最高,但它们在对应行中分别处于第三和第二列。

期望输出:

town         cand_1 cand_2  cand_3   cand_4   

town_1       0.64    0.24    0.12    NaN   
town_2       0.47    0.31    0.21    0.01    
town_3       0.75    0.12    0.08    0.05    
town_4       0.56    0.30    0.14    NaN   

我曾尝试在Excel中操作——选中每行后升序排序,但效果不理想。


编辑说明

有评论指出列名现在失去意义,特此说明:我仅关注城镇的得票分布,不关心cand_1、cand_2等标识(不过我确实关心获胜候选人的身份,这来自另一个数据集)。因此输出DataFrame的列名可以更改;或者cand_1指代对应城镇的得票最高者。

补充说明

我找到了一个使用Python内置函数的解决方案,还能获取获胜候选人的姓名/党派:

hh=[]

for i in results_2008_list:
    cands=[i[4], i[7], i[10], i[13], i[16], i[19], i[22],
                              i[25], i[28], i[31], i[34]]
    hh.append([i[2], i[3], max(cands), sorted(cands, reverse=True)[1],
              sorted(cands, reverse=True)[2], sorted(cands, reverse=True)[3],
              sorted(cands, reverse=True)[4], i[i.index(max(cands))+1],
              i[i.index(max(cands))+2]])

解决方案

方法一:使用Pandas实现行内降序排序

如果数据是Pandas DataFrame格式,用apply结合sorted就能高效实现,同时保留城镇列:

import pandas as pd

# 构造示例DataFrame
df = pd.DataFrame({
    'town': ['town_1', 'town_2', 'town_3', 'town_4'],
    'cand_1': [0.64, 0.21, 0.08, 0.56],
    'cand_2': [0.24, 0.47, 0.12, 0.30],
    'cand_3': [0.12, 0.31, 0.75, 0.14],
    'cand_4': [None, 0.01, 0.05, None]
})

# 对除town列外的每行做降序排序,补全NaN
sorted_df = df.set_index('town').apply(
    lambda x: sorted(x.dropna(), reverse=True) + [None]*(len(x)-x.count()), 
    axis=1
).reset_index()

# 可选:按排序位次重命名列
sorted_df.columns = ['town', 'top_cand', 'second_cand', 'third_cand', 'fourth_cand']

print(sorted_df)

输出结果与期望一致,且比循环列表更适合大规模数据。

方法二:优化现有列表处理方案

你的代码可以简化,避免重复调用sorted,提升效率:

hh = []
for i in results_2008_list:
    # 提取得票率列表
    cands = [i[4], i[7], i[10], i[13], i[16], i[19], i[22], i[25], i[28], i[31], i[34]]
    # 仅排序一次,复用结果
    sorted_cands = sorted(cands, reverse=True)
    # 获取最高得票对应的候选人信息
    max_idx = cands.index(max(cands))
    cand_info = [i[max_idx + 1], i[max_idx + 2]]
    # 组装最终结果
    hh.append([i[2], i[3]] + sorted_cands[:5] + cand_info)

内容的提问来源于stack exchange,提问作者mw1021212

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 18:36:25