如何同时从pandas DataFrame取最大值及索引并优化三类统计代码
优化后完整可运行代码
import pandas as pd import numpy as np # 示例DataFrame构造,你可以替换为自己的数据源 data = [ ["USA",2021,1500],["USA",2018,6000],["India",2019,3000], ["India",2021,5000],["UK",2019,4000],["USA",2019,3200],["India",2018,5000] ] df = pd.DataFrame(data, columns=["Country","Year","Count"]) # 1. 输出全局Count最大的条目 max_total_row = df.nlargest(1, 'Count').iloc[0] print(f"Entry with Max count is ({max_total_row['Country']}, {max_total_row['Year']}, {max_total_row['Count']})") # 2. 输出总Count最高的国家 country_sum = df.groupby('Country')['Count'].sum() max_country = country_sum.nlargest(1).reset_index().iloc[0] print(f"\nCountry with max total count is: ({max_country['Country']}, {max_country['Count']})") # 3. 输出每年Count最大的条目 year_max_rows = df.groupby('Year', group_keys=False).apply(lambda x: x.nlargest(1, 'Count')) print("\nEntry with max count in each year is:") for _, row in year_max_rows.iterrows(): print(f"{row['Year']}, {row['Country']}, {row['Count']}")
问题1解答:是否可以同时获取最大值和对应行,无需先查索引再取行?
可以,直接使用pandas内置的nlargest(返回行数, 统计列名)方法,该方法会直接返回指定列数值最大的前N行,不需要单独计算最大值索引再查询对应行,代码更简洁,大数据量下性能也优于先找索引再取行的写法。
如果不想用nlargest,也可以用布尔索引一行实现:max_row = df[df['Count'] == df['Count'].max()].iloc[0],效果完全一致。
问题2解答:更简洁易读的实现逻辑
优化主要做了三点调整:
- 用
groupby('列名')['统计列'].sum()替代冗余的pivot_table写法,求和逻辑更直观 - 用groupby+apply+nlargest直接批量计算每年的最大值行,替代手动循环遍历分组的写法,减少重复代码
- 所有输出用f-string格式化,替代繁琐的字符串拼接,可读性大幅提升
内容的提问来源于stack exchange,提问作者Arvind Swaminathan
相关产品推荐
相关产品推荐

