如何按天数统计对Pandas DataFrame的国家排序并返回字符串数组
按国家天数总和排序生成字符串数组
原始DataFrame数据
| date | days | country |
|---|---|---|
| 2022-02-01 | 1 | Spain |
| 2022-02-02 | 2 | Spain |
| 2022-02-01 | 3 | Italy |
| 2022-02-03 | 2 | France |
| 2022-02-03 | 1 | Germany |
| 2022-02-04 | 1 | Italy |
| 2022-02-04 | 1 | UK |
| 2022-02-05 | 2 | UK |
| 2022-02-04 | 5 | Spain |
| 2022-02-04 | 1 | Portugal |
需求说明
按国家汇总days列的总和,按总和从多到少排序,最终生成对应的国家字符串数组。
统计结果
| country | count_days |
|---|---|
| Spain | 8 |
| Italy | 4 |
| UK | 3 |
| France | 2 |
| Germany | 1 |
| Portugal | 1 |
Python实现代码
import pandas as pd # 构造原始DataFrame data = { 'date': ['2022-02-01', '2022-02-02', '2022-02-01', '2022-02-03', '2022-02-03', '2022-02-04', '2022-02-04', '2022-02-05', '2022-02-04', '2022-02-04'], 'days': [1, 2, 3, 2, 1, 1, 1, 2, 5, 1], 'country': ['Spain', 'Spain', 'Italy', 'France', 'Germany', 'Italy', 'UK', 'UK', 'Spain', 'Portugal'] } df = pd.DataFrame(data) # 按国家分组求和并排序,提取目标数组 country_summary = df.groupby('country')['days'].sum().reset_index(name='count_days') sorted_summary = country_summary.sort_values(by='count_days', ascending=False) countries = sorted_summary['country'].tolist() print(countries) # 输出: ['Spain', 'Italy', 'UK', 'France', 'Germany', 'Portugal']
内容的提问来源于stack exchange,提问作者Carola
相关产品推荐
相关产品推荐

