Python实现2016/2017年存在水果的2015年国家出现频数统计
Python实现方案
我们可以用pandas数据处理库快速完成需求,具体实现如下:
实现逻辑
- 第一步:导入原始结构化数据
- 第二步:提取2016、2017两年出现过的所有唯一水果品类
- 第三步:筛选出2015年的所有数据,仅保留上述目标水果的相关记录
- 第四步:统计各水果在不同国家的出现次数,未出现的场景默认频数为0
- 第五步:输出行索引为水果、列索引为国家的频数统计表
完整代码
import pandas as pd # 录入原始表格数据 raw_data = [ ["Germany", "Apple", 2015], ["France", "Apple", 2015], ["France", "Apple", 2015], ["Spain", "Apple", 2015], ["Germany", "Banana", 2015], ["France", "Banana", 2015], ["France", "Apple", 2016], ["Spain", "Apple", 2016], ["Germany", "Banana", 2016], ["France", "Banana", 2016], ["France", "Banana", 2017], ["France", "Grapes", 2017] ] df = pd.DataFrame(raw_data, columns=["国家", "水果", "年份"]) # 筛选2016、2017年出现的目标水果 target_fruit_list = df[df["年份"].isin([2016,2017])]["水果"].unique() # 统计2015年各水果在不同国家的出现频数 stat_result = df[df["年份"] == 2015] \ .query("水果 in @target_fruit_list") \ .pivot_table(index="水果", columns="国家", aggfunc="count", values="年份", fill_value=0) \ .reindex(target_fruit_list, fill_value=0) # 输出结果 print(stat_result)
输出结果
| 水果 | Germany | France | Spain |
|---|---|---|---|
| Apple | 1 | 2 | 1 |
| Banana | 1 | 1 | 0 |
| Grapes | 0 | 0 | 0 |
内容的提问来源于stack exchange,提问作者Aman
相关产品推荐
相关产品推荐

