You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

数据分组问题求助:按年份统计男女毕业生Top3课程类型

解决按年份分组统计男女Top3课程类型的问题

核心思路

  1. 获取完整数据集:原代码的limit=100仅能获取部分数据,需先请求总记录数再获取全部数据
  2. 分组统计:使用嵌套字典结构按年份→性别→课程类型的层级累加毕业生数量
  3. 排序取Top3:对每个性别下的课程类型按毕业生数量降序排序,提取前3个课程名称
  4. 格式化输出:按照指定格式打印每年的结果

完整代码

import requests
from collections import defaultdict, Counter

# 数据接口基础配置
API_URL = "https://data.gov.sg/api/action/datastore_search"
RESOURCE_ID = "eb8b932c-503c-41e7-b513-114cffbe2338"

# 第一步:获取总记录数,确保能拿到全部数据
response = requests.get(API_URL, params={"resource_id": RESOURCE_ID})
total_records = response.json()["result"]["total"]

# 第二步:获取所有毕业生数据
full_response = requests.get(API_URL, params={
    "resource_id": RESOURCE_ID,
    "limit": total_records
})
records = full_response.json()["result"]["records"]

# 第三步:构建分组统计字典
# 结构:year -> sex -> {course_type: total_graduates}
year_stat = defaultdict(lambda: defaultdict(Counter))

for record in records:
    year = record["year"]
    sex = record["sex"]
    course = record["type_of_course"]
    # 累加毕业生数量(注意转换为整数)
    grad_count = int(record["number_of_graduates"])
    year_stat[year][sex][course] += grad_count

# 第四步:按年份排序并输出Top3课程
for year in sorted(year_stat.keys()):
    print(year)
    
    # 处理男性Top3
    male_courses = sorted(year_stat[year]["Males"].items(), key=lambda x: -x[1])[:3]
    male_top3 = [course for course, _ in male_courses]
    print(f"Males: {' | '.join(male_top3)}")
    
    # 处理女性Top3
    female_courses = sorted(year_stat[year]["Females"].items(), key=lambda x: -x[1])[:3]
    female_top3 = [course for course, _ in female_courses]
    print(f"Females: {' | '.join(female_top3)}")
    
    print()  # 空行分隔不同年份结果

关键说明

  • 统计逻辑:必须使用number_of_graduates字段累加,每条记录代表对应维度的毕业生总数,而非单条记录计数
  • 数据完整性:通过先获取总记录数再设置limit,避免分页导致的数据缺失
  • 排序规则:按毕业生数量降序排序,确保Top3是人数最多的课程类型

内容的提问来源于stack exchange,提问作者Divakar KN

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 06:15:33