数据分组问题求助:按年份统计男女毕业生Top3课程类型
解决按年份分组统计男女Top3课程类型的问题
核心思路
- 获取完整数据集:原代码的
limit=100仅能获取部分数据,需先请求总记录数再获取全部数据 - 分组统计:使用嵌套字典结构按年份→性别→课程类型的层级累加毕业生数量
- 排序取Top3:对每个性别下的课程类型按毕业生数量降序排序,提取前3个课程名称
- 格式化输出:按照指定格式打印每年的结果
完整代码
import requests from collections import defaultdict, Counter # 数据接口基础配置 API_URL = "https://data.gov.sg/api/action/datastore_search" RESOURCE_ID = "eb8b932c-503c-41e7-b513-114cffbe2338" # 第一步:获取总记录数,确保能拿到全部数据 response = requests.get(API_URL, params={"resource_id": RESOURCE_ID}) total_records = response.json()["result"]["total"] # 第二步:获取所有毕业生数据 full_response = requests.get(API_URL, params={ "resource_id": RESOURCE_ID, "limit": total_records }) records = full_response.json()["result"]["records"] # 第三步:构建分组统计字典 # 结构:year -> sex -> {course_type: total_graduates} year_stat = defaultdict(lambda: defaultdict(Counter)) for record in records: year = record["year"] sex = record["sex"] course = record["type_of_course"] # 累加毕业生数量(注意转换为整数) grad_count = int(record["number_of_graduates"]) year_stat[year][sex][course] += grad_count # 第四步:按年份排序并输出Top3课程 for year in sorted(year_stat.keys()): print(year) # 处理男性Top3 male_courses = sorted(year_stat[year]["Males"].items(), key=lambda x: -x[1])[:3] male_top3 = [course for course, _ in male_courses] print(f"Males: {' | '.join(male_top3)}") # 处理女性Top3 female_courses = sorted(year_stat[year]["Females"].items(), key=lambda x: -x[1])[:3] female_top3 = [course for course, _ in female_courses] print(f"Females: {' | '.join(female_top3)}") print() # 空行分隔不同年份结果
关键说明
- 统计逻辑:必须使用
number_of_graduates字段累加,每条记录代表对应维度的毕业生总数,而非单条记录计数 - 数据完整性:通过先获取总记录数再设置
limit,避免分页导致的数据缺失 - 排序规则:按毕业生数量降序排序,确保Top3是人数最多的课程类型
内容的提问来源于stack exchange,提问作者Divakar KN
相关产品推荐
相关产品推荐

