You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python合并两个CSV文件并转换为指定嵌套字典

问题:将合并后的学生DataFrame转换为指定嵌套字典结构

有两个CSV数据集:

  • student:包含name、matricNo等学生基本信息
  • Records:包含title、unit等选课成绩信息

已通过Python读取为DataFrame并完成合并,但转换为字典时无法得到预期的嵌套结构,预期输出结构如下:

student_records = {
      "name": "James Webb",
      "id": "201003",
      "courses":[
         {"title": "English", "unit": 2, "code": "ENG101", "score": 60, "term": "Fall", "session":'2020'},
         {"title": "Chemistry I", "unit": 4, "code": "CHE101", "score": 70, "term": "Fall", "session":'2021'},
         {"title": "Maths", "unit": 3, "code": "MTH101", "score": 80, "term": "Spring", "session":'2020'},
         {"title": "Chemistry II", "unit": 4, "code": "CHE102", "score": 91, "term": "Spring", "session":'2021'},
         {"title": "History", "unit": 2, "code": "HIS102", "score": 40, "term": "Spring", "session":'2020'}
      ]
   }

解决方案

假设合并后的DataFrame名为merged_df,可以通过分组+字典转换实现预期结构:

import pandas as pd

# 定义学生基本信息列和课程信息列
student_info_cols = ['name', 'matricNo']
course_info_cols = ['title', 'unit', 'code', 'score', 'term', 'session']

# 按学生维度分组,构建目标结构
student_records_list = []
for (student_name, matric_no), course_group in merged_df.groupby(student_info_cols):
    # 将当前学生的所有课程记录转为字典列表
    courses = course_group[course_info_cols].to_dict('records')
    # 组装单个学生的完整字典
    single_student = {
        'name': student_name,
        'id': matric_no,
        'courses': courses
    }
    student_records_list.append(single_student)

# 若仅需单个学生的记录(如示例中的James Webb),可从列表中提取
# target_student = next(item for item in student_records_list if item['name'] == 'James Webb')

代码说明

  • 使用groupby按学生的name和matricNo分组,保证每个分组对应一个学生的全部选课记录
  • 利用to_dict('records')将分组后的课程数据直接转换为字典列表,这就是courses字段的内容
  • 最后将学生基本信息与课程列表组合,得到符合要求的嵌套字典结构

内容的提问来源于stack exchange,提问作者gabidoye

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 03:01:38