You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将含4万条记录的试题数据库表拆分为试题表与答案表?

最简实现方案

核心思路

用Python结合pandas库批量处理,代码量少且高效,适配4万条数据的规模,能快速完成表拆分与JSON导出。

具体步骤

  • 将原数据库表导出为CSV文件(直接操作数据库也可,CSV更直观易调试)
  • 编写简短脚本读取数据,拆分生成试题表和答案表结构,最后导出JSON

示例代码

import pandas as pd
import json

# 读取原数据CSV(直接读数据库可用pandas.read_sql替代)
df = pd.read_csv('original_questions.csv')

# 生成试题表数据:按questionid去重,保留id、questionid、question字段
question_data = df[['id', 'questionid', 'question']].drop_duplicates(subset='questionid')
question_data.to_json('questions.json', orient='records', force_ascii=False)

# 生成答案表数据
answer_list = []
current_answer_id = 1

for _, row in df.iterrows():
    q_id = row['questionid']
    correct_label = row['correct answer'].upper()
    # 遍历四个选项,生成对应答案记录
    for ans_col in ['answer a', 'answer b', 'answer c', 'answer d']:
        ans_content = row[ans_col]
        # 判断当前选项是否为正确答案
        is_correct = 1 if ans_col.split()[-1].upper() == correct_label else 0
        answer_list.append({
            'id': current_answer_id,
            'question ID': q_id,
            'answer': ans_content,
            'correct answer': is_correct
        })
        current_answer_id += 1

# 导出答案表JSON
with open('answers.json', 'w', encoding='utf-8') as f:
    json.dump(answer_list, f, ensure_ascii=False, indent=2)

补充说明

  • 若直接对接数据库,可替换pd.read_csv为pd.read_sql,传入数据库连接即可
  • 脚本自动处理所有数据,每道题生成4条答案记录,完全匹配需求
  • 导出的JSON格式规范,每条记录结构统一

内容的提问来源于stack exchange,提问作者user

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 16:12:14