You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将值为可变长度列表的Python字典写入长格式CSV文件

值为列表的字典转长格式CSV实现方法

核心逻辑

长格式要求每个NCT编号和对应疾病列表内的单个疾病组成独立行,只需要两层遍历字典:外层遍历字典的键值对,内层遍历值列表的每个元素,逐行写入即可。


方案1:使用Python内置csv模块(无额外依赖)

不需要安装第三方库,兼容性最好,适合所有Python环境:

import csv

diseaseDict = {'NCT01266330': ['Oxidative Stress', 'Inflammation'], 
    'NCT01266343': ['Glaucoma', 'Acute Primary Angle-closure Glaucoma'], 
    'NCT01266356': ['Traumatic Brain Injury'], 
    'NCT01266369': ['Mastocytosis'], 
    'NCT01266382': ['Osteoarthritis', 'Spinal Diseases', 'Ligament Rupture', 'Lower Extremity Fracture', 'Neurological Disorders'], 
    'NCT01266395': ['COPD']}

with open('disease_long.csv', 'w', newline='', encoding='utf-8') as csvfile:
    writer = csv.writer(csvfile)
    # 写入表头
    writer.writerow(['nct_id', 'disease'])
    # 逐行拆分写入
    for nct, disease_list in diseaseDict.items():
        for single_disease in disease_list:
            writer.writerow([nct, single_disease])

注意点:打开文件时必须传入newline='',这是csv模块官方要求的参数,避免Windows系统下生成的CSV出现多余空行;指定encoding='utf-8'可避免特殊字符乱码。


方案2:使用pandas快速实现(适合数据处理场景)

如果日常用pandas做数据清洗,代码更简洁:

import pandas as pd

diseaseDict = {'NCT01266330': ['Oxidative Stress', 'Inflammation'], 
    'NCT01266343': ['Glaucoma', 'Acute Primary Angle-closure Glaucoma'], 
    'NCT01266356': ['Traumatic Brain Injury'], 
    'NCT01266369': ['Mastocytosis'], 
    'NCT01266382': ['Osteoarthritis', 'Spinal Diseases', 'Ligament Rupture', 'Lower Extremity Fracture', 'Neurological Disorders'], 
    'NCT01266395': ['COPD']}

# 生成扁平化的长格式记录
rows = [(nct, disease) for nct, disease_list in diseaseDict.items() for disease in disease_list]
df = pd.DataFrame(rows, columns=['nct_id', 'disease'])
# 写出CSV,取消索引列
df.to_csv('disease_long.csv', index=False, encoding='utf-8')

最终输出效果

生成的CSV文件内容如下,完全符合长格式要求:

nct_id,disease
NCT01266330,Oxidative Stress
NCT01266330,Inflammation
NCT01266343,Glaucoma
NCT01266343,Acute Primary Angle-closure Glaucoma
NCT01266356,Traumatic Brain Injury
NCT01266369,Mastocytosis
NCT01266382,Osteoarthritis
NCT01266382,Spinal Diseases
NCT01266382,Ligament Rupture
NCT01266382,Lower Extremity Fracture
NCT01266382,Neurological Disorders
NCT01266395,COPD

常见错误原因

之前代码运行结果不对,通常是踩了以下几个坑:

  • 只做了一层循环,直接把整个疾病列表作为单个值写入单元格,没有拆分列表内的每个元素
  • 写文件时没加newline=''参数,导致CSV出现多余空行
  • 用pandas时直接调用pd.DataFrame.from_dict()生成的是宽格式(每个疾病占单独列),没有做扁平化转换
  • 未指定文件编码,遇到特殊字符时出现乱码

内容的提问来源于stack exchange,提问作者Jose R

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 08:27:20